Evidence synthesis
What the research says when the sources are read together.
The evidence does not support a blanket claim that collaboration, more agents, or more automation is better. It supports a conditional model: divide work by function, preserve visible human authority, ground participants with sufficient evidence, measure the whole workflow, and require the cooperative path to beat the strongest relevant baseline.
The central argument
Collaborative intelligence is a designed relationship, not a model feature.
The strongest recurring signal is that results depend on task shape, division of responsibility, information access, calibration, and recovery. A capable model can still produce a weak system when people cannot see its mode, challenge its premise, recover from failure, or distinguish a recommendation from an authorized action.
Five evidence trails
Where the literature converges, and where it does not.
1 · Outcomes and burden
Collaboration earns its place through comparison.
Human-AI and multi-agent arrangements can improve bounded outcomes, but added roles, messages, and review also add cost. Average quality can hide lost diversity. The right question is whether collaboration beat the strongest practical alternative at acceptable cost and risk.
2 · Workflow evidence
The path matters as much as the final answer.
Generation, selection, review, correction, and closeout can fail separately. Privacy-minimal traces help locate the failure and test recovery. Prior evidence should expire when models, tools, prompts, policy, data, or participants materially change.
3 · Authority and reliance
Human control must be visible and usable.
Stated trust and explanation presence are weak substitutes for observed behavior. A person needs to know the current mode, the limits of an explanation, when to verify, and how to take over. Authority should be separated across information gathering, analysis, recommendation, decision, and action.
4 · Memory and grounding
Useful memory is governed evidence, not accumulated text.
Recall alone does not establish safe memory. Systems also need updates, abstention, time, provenance, correction, custody, isolation, and deletion. Shared grounding should carry the objective, constraints, ownership, and correction state without copying every source record to every participant.
5 · Diversity and alternatives
There is no single best collaborative architecture.
Centralized orchestration, peer coordination, human-led review, specialist deferral, and single-agent work solve different problems. The evidence favors keeping alternatives visible and testing topology against the task, especially when repeated assistance could narrow the range of outputs.
Representative starting points
Follow the evidence behind the synthesis.
Evidence boundary. This is a claim-centered synthesis of the screened catalog, not a meta-analysis or a full-text risk-of-bias review. The claim register records confidence dimensions, null cases, transfer limits, and next tests. Source count is not treated as independent support when papers share an evidence family.