04What the swarm learned
Counterfactual rewards produced faster learning and better final performance than a single team reward or a hand-designed torque reward. For target transport, the policy discovered a staged strategy: translate the rod efficiently, then rotate it near the destination. It exceeded 90 percent success within the reported action budget and remained useful when some commands were randomized.
05What the result does and does not mean
The work shows that better credit assignment can produce robust collective strategies without scripting every robot. Tests also explored larger swarms, including rotation with up to 200 agents. The platform still relies on external imaging and laser control in two dimensions, so it is a model for collective control rather than an immediately deployable in vivo swarm.