swarm-lab
hackathon report
AI Village x Grove Research, AI Swarm Dynamics Hackathon, Oct 3 to 4 2026. Team dmarz, vishesh, shadow. This page is the shadow side's read of the whole weekend: what happened hour by hour, what the experiments found, what Sol's lanes shipped, what went wrong, and what comes next. Every number links to the file it came from.
Repo dmarzzz/swarm-lab · research dashboard swarm-research.pages.dev · live experiment monitor swarm-live.pages.dev
1. Timeline commits per hour on main, UTC, by team
Team is read from the agent id prefix in the commit subject ([shadow/...], [dmarz/...], [vishesh/...]), because vishesh's agents commit under several git names. Bot index rebuilds are shown in grey. Hover a bar for the busiest lanes that hour. The line is cumulative library entries added.
shadow / Sol lanesdmarz fleetvishesh fleetCI botlibrary entries (cumulative)
What happened
2. Findings all three researchers
All results are exploratory, from synthetic task worlds, mostly with one model in the loop. Calls and identities are not independent samples; the unit is in the n column. Status: finding held up and has saved records; lead real but incomplete, scripted or engineering-level; negative valid adverse result; does-not-generalize held on one model, failed on others.
Five of the dmarz rows were independently recomputed from saved records by Sol's xcheck lane with its own code (no model calls): completed-findings-xcheck. 7,512 recorded answers re-scored, zero endpoint mismatches; one small tie-counting defect in the scaling cross-model counts. dmarz's latest-results review and vishesh's PI review (57 findings) are the cross-checks the rest of this table leans on.
3. What the shadow side shipped
Infrastructure
- Collector pipeline (PR #1): X, LessWrong, RSS, YouTube and web search hits deduplicated against the library and published as small claimable GitHub issues. 68 batches worked across both nights.
- Transcription: YouTube captions, then yt-dlp audio plus Deepgram diarised transcripts for talks and podcasts, so talks became read-in-full library entries.
- Dashboards: swarm-research.pages.dev (library graph, X threads view, rebuilt by CI on every push and every 15 min), styled to match dmarz's swarm-live.
- Swarm Lab Discord: channels, roles, a per-team 15-minute repo digest instead of a raw webhook firehose, and review-request pings for inbox items.
Experiments
- Hypotheses (PR #82): N_eff evidence board, capture vs memory, board N-sweep. Proposed, not accepted.
- capture-memory (PR #83): scripted S0, S1, S1b, 48,600 episodes, zero spend.
- capture-memory-mix: 30,000 scripted episodes plus real-model pilots on three models, USD 3.93 under a USD 10 cap enforced in code. Result does not generalize across models; reported as such.
- sol-cm2 (running): corrections for the defects dmarz flagged, and a preregistered Claude Sonnet 4.6 replication of the core cells.
Reviews and tracing
- Discussion benchmark v3 review: pass-with-fixes, two blocking defects (vote metrics zeroed when any ballot is invalid; provider-failure reason dropped) at the commit pinned for a paid run.
- Completed-findings xcheck: five dmarz findings recomputed independently.
- Agent trace spec: TRACE-SPEC v0.2.0 on branch shadow/rsi (envelope vs content split, disclosure tiers private / sealed / public, OTel GenAI conventions).
- rsi-loop R0: offline replay over saved sessions; candidate parser recognised 40/40 nonzero tool statuses vs 0/40 baseline, 0/702 false flags, zero model calls. Engineering only, no research-quality effect measured.
- Submission draft: WRITEUP, DEMO, RESULTS and HACKATHON.md. Not submitted; the humans submit.
Sunday lanes, status at last refresh
Read from the task files and the repo at the commit above. "on main" means the lane's main output file exists on main.
| lane | what | task | output on main |
|---|
4. Problems, honestly
- Reporting defects in our own study. dmarz's review caught that capture-memory-mix's Qwen table selected 180 of 327 raw records (127 invalid) by last-valid selection without saying so, that the README mentions redos while a table footer says no retries, that a caption said round 50 while the config says 30, and that one hub capture fraction exceeded 1. The verdict "mixture rescue does not generalize" is right and we accept it. Corrections are in progress.
- Gateway restarts and overload. A restart at 05:03Z killed the first real-model pilot mid-run (resumed 05:45Z, nothing lost but time). Sunday 13:45 to 14:00Z the orchestrator gateway was overloaded and lane spawns timed out while often landing anyway, so every spawn had to be verified by hand to avoid duplicates. Lesson: checkpoint every paid run and make spawns idempotent.
- Lanes timing out. Several Saturday writer lanes hit runaway web-fetch loops (one used 9.4M input tokens) and timed out mid-rebase. Each was rescued by hand: conflicts resolved, untracked entries committed, claims released. Lesson: cap fetch retries per lane.
- Ledger mishap. A
git stash with live workers briefly reverted and deleted the spend ledger. Five episodes were recorded invalid and kept as invalid; the ledger was reconciled against the provider's account usage. Lesson: ledgers are written by workers and should never be tracked in a tree you stash in.
- Commit identity slip. A wrong noreply email in the git config attributed 19 early commits to an unrelated GitHub account. Found and fixed the same afternoon by rewriting author emails only (trees compared).
- A stray bot post in #general. An orchestrator error message ("exec denied") was posted into the shared Swarm Lab #general channel by mistake. Lesson: lanes never post to shared channels; reports go to the orchestrator only.
- Keys pasted in chat. Over the weekend several API keys and tokens were pasted in plaintext into Discord channels, by more than one person. None were committed to the repo. They are stored locally and flagged for rotation after the event. Lesson: use a secrets channel with a one-time-share tool from the start, and rotate everything shared in chat.
- Rate limits. Semantic Scholar 429'd both teams' boxes and YouTube bot-checked ours; worked around with OpenAlex and a paid transcript fallback under a small cap.
- Synthetic, not in the wild. No submitted result uses the AI Village, collusion.wiki, SwarmTraces or Transluce data the organisers pointed at. Three of those datasets were downloaded locally on Saturday but not analysed.
5. Next steps
Today, before 00:00Z
- Humans submit once through the form in the event Slack, using the submission draft; fill in names and emails. Draft refresh at about 20:00Z.
- Land the capture-memory-mix corrections and whatever the Claude replication produces, reported as is.
- Open the rsi-loop PR from shadow/rsi with one replay-accepted and one replay-rejected bundle.
- Land the fork-merge-security and agent-budgets surveys and the landscape map through the gate.
- Demo: 2 minutes from one shared machine; push everything to GitHub first. No new model calls after 22:00Z.
After the hackathon
- Trace market (RSI). Grow TRACE-SPEC into the thing that lets agents sell and verify improvements: searcher proposes a bundle, builder evaluates by replay against frozen records, a credit ledger pays what replays accept. R0 shows the replay harness works on real records; next is a bundle that changes a research outcome, judged by someone else.
- Extend the Sybil family where it is thinnest: nonredundant specialist facts, fixed attacker resources with varying identities, admission policies compared at equal cost, independent graph families. Report accuracy, attacker share and honest retention together.
- Market splitting: independent market structures and prompts without the registration action in plain view.
- Lineage and copying (Quorum of Mirrors, Phantom Coast): exact lineage deduplication as a baseline, coverage guards.
- Memory and capture: treat the long-window read shape as the hypothesis, preregister it, test on models chosen in advance.
- Go in the wild: run the copying and N_eff measures on collusion.wiki and SwarmTraces.
- Ops: rotate every key shared in chat; per-lane fetch caps; idempotent spawns; checkpointed paid runs.