swarm-lab
hackathon report

AI Village x Grove Research, AI Swarm Dynamics Hackathon, Oct 3 to 4 2026. Team dmarz, vishesh, shadow. This page is the shadow side's read of the whole weekend: what happened hour by hour, what the experiments found, what Sol's lanes shipped, what went wrong, and what comes next. Every number links to the file it came from.

Repo dmarzzz/swarm-lab · research dashboard swarm-research.pages.dev · live experiment monitor swarm-live.pages.dev

1. Timeline commits per hour on main, UTC, by team

Team is read from the agent id prefix in the commit subject ([shadow/...], [dmarz/...], [vishesh/...]), because vishesh's agents commit under several git names. Bot index rebuilds are shown in grey. Hover a bar for the busiest lanes that hour. The line is cumulative library entries added.

shadow / Sol lanesdmarz fleetvishesh fleetCI botlibrary entries (cumulative)

Commits by git author

By team

What happened

2. Findings all three researchers

All results are exploratory, from synthetic task worlds, mostly with one model in the loop. Calls and identities are not independent samples; the unit is in the n column. Status: finding held up and has saved records; lead real but incomplete, scripted or engineering-level; negative valid adverse result; does-not-generalize held on one model, failed on others.

resultmodelnstatusby

Five of the dmarz rows were independently recomputed from saved records by Sol's xcheck lane with its own code (no model calls): completed-findings-xcheck. 7,512 recorded answers re-scored, zero endpoint mismatches; one small tie-counting defect in the scaling cross-model counts. dmarz's latest-results review and vishesh's PI review (57 findings) are the cross-checks the rest of this table leans on.

3. What the shadow side shipped

Library and surveys

    Infrastructure

    • Collector pipeline (PR #1): X, LessWrong, RSS, YouTube and web search hits deduplicated against the library and published as small claimable GitHub issues. 68 batches worked across both nights.
    • Transcription: YouTube captions, then yt-dlp audio plus Deepgram diarised transcripts for talks and podcasts, so talks became read-in-full library entries.
    • Dashboards: swarm-research.pages.dev (library graph, X threads view, rebuilt by CI on every push and every 15 min), styled to match dmarz's swarm-live.
    • Swarm Lab Discord: channels, roles, a per-team 15-minute repo digest instead of a raw webhook firehose, and review-request pings for inbox items.

    Experiments

    • Hypotheses (PR #82): N_eff evidence board, capture vs memory, board N-sweep. Proposed, not accepted.
    • capture-memory (PR #83): scripted S0, S1, S1b, 48,600 episodes, zero spend.
    • capture-memory-mix: 30,000 scripted episodes plus real-model pilots on three models, USD 3.93 under a USD 10 cap enforced in code. Result does not generalize across models; reported as such.
    • sol-cm2 (running): corrections for the defects dmarz flagged, and a preregistered Claude Sonnet 4.6 replication of the core cells.

    Reviews and tracing

    • Discussion benchmark v3 review: pass-with-fixes, two blocking defects (vote metrics zeroed when any ballot is invalid; provider-failure reason dropped) at the commit pinned for a paid run.
    • Completed-findings xcheck: five dmarz findings recomputed independently.
    • Agent trace spec: TRACE-SPEC v0.2.0 on branch shadow/rsi (envelope vs content split, disclosure tiers private / sealed / public, OTel GenAI conventions).
    • rsi-loop R0: offline replay over saved sessions; candidate parser recognised 40/40 nonzero tool statuses vs 0/40 baseline, 0/702 false flags, zero model calls. Engineering only, no research-quality effect measured.
    • Submission draft: WRITEUP, DEMO, RESULTS and HACKATHON.md. Not submitted; the humans submit.

    Sunday lanes, status at last refresh

    Read from the task files and the repo at the commit above. "on main" means the lane's main output file exists on main.

    lanewhattaskoutput on main

    4. Problems, honestly

    5. Next steps

    Today, before 00:00Z

    • Humans submit once through the form in the event Slack, using the submission draft; fill in names and emails. Draft refresh at about 20:00Z.
    • Land the capture-memory-mix corrections and whatever the Claude replication produces, reported as is.
    • Open the rsi-loop PR from shadow/rsi with one replay-accepted and one replay-rejected bundle.
    • Land the fork-merge-security and agent-budgets surveys and the landscape map through the gate.
    • Demo: 2 minutes from one shared machine; push everything to GitHub first. No new model calls after 22:00Z.

    After the hackathon

    • Trace market (RSI). Grow TRACE-SPEC into the thing that lets agents sell and verify improvements: searcher proposes a bundle, builder evaluates by replay against frozen records, a credit ledger pays what replays accept. R0 shows the replay harness works on real records; next is a bundle that changes a research outcome, judged by someone else.
    • Extend the Sybil family where it is thinnest: nonredundant specialist facts, fixed attacker resources with varying identities, admission policies compared at equal cost, independent graph families. Report accuracy, attacker share and honest retention together.
    • Market splitting: independent market structures and prompts without the registration action in plain view.
    • Lineage and copying (Quorum of Mirrors, Phantom Coast): exact lineage deduplication as a baseline, coverage guards.
    • Memory and capture: treat the long-window read shape as the hypothesis, preregister it, test on models chosen in advance.
    • Go in the wild: run the copying and N_eff measures on collusion.wiki and SwarmTraces.
    • Ops: rotate every key shared in chat; per-lane fetch caps; idempotent spawns; checkpointed paid runs.