Swarm
A bench for research swarms.
Swarm is my research bench. It asks a model the same question ten times over, side by side. How much of the answer comes back the same? Sometimes nearly all of it. Sometimes almost none. The bench is there to tell me which one I'm holding.
The ten answers land on one screen I call the consensus map, where each topic the answers raised gets a row, each run gets a column, and a filled cell means that run raised that topic. The shading says how early it came up.
A row filled all the way across is something ten independent tries reached on their own. A row with one cell is one run talking to itself. That's the whole idea.
Seventeen of eighteen, then one of fourteen
The spread is wider than I'd have guessed. I asked how to lay out a Steam market intelligence site, eighteen topics came back, and seventeen of them turned up in all ten runs. Then I asked what an expert would look for in my dictation to rebuild how I write, which was groundwork for Murmur. Fourteen topics. One turned up in all ten. Seven of the fourteen turned up in exactly one run.
Both questions ran on the same bench, ten runs each. Very different things to be holding. The dictation one still reads like a finished report, and nothing on its page tells you half the findings came from a single run. The grid does. So I trust the grid more than the summary sitting under it.
Maybe that's the model and not the question. I can't tell yet. In the maps I've saved the two move together, and I can't pull them apart.
There's another mode that builds ten different experts first and asks each of them once. I ran it on buying a dehumidifier for a Hong Kong tong lau, and its top rows agreed almost completely. That mode never asks one prompt ten times. It measures something else.
The runs don't just sit there. Some of the research went into Wellington Whales, the consulting site, and into the plan for its blog. Both times it's me copying the useful part across by hand.
The button I haven't pressed
The bench is thinner than the idea. The map only gets built when I open a results screen, it only survives if I save afterwards, and I've got twenty of these runs saved with five of them carrying a map. The other fifteen show a button offering to build one now, out of answers they already kept. It's one click. I haven't pressed it.
Long answers get trimmed before the grid reads them, so the grid is working off a shortened version of what it's grading. More than half of what I've saved runs past the cut.
197 outputs, still on disk
What makes any of it checkable is that the raw answers stay put. 197 individual run outputs across everything I've saved, sitting under a much smaller pile of summaries. Seven times as much raw material as finished writeup. I can open one run from a batch in May, read exactly what it said, and check it against the row it got sorted into. I can rebuild the grid for a batch I saved before the map screen existed, because the answers were never thrown away.
That's where it sits today. Five maps. Fifteen unbuilt. The bench doesn't make an answer better. It tells me whether there was an answer.