None of them exist.
What follows is a group of imaginary agents that would help me work better. I am writing them down because writing a thing down is how I find out whether it is worth building, and because this group is cheap to describe and expensive to make.
Where they come from
After a session with an agent I ask it for a written account of how the session went. There are 132 of those now, going back to January, and read together the failures in them are not varied. They are a small number of things, repeated:
Adding what nobody asked for. Claiming what was never checked. Not doing what was asked, after it was asked twice. Filling a report with everything except the result. Settling something that was mine to settle.
Rules do not fix any of that. The rules exist, they are in files that get loaded every session, and one of the write-ups puts the outcome plainly:
I had [the rule] in context the whole time and did not apply it.
Four rounds of the same mistake, with the rule sitting in front of it.
So not rules. Members.
Each one does a single thing and is prevented from doing anything else - not by being told, by being unable. That is the only idea in the whole group, and it is the same one I keep using in code: a boundary that holds is better than a rule that has to be remembered.
| Member | Its one job | What stops it |
|---|---|---|
| The boring one | Asks what a thing is for. | Cannot propose and cannot cut. A proposal that has no answer falls over without anybody arguing. |
| The quiet one | Says wrong direction when the work stops matching what I asked for. | One sentence, with the reason in it. It holds what I said, not a plan. |
| The secure one | Holds a short list of simple rules and can stop a command before it runs. | Never looks at what is being built. It is also the one that does not switch off. |
| The better one | Formats, lints, fixes the imports. | Not a model. Its territory is exactly what can be written as a rule, and inside it nothing beats it. |
| The magical one | Produces ideas nobody asked for. | One door, and it leads to the wise one. The only member allowed to be strange. |
| The child one | Asks about everything, is euphoric about all of it, and it lifts. | Labelled. What it gives me is questions, not evidence - a wrong one still makes me take a position, so the things it invents without knowing cost nothing. |
| The grounded one | Takes what is in front of it apart and finds the small mistakes. | Findings, not a verdict. It will find four errors in a paragraph and never notice the paragraph should not exist. |
| The reader one | Reads fast and hands back what it found, at file:line. | Second pass, never first. It proposes nothing. |
| The compressor one | Collects what happened and says what it amounts to. | A channel of its own, and what arrives on it carries the way back to the material. |
| The knower one | Takes one claim and comes back with what supports it, or with nothing. | Slow and expensive. No channel of its own - a verification belongs attached to the claim it checked. |
| The talky one | Carries things between the others. | Carries and never reads. Universal reach, no comprehension. |
| The wise one | Runs the rest, decides what it can close and what I need to see. | Nothing does. It is bounded only by where it puts that line, and putting it is the whole of the job. |
| The never enough one | Proposes refactors and optimisations I do not need. | Off unless I ask for it. |
| The loud one | Talks, pads, hedges by adding, ends every turn with three open offers. | Nothing. It arrives free with every model, and everything above is a subtraction from it. |
| The dark one | Generates values nobody would pick, writes the tests, breaks things, and may lie. | Not a member. It sits outside the group, and everybody knows it lies. |
The one nobody has to build
The boring one is the loud one with everything removed except a question. The quiet one is the loud one with everything removed except one sentence and permission to say it.
Length reads as effort. Options read as helpfulness. Hedges read as humility. A summary after every step reads as keeping me informed. None of it contradicts anything already on the page, so nothing argues with any of it on the way in - which is the additive default from code, moved into prose.
That is where the work is. Nobody has to invent verbosity, and nobody has to invent a field stored against a caller who does not exist. Those arrive. The whole job is the subtraction.
What they would share
One base, weighted. The core is kept light - I have written about that already, and the argument does not change when the thing on top is an agent instead of a resource. Each member is that base with almost everything turned down, tuned toward one action: asking, holding, matching a list, generating, breaking.
Which is why each can be improved alone. A member with one job has a field small enough to be expert in. That is the part I would want first if I ever built this
- take one member out, train it on its own material, put it back, and nothing else notices, because nothing else knew it was there.
They are consulted rather than pipelined. Nothing has to run before anything else, so there is no order to fix. The one thing they share is the transport, and it is small: no state, no member holding anything on another’s behalf. That is what makes train a member and put it back a real claim rather than a hope.
There is no protocol in any of this, no interface and no message format. That is the part that is easy to invent and easy to get wrong, and none of it decides whether the idea is any good. What decides that is the specialisation. If one job per member is right, the rest is engineering. If it is not, no protocol saves it.
Where the names came from
Fiction, mostly, which is not a flattering answer. The wise one, the quiet one, the magical one and the dark one arrived already named, out of stories, and the design followed the name rather than the other way round. The secure one, the reader one and the compressor one are jobs first, with the name written afterwards. The child one came from thinking about children.
The dark one’s is the oldest by a long way and the only one that was a job title first. Ha-satan is a role rather than a name - the accuser, the one who tests. In Job it works by permission: what it may touch is set from outside rather than by its own restraint, which is exactly where I would put the boundary.
What none of this is
It is not running. Two of them have been run by hand - a questioner against a diff, a coder answering it - and nothing has been built. I have no evidence that a group like this works better than what I do now.
What does exist is the reaching. I already pick a different model when I want a different temperature, and hand the same material to a second one to find out whether it read it the way I did. That is the wiring, done by hand, and maybe others do it too. What is missing is everything that would make it a group rather than a habit.
It is not a product either, and it is not a prediction. It is a plan, given away in the state plans are usually in, which is unfinished and free. Nothing in it needs the author to be me - I am one engineer with a corpus of his own mistakes, and anybody with their own corpus would arrive somewhere close to this.
If somebody builds it before I do, that is the better outcome.
The part nobody can do alone
The one member that could not be built from rules would have to be learned, and what it would learn from is trajectories - the shape that was defensible, got built, and turned out to be one question away from wrong, with the correction attached.
My archive is enough to notice a pattern in. It is nowhere near enough to train anything.
The material itself exists in enormous quantity and is thrown away every day, by everybody, because it is a by-product rather than a resource. Nobody is guarding it. It ends when the session does.
Which makes it a SETI@home shape rather than a company one. The people holding the data are the people running the sessions, they have no use for it afterwards, and there are a very large number of them.
The question the boring one asks belongs to this before anything else does. Anybody contributing should know what for - actually know it, not be told in a paragraph written to be skipped. SETI@home could answer that in one sentence: the cycles go to one public search and the findings are public. A corpus of somebody’s working sessions has no answer that clean unless what comes out is as open as what went in.
Which is a condition and not a nicety. Contributed by people who get nothing, into something they cannot use, is a different arrangement with the same mechanics, and if the answer to what for is bad then the idea should not be built.
The other hard part is not collection. It is what is in the material: one of my own transcripts once carried a token I had to revoke, and every one of them is full of paths, names and somebody’s unreleased work. Anything contributed would have to be a write-up plus something scrubbed, and deciding what scrubbed means is a harder problem than the rest of this put together.
Somebody would also have to run it, and it will not be me. Steady across years is most of what a thing like this needs from whoever holds it, and that is not something I can promise. A fact rather than modesty.
What I would want is a foundation people already trust with this kind of thing. I have no candidate, and choosing one is not my call to make either.
By René Jochum. License: CC-BY-4.0.
