Every other agent product shows you the output. We want to
show you the work.
This is not a product page. There is nothing here to buy and nothing to download.
It is a place to think out loud, in front of a few people, about a piece of work
that is still being made — the arguments behind it, the reading that shaped it, and
the parts we have not settled.
We will say plainly which of these is which. Where we have an intention rather
than a result, it will read as an intention.
In development · not open to the public yet
The composition is the product
The industry is running a single race: make one model general enough to do anything.
The whole economic structure of the field is a bet on that premise, and it may well
pay off. We are not running that race, because we do not think the model is the
interesting unit.
The composition is the product — not the model underneath. The same base model
becomes many different composed intelligences depending on which specification it
loads. Change the specification and you have changed what the thing is, what it may
touch, what it is for, and what counts as it having done its job — without touching
the model at all.
"Composed" is doing deliberate work in that sentence, and it is not a synonym for
"narrow." Narrow implies a smaller version of something that ought to be bigger.
Composition is not a concession made while waiting for generality to arrive; it is
the mechanism by which capability gets arranged, directed and governed. You compose
the parts into an ensemble. Each part is itself composed — assembled from a role, a
boundary and a budget rather than handed raw capability. And when the arrangement
turns out to be wrong, you recompose it.
Three things follow, and they are the reason we think this is the better bet:
The arrangement is the design surface. If capability
comes from composition, you improve the system by rearranging it — something you can
do this afternoon, deliberately, and reason about afterwards. The alternative is to
wait for someone else's next model and hope the improvement lands where you needed
it.
Governance becomes tractable. Governing an unbounded
general agent means governing emergent behaviour: what can it reach, what may it
change, what will it do in a case nobody anticipated? The answer is always "it
depends," and the dependency is on runtime context that is hard to audit and harder
to explain. Governing a composition means governing an arrangement you wrote down.
One of those is a tractable problem.
The person is a composer, not a user. A user consumes
capability that someone else decided the shape of. A composer decides what each
intelligence is for, how the parts fit, and when the fit has stopped working. That is
a materially different relationship to a machine, and it is the one this work is
designed around.
The industry keeps throwing away the work
Agent systems are judged on their outputs. Did the answer come out right, did the
tests go green, did the customer accept the ticket. Those are outcome measures, and
outcome measures cannot tell three very different situations apart: a system that
reasoned well and got the right answer, a system that reasoned badly and got the
right answer anyway, and a system that reasoned well and was defeated by something
outside its control. From the outside all three look identical. Only one of them is
worth repeating.
This is tolerable when the work is small and cheap to redo. You read the answer, you
judge it, and if it is wrong you go again. It stops being tolerable as soon as the
work is long, or several things worked on it, or the question you actually need
answered is not "is this right?" but "should I trust the next one?"
The gap is that the reasoning which produced the output is not something you can go
and look at. It happens — it is genuinely the most informative thing the system
produces — and then it is gone, and what survives is a summary written by the same
process you were trying to check. Asking a system to account for itself after the
fact gets you a plausible account. That is not the same as a record, and the
difference only shows up in the cases where it matters most.
We find it useful to separate three things that get run together. Something is
expressed — the reasoning happens. It has to be
captured, or it evaporates the moment it is done. And it has to be
stored, or capture buys you nothing beyond the present moment. Most
of the category solves for expression alone, because expression is what the customer
sees. The consequence is that these systems have no past. Each session begins in the
same fog as the last one, and nobody can look back at how a conclusion was actually
reached, because there is nothing to look back at.
It is worth being clear that this is not a complaint about one vendor. It is
structural. The transcript is built as a delivery mechanism — something to read once
and move on from — and a delivery mechanism is not a record. Nobody chose to discard
the work; it was simply never anyone's job to keep it.
We want to close that gap. We are not claiming to have closed it, and we are wary of
anyone who says the problem is easy — keeping a record of reasoning raises hard
questions about volume, about what is worth keeping, about who gets to read it
afterwards, and about whether a record that is never inspected is worth anything at
all. What we are claiming is that it is the right problem, and that a field which
measures only outputs has agreed not to look at the only thing that explains them.
Why a transit map
When we needed a way to picture this, we did not reach for an architecture diagram.
Boxes and arrows describe how a thing is built, to people who already know how it is
built. We reached for transit cartography instead, and the case for it is mostly
other people's — a century of argument about how to draw a network you move through.
It is the most interesting reading attached to this project, so here it is at
length.
Beck, 1933 — throw away the geography
Harry Beck was a draughtsman in the London Underground's signals office, and he drew
the network in his own time, as a circuit diagram. He threw away the geography:
stations evenly spaced whether they were half a mile or five miles apart, every line
running horizontal, vertical or at forty-five degrees, the congested centre expanded,
the suburbs pulled in, and exactly one concession to the real world — the Thames.
Management turned it down as too radical. When they relented and printed a trial run
in 1933, the public took it faster than it could be reprinted. Beck was reportedly
paid five guineas.
The idea underneath is the one worth having. For a passenger underground, geography
is not information. What you need is which line, in what order, and where to change.
Beck's drawing is not a poor map of London; it is an accurate map of the network's
topology, and topology is what a rider is actually navigating. Every schematic since
is an argument about which truth to keep and which to sacrifice.
Vignelli, 1972 — and the instructive failure
Massimo Vignelli took the same logic further for New York. Forty-five degrees only,
one colour per line, the city itself flattened to a field, Central Park a neat square.
It is a beautiful object and a canonical piece of modernist design, and it lasted
seven years before the authority replaced it with something more geographic.
The reason it failed is the useful part. New Yorkers, unlike Londoners, spend a good
deal of a journey above ground, where the diagram can be checked against the street.
A schematic is a contract with its reader — I am going to lie about these things
so that I can be exact about those — and the contract holds only as long as the
reader has no cheap way to audit the lie. Vignelli's contract was defensible in
principle and broke in practice. Versions of his diagram have since come back for
contexts where the reader cannot check it: planned service changes, small screens,
anywhere the abstraction is all you have. He was not wrong. He was wrong about
where.
Marey and Ibry — a diagram you can think in
In La Méthode Graphique, Étienne-Jules Marey published a graphical schedule
for the Paris–Lyon line, a design he credited to the engineer Ibry; Tufte reproduced
it a century later in The Visual Display of Quantitative Information, and it
may be the most instructive transit graphic ever drawn. Distance runs down the page,
stations spaced by real distance. Time runs across. Every train is a single diagonal
line. Its slope is its speed. Where two lines cross, two trains meet.
Nothing is annotated because nothing needs to be — delays, slack and conflicts are
visible as geometry, and you see them rather than calculate them. Tufte's point is
that this is not a picture of a timetable. It is the instrument the timetable was
designed with. That is the standard worth aiming at: a diagram you think in,
not one you read after the thinking is over.
Roberts — the rule is not the reason
Maxwell Roberts, a psychologist at Essex, spent years doing the thing the design
field mostly did not: testing. He drew alternative Underground maps, including frankly
curvilinear ones — arcs and circles, Beck's octolinear rule broken outright — and then
timed people planning real journeys on them.
The findings are chastening for anyone who assumes the rule is the reason the map
works. Octolinearity by itself is not what produces usability; coherent line
trajectories, simplicity and consistency are, and a well-designed curvilinear map can
beat a badly designed octolinear one. He also found that what people say they prefer
and what they actually perform better with come apart. Preference is not performance.
That is a useful thing to know before defending a house style on the grounds of taste
— including ours.
Gallotti, Porter and Barthelemy — the ceiling
In 2016 the three published a study in Science Advances asking how much
information a traveller needs in order to plan a trip through a multimodal network.
They found a limit. Past roughly two hundred and fifty connection points, the
information required to plan an unfamiliar journey exceeds what an unaided person can
hold, and the map stops functioning as a map. Several real networks are already past
it.
That result does something specific to how we think about drawing one. The number of
things you may put on a map is set by what a reader can hold, not by how many things
exist. A map is not an inventory. The moment it becomes one it has failed at the only
job it had — and the failure is invisible to the person who drew it, because they can
read it fine.
Guo — and why this is a claim, not a decoration
One more, because it raises the stakes. Zhan Guo's 2011 study of transit map effects
found that map design measurably changes the routes people choose: riders will take a
path that is shorter on the diagram and longer in life, and in the network he studied
the drawing influenced path choice more than the actual travel time did.
A map does not merely describe a system. It teaches people what the system is, and
they believe it over their own experience. Drawing one is a claim you are making about
your own work, and you are answerable for what people conclude from it.
So: why this form
Because what we are building is something you move through rather than something you
look at. What matters is which line, in what order, and where things connect — the
topology, not the floor plan. And because, like Beck, we have thrown the geography
away.
That is the honest choice for a page like this one, and it is also the reason a
drawing like this can be shown at all. A picture of what it is like to move through a
system is not a picture of how the system is built. The lines carry an identity and
nothing else. Read it as a promise about character, not as a diagram of parts — and if
you try to count something in it, you will be counting a composition decision.
Local-first, and what it costs
The work stays on the machine where it was done. That is the default, not a setting
you go and find.
The reason is not ideological. It follows from everything above. If the point of the
exercise is a record you can inspect and disagree with, then a record held by someone
else is a weaker version of the thing — it can be changed, lost, priced, discontinued,
or handed to a third party, and none of those require anyone to act in bad faith.
Putting the record where the work happened is the only arrangement in which
inspection means what it sounds like it means.
Here is what that costs, because pages like this one usually skip this part.
Sync is your problem. Two machines means you solve
it, or you accept that it does not happen. There is no invisible service quietly
making your work appear everywhere.
Backup is your problem too. A dead disk is a dead
archive. We cannot restore what we never had, and there is no support ticket that
changes this.
Collaboration is genuinely harder. Sharing something
becomes a thing you do, rather than a link that already exists. Some of that friction
is the feature — it makes disclosure deliberate. Some of it is just friction, and we
are not going to pretend we can always tell you in advance which one you are
hitting.
We learn slowly. Products get better fast partly
because they watch everybody use them. We have given that up. It means we improve from
fewer signals and more slowly, and mostly because somebody took the trouble to tell us
something.
Your hardware is the ceiling. What runs, runs where
you are. That bound is real and you will meet it.
And the honest one, which gets left out of most claims of this kind: local-first is
not the same as offline, and it is not privacy in the absolute. If you use a model
that runs somewhere else, what you send it goes somewhere else. What stays with you is
the record — the work and its history. That is a real boundary and we think a
worthwhile one, but it is a narrower claim than "your data never leaves," and we would
rather draw it precisely than let you assume the larger version.
We would rather name the trade than sell around it. If those costs are disqualifying
for you, then they are disqualifying, and that is a perfectly reasonable place to end
up.
Notes
Short entries, most recent first. Things being worked out rather
than things concluded. Some of these will turn out to be wrong; they are left up
anyway.
2026-08-06 — A map is not an inventory. Still thinking about the
two-hundred-and-fifty-connection ceiling. The useful consequence is a prohibition:
however much exists, the drawing may only carry what a person can hold, so a map can
never be a census of the thing it depicts. Which means the honest response to "what
is on the map" is that the map is a composition, and that the pull towards
completeness — draw everything, it is all true — is the exact instinct that makes a
diagram useless.
2026-08-02 — Preference is not performance. Roberts' experiments
are a standing rebuke to how we are choosing this visual language. We like it. We
have not tested whether anyone reads it faster, or reads it correctly, and liking a
drawing is not evidence that it works. This is a gap we know about and have not
closed.
2026-07-24 — "Observability" is already taken. The word means
something specific in software: metrics, traces and logs, emitted by machines for
machines, so an operator can tell whether the system is up. What we are after is not
that. It is nearer to a lab notebook than a dashboard — the record of a piece of
reasoning, kept because someone may need to disagree with it later, possibly months
later, possibly the person who wrote it. We keep reaching for the word and keep
having to put it back. Suggestions welcome.
2026-07-18 — A control that is never tested is a decoration.
Something we hold ourselves to, and mention because it applies to this page too: a
safeguard is only worth having if you have watched it fail on purpose. A check that
has only ever passed tells you nothing — it might be enforcing the rule, or it might
be doing nothing at all, and from the green result those look the same.
2026-07-11 — Vignelli's contract, applied to this page. His
diagram broke because the reader could step outside and check it. We are abstracting
at least as hard here, and a reader has no way at all to check us — which is a more
comfortable position and a less honest one. The obligation that comes with it is to
be clear that this is an abstraction, rather than to let it be mistaken for a
specification. Hence the "read it as a promise about character" line in the map
section, which is doing real work and should not be cut for length.
2026-06-30 — What we owe a dead laptop. Local-first puts the
archive in your hands, which is the point, and also means we cannot get it back for
you. Every answer we have sketched so far quietly re-centralises the thing it was
meant to avoid. Current status: unsolved, and we would rather say so than ship a
reassuring sentence.