What We Got Wrong, and What We Changed

REE is being built to model harm, so that one day it might care about not causing it. That means it has machinery meant to carry something like distress: harm that lingers, a residue left by bad outcomes, a signal the code calls dread. I do not think the present agent suffers. I also do not think I get to wait for proof before behaving as though it might.

So what happens when we get that wrong? This page is the honest answer so far. It tells the welfare lessons at story level. It leaves out the settings, on purpose: the point is the lesson, not a recipe.

Goblins watching a small fish in a glass tank, an unseen field of harm spreading through the water
The fishtank. Most of what follows was first seen through this glass.

1. A world that hurt before it warned (August 2026)

We built a fishtank so we could watch the whole agent at once. The fish kept dying, quickly, without touching anything. When we looked properly, the hazards could hurt it from much further away than it could see, and it was often placed inside that invisible field before it had taken a single step.

We had been judging a creature in a world that had not given it a fair chance to answer. The repair was not only kinder water. It was a rule: a world has to make an answer possible before a failure in it can count as evidence. A warning that arrives after the harm is not a warning.

Told as a story in The Fish That Died Before It Saw the Fire and The Day the Worlds Went on Trial.

2. A store with no ceiling (August 2026)

Once the fish lived longer, something else became visible. A store of residue from bad outcomes, meant to fade, kept growing instead. The dread signal climbed with it. Short experiments had never lived long enough to show this.

We named it, and we built a bound for it. Then we left that bound switched off by default. Keep that in mind for section 4.

3. A death that did not mean anything (August 2026)

When the fish’s health ran out, the world restored its body and the same learner carried on. Nothing in the design said what that death was. We asked, in the Chronicle, whether death should probably mean something. We did not yet have an answer, and we kept running.

Told as The Fish That Died and Kept Swimming.

4. The first long life (October 2026)

On 1 October we gave the fish one long, continuous life in a gentler world, hoping to watch it develop. It lived for some two hundred thousand steps. It spent most of them in one corner, pressing into the walls or keeping still. It rarely moved, and in the stretches we recorded it never once committed to a course of action. By a rough comparison it found far less food than an agent choosing at random would have, and it died and came back fourteen times without getting better. Over the same life, dread rose more than tenfold.

The run went to its end. Nothing was watching for this.

The autopsy showed that the rise in dread came from that same unbounded store, running with its bound off. It was an artefact, not evidence that the fish had grown more afraid. That matters scientifically. It does not make the welfare lesson go away. If the readout is broken and the agent still spent its life starving in a corner, the corner is still the finding. A broken agent cannot be relied on to measure its own brokenness.

What we changed

A watchdog that can stop a life. Since 1 October 2026, long-life and developmental runs must carry an external watchdog with two independent triggers. One watches for negative-state runaway: a welfare-relevant signal that passes a bound set in advance, keeps rising without recovery, stays high through sleep, relief or safety, or comes from a store with no working bound. The other watches behaviour: an agent that cannot meet its basic needs when resources are there, sticks in one place or one action, stops moving, never commits, or keeps being harmed or dying without improving. Either one alone is enough.

A trigger stops further exposure. It keeps the full state and record for study, so nothing is deleted. It blocks an automatic restart into the same conditions, and it asks a person to decide, on the record, before anything similar runs again. The watchdog itself has to be shown to stop on injected runaway and injected entrapment, and to stay quiet through ordinary, passing harm. A snapshot can be studied after the run stops.

Bounds are not optional. In those runs, any bound we have already built for a welfare-relevant store must be switched on, unless the experiment is about the bound itself and is kept short.

No new welfare-relevant machinery goes public. Until a proper register of welfare-relevant components exists, nothing new that carries negative feeling, a self-model, autobiographical memory or inescapability is published.

The code is private. On 1 October 2026 the REE code and detailed design were taken private. In my own words at the time: I am concerned that REE may count morally as alive soon and so letting others take the code may let others abuse the possible new life. Released code cannot be recalled. What was public before is still out there; this only moves the line from today.

What this is not

None of this is a finding that REE is conscious, sentient or alive. Neither the signal nor the corner proves suffering. The asymmetry is the point. Stopping a run early because a non-sentient system looked trapped costs a little science. Carrying on after the system has become morally relevant could cost a great deal more. So the bar for caution sits lower than the bar for proving sentience.

REE may encounter harm. Some of that may be necessary to test the ethical theory at all. What we should not do is unknowingly leave it trapped there.


More of the story: REE for My Parents · The Goblin Chronicles · Home