Glossary

These are the words that keep turning up on this site, in plain language. They describe what each idea is for, not how it is built. The detailed design is held privately (see the research status note), and nothing here claims more than the evidence allows.

The project

  • REE (Reflective-Ethical Engine). A research programme asking what a mind needs, at a minimum, to act well in a world it shares with others and cannot fully predict, and whether that can be worked out from first principles rather than trained in and hoped for.
  • The creature, the agent, the fish. The small simulated agent REE is built and tested on. It lives in a grid world. Nobody is claiming it is conscious, sentient or alive.
  • V1, V2, V3. Successive versions of the working agent. V3 is the current one; V1 and V2 are earlier versions kept as a record.
  • The fishtank. A long, watchable run of the whole agent in a richer world, used to see what it actually does over time, rather than what we hoped it would do.
  • Lab Window. The public, read-only summary of the reviewed evidence record, in aggregate.

The foundations

  • The axioms. A handful of starting commitments that REE does not try to prove. Give them up and, I would argue, coherent thought about acting in a shared world goes with them. Everything else is meant to follow from them.
  • The derived ethical objectives. What the axioms seem to require of any agent that takes them seriously: preserve minds; preserve future options; reduce unnecessary suffering; increase shared joy; stay correctable; keep seeking the truth; keep the ability to love and be loved; look after the shared world; keep open the possibility of future minds and future love; communicate honestly. They are meant to be consequences, not a list of preferences.
  • Ethics in the machinery. The central bet: that ethics has to shape how possible futures are generated in the first place. A filter added at the end tends to be learned around.
  • The cognifold. The whole of the agent’s changing inner state together with the rules by which one moment becomes the next. It is the running system, not any single part of it.

The parts of the mind

  • E1. The slow, deep part: a lasting model of the world and of the self.
  • E2. The fast part: a quick guess at what an action will do next.
  • E3. The part that weighs possible futures and commits to one.
  • Control plane. What decides how much to trust each signal, how much effort to spend, and which mode the mind is in, for example looking outward at the world or turning inward to imagine.
  • Hippocampal systems. Imagining possible futures, keeping a map, replaying the past, and carrying what earlier harm has left behind.
  • Imagined futures (rollouts). Short what-if sequences the agent plays forward before it acts.
  • Commitment. The point at which the agent stops weighing and acts, and after which it is answerable for what follows. Imagining something carries no responsibility; doing it does.
  • Precision. How much weight a signal or an error is given. Many psychiatric conditions can be read as precision going wrong.
  • Sleep and replay. An offline phase in which experience is gone over again and consolidated. Whether it changes behaviour in REE is still an open question.

Harm, care and residue

  • Harm. Signals that tell the agent it is being damaged, or is close to being damaged. REE models harm so that it can one day care about not causing it.
  • Residue (moral residue). A lasting cost left by harm, even justified harm, that keeps past actions alive as constraints on future ones. Closer to conscience than to a rulebook.
  • Dread. The name of one internal signal of anticipated harm. It is a measurement, not evidence of fear. One reading of it went badly wrong; see What We Got Wrong.
  • The caregiver requirement. The idea that the capacity for ethics may be present from the start but only becomes motivating through being cared for, through being loved and treated as worth loving. Testing it needs more than one agent, which is future work.
  • Love-exclusion failure. A way development could go wrong: an agent learns that love exists but concludes it does not apply to itself. What is left is goal-pursuit without moral motivation, the shape of ethics without its force.
  • Moral progress between generations. The hypothesis that each generation has to be handed its ethical orientation through care. Leave it to chance and it has to be rediscovered from scratch.
  • The welfare watchdog. An external check on long runs that stops the run if a negative signal runs away or the agent looks trapped, keeps the record, and needs a person to decide before anything similar runs again.

How the work is kept honest

  • Claim. A registered statement about how the mind works, with a place for the evidence and a place to mark it wrong.
  • Supports, weakens, mixed. The direction a reviewed result points for a claim. “Supports” means a test was passed; it does not mean proven.
  • Failure autopsy. Opening a failed experiment and asking what exactly failed: the agent, the mechanism, the measuring rod, or the world it was tested in.
  • World adequacy. Whether a test world gave the agent a fair chance to answer. A failure in a world that could not be answered is not evidence.
  • Closure map. The private dashboard of what is finished and what is not. It measures the roadwork, not the soul.
  • The Goblin Chronicles. Real turns in the project retold as stories, with goblins, forges, rulers and a fish tank. In them, the magics are the AI tools the work is built with: powerful, tireless and forgetful, and in need of governing.

Neighbouring work

  • World-model research. Other programmes build agents that learn a predictive model of their world. REE could borrow that kind of model for its slow, deep part. What REE adds is the parts those programmes usually leave out: deciding and committing under uncertainty, responsibility for what follows, and harm and care built into the machinery.