off policy

Neuroscience Machine learning14 September 2026

Fast-weight continual learning in a fly

Peter Wang and Nico Christie

Peter’s notebookCode on GitHubNico’s training plan

In our words

I and @nicochristie ran the fly connectome and found the group of neurons (hΔH, hΔA, hΔI and hΔG) that could allow the fly to navigate using fast synaptic weight updates, not neural activations. This is fast-weight continual learning in a fly, something current LLMs don't do!

@BrainsAndTennis on X · 14 September 2026

The navigational system of the fruit fly is a crown jewel of systems neuroscience due to the work of some seminal neuroscientists (Larry Abbott, Gaby Maimon, Vivek Jayaraman, Barbara Webb), but it is an unfinished story.

A fly that leaves a drop of food and wanders in the dark can always find its way back. To do that it has to keep a running sum of every step it has taken, a process called path integration. The neurons that report each step are known, but the neurons that add the steps up have never been found.

There are two ways a brain can hold a sum like that. The usual answer is that some neurons holds it in activations and sustains these activations by exciting each other in a loop tuned so precisely that the signal neither fades nor blows up. Most models of navigation assumes this, and it is how RNNs and LLMs hold state too. The other answer is that nothing keeps firing at all. Each step is encoded into synaptic strengths (aka weights), and the sum of all synaptic strengths is the sum of the journey. A few papers have suggested the fly works this way but nobody has pointed to any candidate neurons, until now.

Four neuron types, hΔH, hΔA, hΔI and hΔG, have no known functions, but we found that they have every ingredient option 2 needs. They receive input from neurons that report each step taken, they receive velocity-sensitive dopamine input that could gate memory writing, and they receive a reward-sensitive octopamine neuron that could reset the synaptic weights at food arrival. Simulations confirm this is a viable candidate for path integration.

This is the key finding, but we have posted 3 other findings in links below. We tried to be exhaustive with published papers but may have very well missed some key published results, so inviting the cogniscenti to engage.

@BrainsAndTennis on X · 14 September 2026

What started off as trying to turn the fly gay might have led to the first meaningful scientific finding from the connectome release

@nicochristie on X · 14 September 2026

Finding 1 of 4 · the missing middle of the circuit

The fly's running sum of its own steps may be held in synapse strengths, on two hΔ neuron types nobody has recorded

What kind of claim each statement is. This page reads the connectome and the literature; every claim is one of the following.
  • Already published: the idea that the columnar inputs to FC2 store vectors is Maimon & Abbott 2026; synaptic vector memory is Hulse 2021, Goulard 2023, Dan 2024. hΔG (Janke 2025) and hΔA (Avritzer 2026) have been recorded as leaky integrators over seconds; the mechanism, activity or synaptic, is not determined.
  • Measured here in the wiring: which hΔ types receive hΔB in the matching column, and per-cell dopamine and octopamine coverage, in two connectomes.
  • Simulated here: a toy agent showing a vector sum can home.
  • Proposed here, untested: a synaptic store at hΔH or hΔI, both unrecorded.

Summary. A fly that leaves food and wanders in the dark can walk straight back, so it keeps a running vector sum of its steps. The neurons that report each step, hΔB, are known. Maimon & Abbott 2026 proposed that the columnar populations feeding the goal neurons FC2 could each store a vector, for different durations, as activity or as synaptic weights; two of those populations have since been recorded in the Maimon lab and both integrate over seconds, with a leak (hΔG: Janke 2025; hΔA: Avritzer 2026); whether the integration is done by firing or by synapses is not yet determined. A minutes-long sum that survives a detour has not been found in any of them. This page asks what the wiring says about a synaptic version: hΔB's travel-direction bump lands on hΔH and hΔI in the matching column, so each step would strengthen only the synapses for its direction and the weight pattern across columns would be the sum; dopamine types cover every cell of both as a write gate; and their outputs reach FC2 and the steering neurons PFL3. The connectome cannot decide between activity and synaptic storage, and it cannot tell a leaky integrator from a lasting one. What it can do is say where to look, and hΔH and hΔI have not been looked at.

The walk food hΔB right now Synapse strengths (the store) What the store holds N read from the weights
The fly walks 6 steps north, 4 south, 3 east. Each step, hΔB's bump strengthens the synapses in its own column. The weights are the running sum; the arrow is what a downstream reader sees. Step-by-step version →

The question

To return to a place it cannot see, the fly must know the vector from that place to itself: the sum of every step since leaving, each step counted in world coordinates. The fan-shaped body has the neurons that report each step: about twenty hΔB cells whose bump of activity sits at the column matching the direction the fly is moving, with height proportional to speed (Lu 2022; Lyu 2022). Each column is a compass direction, so the bump is the current step as a vector. The question is what adds these steps up, and for how long.

Two ways to hold a running sum

A · activity integrator hΔB bump integrator populationbump = running total recurrence, gain ≈ 1 required Memory = firing pattern. Needs strong, structured self-connections. B · synaptic integrator hΔB bump store populationfires only when driven memory = these synapse strengths "walking" modulator "at food" reset Memory = weights. Needs a write gate and a reset, not recurrence.
Left: an activity integrator must feed back on itself with gain near one. Right: a synaptic integrator needs no feedback, but needs a write gate and a reset. The connectome can tell these apart.

The first is in activity. A population keeps firing the current total and adds each new step to it. For the total not to leak away, the population must excite itself with a gain of almost exactly one. This is how most models of path integration work.

The second is in synapses. Nothing keeps firing. Each step makes hΔB strengthen the synapses it is currently using. After a walk, the pattern of synapse strengths across the columns is the sum: six steps north and four south leave the north synapses at six and the south ones at four, and the difference is a vector of length two pointing north. This needs no self-excitation. It does need hΔB to land on its target in the matching column, a signal that says "write now" while the fly walks, a reset at food, and a path from the store to the neurons that steer. Hulse 2021, Goulard 2023 and Maimon & Abbott 2026 have all described vector memories of this kind, and Dan 2024 modelled a goal angle stored as a set of synaptic weights onto the steering neurons.

What has been recorded

Two of the hΔ populations that feed FC2 now have physiology, both from the Maimon lab and both available as thesis abstracts.

  • hΔG (Janke 2025): activity drops sharply when the fly reaches sugar, then its mean rises over minutes while the amplitude of a sinusoidal bump tracks the distance walked over the past few seconds. The bump is built by integrating input from vΔE, with a continuous leak, so it is not a perfect path integral. The same vΔ→hΔ motif recurs across layers of the fan-shaped body.
  • hΔA (Avritzer 2026): integrates the fly's recent travel direction over a window of 7–10 s and feeds the steering circuit as an inertia term that promotes continuing in that direction.

Both are integrators, both leak over seconds, and both are read out in calcium; neither abstract determines whether the integrating step is recurrent firing, a cellular process or a synapse that facilitates and decays, and Maimon & Abbott 2026 list all of these as candidate mechanisms. hΔK adds a third measured case, a persistent bump that accumulates odour evidence and holds an upwind heading for tens of seconds (Lanz 2025; Kathman 2026). None of them has been shown to hold a displacement across a detour of minutes and point home afterward, which is what the behaviour requires (Kim 2017; Behbahani 2021; Titova 2023).

What the wiring cannot decide

Could a lasting sum be held in activity somewhere in the fan-shaped body? Counting each population's synapses onto itself, binned by column offset, and taking the cosine component as a share of total input gives the same answer for the fan-shaped body as for the compass, whose bump does persist for minutes: vΔA_a 0.145, EPG 0.050, the hΔ types 0.01–0.02, the EPG↔PEN shifter loop 0.003. Synapse shares are not gains and set no threshold. The wiring leaves the activity option open, and it cannot distinguish the leaky seconds-scale integration that has been measured from the minutes-scale integration that has not, nor say which mechanism does either.

What the wiring points to for a synaptic store

hΔH: stores it hΔB lands almost entirely in hΔH's own column, so north steps write north, south steps write south, and the difference survives. 1 · where hΔB's synapses land opposite same 82% column, relative to hΔB's own column 2 · weights after the walk W S E N dark: written by the 6 north stepslight: written by the 4 south steps 3 · the vector it holds 1.7 northgrey: the truth, 2 north hΔJ: cancels (rejected) hΔB lands in hΔJ's own column and the opposite one, so every step writes both north and south, and the two contributions cancel. 1 · where hΔB's synapses land opposite 28% same 35% column, relative to hΔB's own column 2 · weights after the walk W S E N dark: written by the 6 north stepslight: written by the 4 south steps 3 · the vector it holds 0.3 northgrey: the truth, 2 north
Worked example computed from the measured wiring. The fly walks 6 steps north then 4 south, so the truth is 2 north. Top: hΔB lands in hΔH's own column, so the weights it leaves point north with about the right length. Bottom: hΔB lands in hΔJ's own column and the opposite one, so every step writes both directions and the total cancels. hΔA, hΔI and hΔG behave like hΔH.
hΔBtravel direction bump hΔH · hΔIunrecorded candidates (hΔA, hΔG: recorded, leaky)every cell receives hΔB, column-matched 1,000–3,600 each FB4M, FB5H (dopamine)FB4M driven by PFNv, hΔB, PFNd and the ascending AN19B019: on while walking FB4M on every hΔA and hΔI cell; FB5H on every hΔH cell OA-VPM3 (octopamine)gates recent memory in the mushroom body (Kapoor & Waddell 2024); reset role untested hΔH 361 (all cells) · hΔI 72 FC2 goal / PFL3 steering1,400–5,700 each read-out All numbers are synapse counts in MaleCNS; every edge shown replicates in the hemibrain.
Everything a synaptic store needs meets on hΔH; hΔI lacks the reset. Synapse counts from MaleCNS; every connection shown is also present in the hemibrain.

Scoring every hΔB target for the four requirements of a synaptic store, and setting aside the two populations already recorded, two candidates remain.

typehΔB synapsessame columnwrite gate (cells covered)resetsends to
hΔH (8 cells)1,01082 %FB5H dopamine 249 (8/8)OA-VPM3 361 (8/8)FC2, PFL3
hΔI (17)3,57548 %FB4M dopamine 380 (17/17)nonePFL3, PFL2
hΔJ (31), rejected4,23235 % (+28 % opposite)FB1H 1,184 (31/31)OA-VPM3 377FC2

Column matching is what makes a synaptic store work. hΔB lands on hΔH in hΔH's own column 82 percent of the time, so a step north strengthens the north synapses and little else. hΔJ receives more hΔB synapses than any candidate, but in two opposite columns, so every step strengthens both north and south and the sum cancels. In a simulation that writes a random walk through each type's measured landing pattern, hΔH retains 86 percent of an ideal store, hΔI 56, hΔJ 14. (The recorded hΔA and hΔG score 76 and 45; whether a synaptic component sits under their calcium signals is untested.)

A write gate is present on both. FB5H, a dopamine type, contacts every hΔH cell; FB4M, another, contacts every hΔI cell across all columns and is fed by the velocity neurons PFNv, PFNd and hΔB and by an ascending neuron from the ventral nerve cord, so it is placed to fire while the fly walks: the translational counterpart of ExR2, which gates compass learning by rotation speed (Fisher 2022).

A reset is the weakest link. OA-VPM3, the octopaminergic pair that supplies most of the mushroom body's octopamine and gates recent-memory expression there (Kapoor & Waddell 2024), contacts every hΔH cell uniformly across columns; nothing about its known function says it fires at food. hΔI has no uniform octopamine input. Janke's hΔG reset at sugar shows a reset signal exists in this circuit; what carries it is unknown.

The output goes to the right place: FC2 and PFL3, which turn a stored direction into walking (Mussells Pires 2024; Westeinde 2024). Every edge here is present in both connectomes. hΔH meets all four requirements; hΔI lacks the reset.

How a synaptic store is read

There is no recall step. A store cell fires as hΔB's drive times its synapse strength. hΔB is active whenever the fly moves, so as soon as it walks, the store's output across the columns is the stored pattern, flowing to FC2 and PFL3. The same walking that reads also writes: walking home strengthens the opposite columns until the difference is zero, which is arrival. The output points from food to fly; turning it into "go back" is a separate neuron, and is finding 2.

What the simulations show

Only a running vector sum reproduces the fly's search behaviour. In Kim's and Titova's assays the search-centre error is 0.8 units for a vector sum, 12.7 for a random walk, and 26.8 for either remembering the direction at the food or the distance walked. In closed loop, an agent using the store returns to the food (median closest approach 1.1 units versus 6.8 without memory), provided the readout is rotated by half a turn before steering; without the rotation it walks away every time.

What is not known

Nobody has recorded plasticity at these synapses, or dopamine or octopamine acting on them. If synapses only strengthen they saturate over trips; either a reset or a rule that weakens the opposite synapses on the return is required, and the connectome cannot tell which. A fly standing still has no hΔB drive and so no readout. The behavioural premise itself has a caveat: silencing the compass did not reduce direct returns to food in local search (Goldschmidt 2026). And the measured hΔ integrators all leak over seconds, which is a reason to expect the same of hΔH and hΔI.

What would settle it

  • Image hΔH in a fly re-zeroing its home vector (Behbahani 2021), with a cAMP sensor alongside calcium, as Gorko 2025 did for the PFL1 goal circuit where the stored direction proved absent from calcium: a synaptic store grows with distance in cAMP or weights; an activity store grows in calcium; either collapses at re-zero.
  • Block dopamine or cAMP signalling in hΔH or hΔI: the fly should stop returning while hΔB and the compass stay intact.
  • Silence OA-VPM3 while the fly feeds: if it is the reset, subsequent search should no longer centre on the food.

Prior work

hΔB as the travel-direction signal: Lu 2022; Lyu 2022. Vector memories in the columnar inputs to FC2, with durations set by mechanism: Maimon & Abbott 2026 (Figure 5). Synaptic storage of a vector or goal angle: Hulse 2021; Goulard 2023; Dan 2024. hΔG as a leaky distance integrator reset at food: Janke 2025 (thesis, abstract; full text embargoed to October 2026). hΔA as a 7–10 s travel-direction memory: Avritzer 2026 (thesis, abstract; embargoed to May 2027). hΔK persistence: Lanz 2025; Kathman 2026. What is added here is the cell-by-cell wiring test of the synaptic version: which hΔ types receive hΔB in the matching column, which have uniform dopamine and octopamine coverage, and the readout and sign analysis that follows in finding 2.

Methods

Column offsets: every fan-shaped-body neuron in MaleCNS v1.0 carries a column label; angular offsets use the true column count per type (12 for hΔA/B/C/I/J/K/L, 8 for hΔD/E/G/H/M and PFNd, 6 for hΔF; C9≡C1 for nine-label types). hΔ cells are named by their dendritic column (Hulse 2021), and hΔB's bump sits on its axonal arbor (Lyu 2022). Recurrence: within-type synapses binned by angular offset, normalised by total input, signed by predicted transmitter, Fourier-decomposed; compass cells binned by ellipsoid-body wedge. Site screen: every synapse from a travel- or heading-tuned population onto every columnar population scored for column matching, per-cell modulator coverage, and output to FC2/PFL3. Simulations: rate-based agent, eight columns, rectified-cosine hΔB bump written at the axon column, weights incremented through each candidate's measured landing kernel, readout as weights times drive through its measured output kernel. Tables: data/derived/recurrence_modes_v2.csv, fb_column_offsets.csv, synaptic_site_screen.csv, cx_ext_cell_edges.csv. Scripts: scripts/recurrence_modes_v2.py, synaptic_site_screen.py; simulations: simulations/closed_loop_routes_tspace.py, strategy_discrimination.py, closed_loop_return.py. Notes: docs/exhaustive-search.md, docs/audit-2026-09-13.md.

The other findings

Working with Fable 5.1

The same way levent gets to do better math because of ai, someone should be doing better neuroscience because of ai. I have been unplugged from systems neuroscience for the last two years and but have iterated closely with fable 5.1 over the course of two days.

the good: fable 5.1 was creative and took a lot of scientific paths that were in high taste and most of the time found good sanity checks. it is really not emphasized enough how far along fable's research and scientific capabilities are compared to other models, probably because it can't be felt by making threejs games

the bad: it needed a lot of pushing to go outside of what is currently known, and even then, it still clung on to existing hypotheses (storing spatial memories in synaptic weights have previously been hypothesized). also, getting it to communicate effectively to me, to layman, and to neuroscientists, was extremely triggering. it does not seem to understand a theory of mind over what people know and don't know and how to properly define words that are used and staying on a certain level of abstraction throughout the discourse. lastly, claude code should compact at ~400k, not 1M. responses go haywire to a point of no return after this.

@BrainsAndTennis on X · 14 September 2026

More work

1 October 2026Skeleton TennisA game: ten rallies drawn only as two skeletons, the ball, the court and the net. Name both players.Peter Wang25 September 2026Warcraft III as an RL environmentSo Nico and I present wc3env, an open-source gym environment for Warcraft III: The Frozen Throne.Peter Wang and Nico Christie18 September 2026A general and seven platoon leadersA composite agent harness composed of system1 (multi-agent jevs) + system2 (astra) out-micros an insane computer in my WC3FT RL env.Peter Wang and Nico Christie10 September 2026Turning the fly bisexualInspired by Kallman et al ('15), I blocked mAL output in a 166,606-neuron fly model and measured responses in 8 candidate P1 male courtship neurons.Nico Christie and Peter Wang