Anatomy of a memory access on the ET-SoC-1
Companion · interactiveEach level, interactively: L1, L2, L3, scratchpad, DRAM down to the transistors →Anatomy of a memory access, interactively: an animated diagram of each level, from the part that serves a load down to the transistors that switch, each access played step by step with the cycles and energy measured on three cards; every part marked documented, generic or unknown, with what each level is built from, SRAM or not, and what the team is asked.How finely can this chip show you where a load's time and energy go? For time, the answer is one cycle, for a single load. You can put any one line in L1, L2, L3 or DRAM, time one load of it, and split the result into cache lookups, mesh hops, the memory controller and the DRAM row state. A model built from those parts, fitted with one constant and eight memory shire positions, matches each line's DRAM latency to within ±3 cycles for 97% of lines (93–97% on each of three cards a week later, with the constants left as fitted). Of a typical DRAM load's ~300 cycles, only about 28 are the DRAM chip's own timing; the rest is paths on the chip. A DRAM row stays open until a refresh or a load to another row of its bank closes it; no idle timer was seen to close it.
For energy, the finest view is four destinations: the minion cores, the SRAM arrays, the mesh (NoC), and the rest of the board, which is mostly the DDR side. The first three are metered rails; the rest is the board total minus those three. The meters gave a new value every 133 ms on this card (126–189 ms on aifoundry2 and aifoundry1's card 1 in the three-card check, 223–322 ms on aifoundry3, slower the harder the host polls: every 0.1 s, at 10 Hz or at 20 Hz), so each energy number is an average over billions of identical loads.
Checked on three cards (26 September 2026). This page's measurements are from one card, aifoundry2, most of them from one launch of each probe on 19 September. On 26 September its claims were re-measured under a pre-registered plan, five passes of every latency probe on each of aifoundry2, aifoundry3 and aifoundry1's card 1 (the record). Of 66 claims tested here, this page counts 49 held, 11 corrected, 5 differ by card and 1 not confirmed; the hub’s scoreboard, 22 “proven on the cards”, 35 “a test behind it failed”, 8 “differs by card” and 1 “fewer than three repeats”. The charts in §2–4 put each card's latencies beside the 19 September session's. The registered test of the row-state claims in §5 read one series without the later build's 7-cycle timer offset and failed; re-referenced to each pass's own L1 hit, as everything else here is, those claims hold on every card (§5). The energies (the explorer in §1, §6 and the DRAM figure below) are the energy manual's (§4), measured in the same check on the three cards at 600 MHz (26 September).
Terms used on this page
The compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts) and an L1 cache; 32 minions form a shire, whose 4 MB of SRAM is a 512 KB L2, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB scratchpad. The 32 compute shires and the master, spare, I/O and PCIe shires sit on a 6×6 mesh (the NoC, 400 MHz), and eight memory shires, which hold the LPDDR4X controllers, sit outside it on two sides. Every cache line has one L3 home shire, chosen by its address bits, whose L3 slice holds it. Kernels run in user mode (U-mode) and firmware in machine mode (M-mode), and the service processor (SP) is the on-die core that reads the power sensors; more in the hub's glossary.
1. Trace one load
How this explorer works
This explorer is built from the measurements below. Pick the shire that issues the load and the L3 home shire of the address. Physical-address bits 10:6 give the home shire, and bits 8:6 give the memory shire, so the memory shire is always the home shire's low three bits. The bar breaks the latency into its parts. Latency comes from the model, which is within ±3 cycles for 97% of lines (the fastest of three loads of each line; 83% of single loads). The strip under the bar puts the path among the loads actually measured: for every home shire the model sits within 1 cycle of the measured median, as it should, since it was fitted to those loads, and, with the memory shires placed one hop outside the grid, the memory shire's position adds 12–84 cycles. Energy is the energy manual's for that level (§6): the mean over every pass on three cards at 600 MHz (26 September), above idle, per 64 bytes read (one line from the L2, the L3 or DRAM; two 32-byte vector loads from the L1), with each card's beside it. The L3's is at the mixed distances of its run, and the cells give what each mesh hop adds. All loads were timed from shire 0; for other requesting shires the explorer assumes the same 12 cycles per hop, which the on-chip communication report measured between every pair of shires on the same card. The three-card check (26 September) timed the same kind of loads from shires 7, 24 and 31 as well, and the model held from each of them on every card: L3 hits within ±4 cycles for 99.5–99.9% of lines and DRAM loads within ±3 for 91–96% (mean of five passes, per card and requester).
How to read the map
Map: marty1885's logical shire layout, as in the on-chip communication report. The unlabelled grey cells are not compute shires (the master and spare shires and the PCIe and I/O shires). In this map memory shires 0–3 sit above the grid and 4–7 below it; their positions come from the fit in §4 (19 September), except that memory shire 2's is a tie-break: its four home shires all sit in column x = 4, so x = 3 or x = 5 above the grid, or (6, 0) beside it, fit equally. The map appears to be the die turned a quarter: on the die (Heat per millimetre) and in the firmware's naming (0–3 west, 4–7 east) the memory shires are the two side columns. Shades show the model's latency from the requesting shire, an L3 hit or a DRAM load, 20 cycles per step. The requester has a solid outline, the L3 home a dashed one, and the memory shire that serves the home is filled. No route between them is drawn: the order in which the mesh routes a request was not measured. From the keyboard: Tab to the map, arrow keys move between shires, Enter sets the requester or the home.
2. The ladder, one load at a time
For each of 64 random lines, one hart loaded the line, evicted it to a chosen level, fenced, waited 300 cycles, and timed one load. Values are load-to-use cycles: the raw reading minus 5, an offset chosen so an L1 hit reads the 5 cycles the memory-hierarchy report's pointer chase measured on the same card. (On 19 September a timed no-op and a timed L1 hit both read 10 raw. The build the three-card check ran on 26 September reads an L1 hit at 17 raw, 7 cycles more, so its readings are re-referenced to its own L1 hit the same way. The offset moves every absolute number but none of the differences between levels.) The spread at L3 and DRAM is almost entirely distance on the mesh (§3, §4), not noise.
Bar: 10th–90th percentile; thin line: fastest to slowest; dot and number: median; all from aifoundry2 on 19 September. Under each bar, each card's median on 26 September: the median of its five passes' medians (hover or focus for their range). Evicting a dirty line and fencing, which is not a load, took 112–480 cycles (median 241), the write-back; one of 128 took 16, as fast as a clean one. On the three cards its median was 249–263 per pass. There the median of the loads left in memory was 299–309 from pass to pass, near one of two values, 299–300 or 307–309 (309 here), while L2 read 49 in every pass and L3 170–172.
Each level of this ladder drawn down to the transistors: the L1, latch RAM inside the minion, where a hit takes 5 cycles; the L2, in the shire cache; the L3 and the DRAM are linked from §3 and §4.
The numbers (min, median, 90th percentile, max, median in ns)
Row-open behaviour after an evict
Evicts keep working after the instruction retires, but none of the 256 loads timed straight after an evict to memory came back in L3 time: every one went to DRAM. What changes is the timing. With no fence, the load queues behind the evict (median 365, about 60 cycles above the §4 model; 352–361 per pass and about 52 above the model on the three cards). With a fence but no wait, 49 of the 128 loads (54–65 per pass on the three cards) find the row that the previous load of the same line opened still open, and read 6–12 cycles fast (median 10; minimum 213, on a line whose closed-row time is 225). In this run that was common for memory shires 0, 1, 2 and 4 (42 of 70 loads) and rare for 3, 5, 6 and 7 (7 of 58). On the three cards the difference had the same sign in every pass but was smaller and varied more: the row was found open for 25–34 percentage points more of the loads to memory shires 0, 1, 2 and 4 (mean of five passes per card; 48 points here), too uncertain for the registered test, which asked for at least 20 points at 99% confidence. The loads that missed the open row read a median 16 cycles above the closed-row time here and 9–12 on the three cards (mean of five passes, per card), whichever the memory shire: even after the fence, the load still waits a little behind the evict. With no fence that wait runs from about 10 cycles for memory shire 0 to 92–94 for memory shires 3 and 7. With the fence and the 300-cycle wait, none read fast (the line's row was last opened by the first load of its sequence, some 3,000 cycles and at least one refresh earlier): 115 of 128 land within ±6 cycles of the §4 model, which is the closed-row time (110–122 per pass on the three cards), and the rest are slower.
Cross-check against the pointer-chase report
The memory-hierarchy report's DRAM chases agree: 287–297 cycles at 600 MHz, on this card the day before and on three cards on 26 September. Every load here ran at a steady 600 MHz (§5's refresh period, 2,325 cycles in every pass, fits the DRAM's refresh interval only at that clock).
3. L3: 110 cycles plus 12 per hop
How the fit was done
Each line's L3 home shire is picked by physical-address bits 10:6. From shire 0, 1,500 random lines were timed as L3 hits. Latency depends only on the home shire's Manhattan distance on the mesh: 110 + 12 × hops, within ±4 cycles for 1,498 of 1,500 lines (the median of three loads per line; 1,309 are within ±2). On the three cards (26 September, five passes each) the fitted line was 110.5–110.6 + 11.99 × hops on every card, with 99.4–100% of lines within ±4 cycles. So a local L3 hit costs 110 cycles: ~48 for the L1 and L2 lookup and miss, if an L2 miss costs what an L2 hit does, and ~62 for the home slice. Each hop adds 12 cycles (20 ns) for the request and the reply. The memory-hierarchy report's pointer chases at 600 and 800 MHz, in one session on the same card (aifoundry3 never leaves 600 MHz: a boot service sets its TDP to 0 W at every boot), split the 110 by clock: about 72 cycles scale with the minion clock and about 61 ns do not (72 cycles + 61 ns = 109 cycles at 600 MHz), and a hop is 20 ns at either clock.
L2 hits in the same run come in two flavours: 37 cycles when the line is still in the bank's 8-entry read buffer, and 48 from the array. The split held on all three cards (26 September): 98–100% of lines read 44–52 cycles on their first L2 load and 98–99% read 40 or less on the repeats.
Drawn: the L3, a hit played from the requesting shire across the mesh to the home shire's data panel and back, with this fit's cycles; the two L2 flavours are a read-buffer hit and a load from the array in the L2's diagram.
4. The memory shire leg: 91 cycles plus 12 per hop
How the fit was done
The same 1,500 lines were also timed from DRAM, with no contention, and each line's L3 time was subtracted. What remains is the leg from the home shire to the memory shire and back. Every one of these loads found its row closed and paid an activate. Each came about 2,600 cycles or more after the line's row was last opened (by the untimed load that starts the line's sequence, or by an earlier pass), longer than a refresh period; home shire 0's lines read 225 cycles, the refresh series' closed-row time for a line of the same home and memory shire (226), not its open-row time (215); and no line read 11 cycles fast. The memory shire is picked by bits 8:6. The hardware counters confirm it: for all 64 lines tested, 16 DRAM loads raised the read count of the shire named by those bits, by 16 for 60 lines and by 15 for 4. The other seven shires' counts didn't move (one stray count of 1 in 448). In the three-card check the counter that rose was the one PA[8:6] names for 64 of 64 lines in every one of the 15 passes.
The fit uses one shared constant plus 12 cycles per hop. Assuming each memory shire hangs one hop outside the grid (the only positions the fit tried; moving all eight one hop further out fits equally well with a constant of 79), the best fit puts memory shires 0–3 above the grid and 4–7 below it, at x = 1…4. The total squared error is 2 cycles² over 32 home/memory-shire pairs; refitted pass by pass on the three cards, the constant reads 90–91. The whole DRAM load is then
latency = 110 + 12·hops(requester → L3 home)
+ 91 + 12·hops(L3 home → memory shire)
and the model's error is within ±3 cycles for 1,450 of 1,500 lines and within ±5 for 1,490 (of all 4,500 single loads, 3,721 (83%) are within ±3: taking each line's fastest drops refresh waits and a few cycles of jitter, though it keeps the timer glitches). Left exactly as fitted here, the model was within ±3 cycles for 97% on aifoundry2, 93% on aifoundry3, 94% on aifoundry1's card 1 in the three-card check (26 September; 1,500 lines in each of five passes per card), and from shires 7, 24 and 31 as well (§1). The other ten lines here are outliers: seven read about 128 cycles short (counter glitches the correction missed, §7), two are 44 and 71 cycles slow (probably refresh), and one is 29 cycles fast.
Of the 91 cycles (152 ns), the DRAM chip itself accounts for about 28: an activate (the activate-to-read delay tRCD, 11 cycles = 18 ns, §5), the 19.3 ns read latency, and the two 4.3 ns BL16 bursts that carry a 64-byte line (17 cycles together), from the controller's timing registers. The remaining cycles, at most ~63 (~105 ns), are the memory shire: controller queue, PHY, and the crossings between the 933 MHz DDR clock and the mesh. At most, because the fit puts the memory shires one hop outside the grid: each hop further out moves 12 cycles from this constant into the mesh legs, and latency alone cannot tell them apart. The frequency experiment in §9 would split this further.
Drawn: the DRAM, from the memory shire's two controllers and PHY down to the DRAM cell, with the ~63 cycles no measurement splits left hatched and asked of the team (the hub's rung 26).
5. Inside the DRAM: rows, banks and refresh
The rows, banks and refresh below, drawn down to the cell and the sense amplifier: a row hit, a row conflict and a refresh.
Which address bits share a row
A second load B was issued one cycle behind load A, with B = A with one physical-address bit flipped. Both were cold in DRAM, 150 random A per bit. When B shares A's row, or sits in another bank or channel, it rides along and arrives with A (~300 cycles). When B is in the same bank but a different row, the controller must wait for A's row to finish before opening B's: about +38 cycles (64 ns) for every bit from 18 up (38.5 here; on 26 September 36.9, 37.4 and 36.5 on aifoundry2, aifoundry3 and aifoundry1's card 1, the mean of five passes each, with aifoundry3's passes scattered more widely, 35–42.5). That matches the controller's mapping: bank = PA[12:10], column = PA[17:13], row from PA[18]. Issued one after the other instead, B pays about 9 cycles more (median) when it needs another row of A's bank (the precharge; tRP = 18 ns = 11 cycles), saves about 12 when it shares A's row, and costs nothing extra in another bank. On the three cards the first two were +10 and −11 in every pass, and the third within a cycle of 0.
Try it: decode any physical address
Which bits of an address pick each part of its path, as this page reads them: the L3 home shire, the memory shire, and the bank, column and row inside it. Hover, focus or tap a bit for what flipping it cost in the sweep above; click it, or press Enter, to flip it.
Refresh, caught in the act
Loading the same line from DRAM 19,000 times at random intervals, and folding the timestamps on a period of
2,325 cycles = 3.88 µs, shows LPDDR4X refresh (the controller is programmed for one every 3.87 µs, tREFI;
docs/research/counters-and-dram.md; the three-card check found the same 2,325 cycles in every one of its
15 passes):
- A load that arrives while refresh is running waits up to 208 extra cycles (347 ns; 208–209 per pass on the three cards). That's close to the 281 ns all-bank refresh time, and the wait shrinks as the load arrives later in the refresh window.
- Refresh also closes the open row. The next load pays an activate, 215 → 226 cycles: 11–12 cycles (about 19 ns), consistent with tRCD (18 ns), the activate-to-read delay. On the three cards the two clusters' means were 11.0–11.6 cycles apart.
- A loop that repeats the same work locks onto refresh: every fourth load is the slow one (below).
Why every fourth load is slow
The probe's locked series loads one line over and over, each turn the same four ops (evict, fence, time stamp, timed load), each fetched from the scratchpad in about 90 cycles. A turn whose load finds the row open takes 576 cycles, 214 of them the load, so four turns take 2,304, 21.4 cycles short of the refresh period. The first load after each refresh makes up the shortfall: it reads 235 cycles against 214 (medians), the activate (11.6) and about 9 cycles more, never 250 or more (at most 242), and the slow loads come exactly four apart in 5,999 of 6,000 gaps. The build the three-card check ran fetches an op in about 86 cycles and turns in 557–558 with the row open: four turns fall 93–97 cycles short, so its slow load lands well inside the refresh (89,470 of 89,521 slow loads read 250 cycles or more; median 308–311, at most 442), 93–97 cycles above an open-row load, with 24.6–25.0% of loads slow on each card and 98.1–100.0% of the gaps exactly four per pass. In every series the slow load's extra time equals the refresh period minus four open turns to within 0.6 cycles: the loop, not the DRAM, sets how slow the slow load is. Every comparison of two loads below counts the probe's own time.
How the model works
How long a row stays open, and what closes it
How the fit and the counts were computed
Open-row hits save 11 cycles. A row stays open until a refresh, or a load to another row of its bank, closes it.
That is what the controller's open-page policy with no idle-close timer implies
(docs/research/counters-and-dram.md). In the refresh series, 96% of loads whose previous load
came after the last refresh were row hits (3,556 of 3,713, after 450–2,000 idle cycles), against
4% of those with a refresh in between (201 of 5,382). The three-card check found
the same on each card, 95–96% against 3–4% (five passes per card), with its readings
re-referenced to each pass's own L1 hit. Its registered test of this split, which keeps only pairs far from the
refresh start, read the series without that re-reference, counted most open-row hits as closed, and failed;
re-referenced, the same test finds 99.9–100% hits with no refresh in between and at most 0.05% across one, on every
card. Pairs of loads to two lines
of one random row agree once the probe's own time is counted. Between the two loads the program spends about
180 cycles fetching its next ops from the scratchpad, on top of the first load's ~300 cycles, and
7% of loads land inside a refresh. Refresh alone then predicts about 72% hits with no added
idle time, 29% after 1,000 cycles and 7% after 1,500; the pairs gave 67%, 23% and
5% (60 pairs each). On the three cards the pairs again agreed with refresh alone.
One closer: refresh (or a conflict). With no refresh in between the row stays open; with one, it is closed; the random-row pairs fall where refresh alone puts them.
Notes on this chart
Refresh series: one line (memory shire 0), each load compared with the one before it; dot area follows the number of loads, and loads that were themselves held up by a refresh, or glitched, are left out. The hollow point of the no-refresh series (1,750–2,250 idle cycles) is uncertain: that close to a whole refresh period, whether a refresh fell in between depends on exactly where the refresh starts. Pairs: two lines of one random row, 59–60 pairs per delay, with 95% intervals; x is the programmed delay, which comes on top of about 180 cycles of probe time and the first load. From 3,000 to 20,000 cycles, 1 of 300 pairs found the row open; in the three-card check 0–3 of 300 per pass on aifoundry2 and aifoundry1's card 1, but 2–9 on aifoundry3.
The numbers behind the chart
6. Where the energy goes
The SRAM rail feeds the shire cache's arrays, drawn down to the cell in the diagrams of the L2, the L3 and the scratchpad, the 2.5 MB of each shire's cache that software manages. Which cell they use is asked of the team (the hub's rung 37): the documents call them SRAM, and the lab lead says the chip is not using SRAM.
What a byte read from each level costs above idle, at 600 MHz on three cards (the energy manual, §4, 26 September): an L1 hit 0.54 pJ (32-byte vector loads, random data), the L2 3.11, the shire's own scratchpad 2.25 holding zeros and 4.40 holding random data, the L3 14.7, another shire's scratchpad 5.10 and 11.8, and DRAM 114.6. A 64-byte line costs 199 pJ from the L2, 0.94 nJ from the L3 and 7.3 nJ from DRAM. The chart splits each way of reading a level between the four places the energy goes.
How the energy was measured
Both sets of numbers are the version-3 check's (26 September), at a steady 600 MHz on each of aifoundry2, aifoundry3 and aifoundry1's card 1, every burst against the idle on either side of it and corrected for the die's warming. The levels (the paragraph above, the explorer in §1 and the second table): hart 0 of every minion streaming 1 KB tensor loads, which skip the L1 but are cached in the L2 and L3, over a working set sized to one level (sized by design: no counter a user-mode kernel can read reports L2 or L3 hits and misses), six passes on each card. The contents of the L2, L3 and DRAM buffers were not set; the scratchpads were filled with zeros or random data before they were read. The L1's figure is the catalogue's vector loads, next. The paths (the chart and the first table) are the manual's catalogue, three passes on each card, each with zeros and with random data in the memory it reads: both harts of every minion re-reading the L1 with 32-byte vector loads, and 1 KB tensor loads from the shire's own scratchpad, from the L3 (a working set that fits it, its lines homed across the chip, so read through the mesh), from the scratchpads of shires exactly 1, 3 and 6 mesh hops away, and from DRAM. The split comes from the service processor's trace of the three metered rails, minion cores, SRAM (the L2/L3/scratchpad arrays) and NoC (the mesh), read at the end of each burst and corrected for how far the power-management controller's running average had risen; the rest = board power − those three rails. The rest is mostly the memory shires, DRAM, PHY and I/O. It also holds the regulators' delivery losses on the three metered rails (10–19% of the minion rail's power, by card; hub §4.2).
First measured here, on 19 September on aifoundry2 alone: hart 0 of all 1,024 minions walked a working set sized to one level with 8-byte loads, one per 64-byte line below the L1, in the same six-instruction loop for every level, for 6 s per pattern and twice each, and the power above the level just before each run was split by the same trace. What they showed about the meters: the host's board-power log read 14–25% higher than the trace in every one of the 12 runs, from the edges of each window, because the trace is the power-management controller's running average (τ ≈ 1.13 s on aifoundry2, 1.24 s on aifoundry3 and 1.15 s on aifoundry1's card 1; hub §4.1): its baseline reads high and its run average low.
The numbers: each path by rail on the chosen card, and every level on each card
pJ per byte above idle at 600 MHz (26 September): the mean over every pass on every card, the range those passes spanned, and each card's mean. The levels are hart 0 of every minion streaming 1 KB tensor loads (six passes on each card); the rows with the data named are the catalogue's (three passes on each card: the L1 by 32-byte vector loads, the L3 through the mesh and DRAM by 1 KB tensor loads). From the energy manual, §4.1–4.3.
What these measurements show:
- An L2 read is mostly SRAM. Reading the shire's own scratchpad, the L2's SRAM without the tag check, puts 68–69% of the power on the SRAM rail on aifoundry2 and aifoundry3, and the cores take 21–24% (they carry 80–87% of an L1 hit's). On aifoundry1's card 1 the SRAM rail carries 49% and no metered rail 29%, against 7–9% on the others: on that card the power on no metered rail rises 0.54 W with each watt on the SRAM rail, against 0.03–0.04 W on the others (the energy manual's fit of the unmetered power). The L2 itself costs 3.11 pJ per byte [1.42–4.99 over its passes], 199 pJ per 64-byte line.
- An L3 byte costs 4.7× an L2 byte (14.7 against 3.11 pJ/B), 0.94 nJ per line. Read through the mesh by tensor loads, the L3 puts 28–39% of the power on the SRAM rail, for the L2 miss, the L3 read and the L2 fill, and 39% on the mesh; random data cost 2.5× zeros (7.6 and 19.3 pJ/B on aifoundry2, 7.5 and 18.8 on aifoundry3 and 9.0 and 22.8 on aifoundry1's card 1).
- Each mesh hop adds 1.80 pJ per byte on random data and 0.67 on zeros, 115 and 43 pJ per 64-byte line: the mean of the manual's straight lines through another shire's scratchpad at 1–8 hops (1.72–1.91 and 0.62–0.73 card by card; Heat per millimetre splits it into bits that flip and ones carried). The mesh rail's share of the power grows with the distance: 27–28%, 39–40% and 50–52% at 1, 3 and 6 hops.
- A DRAM read costs 114.6 pJ per byte [89.0–141.3], 7.3 nJ per 64-byte line (111.5 on aifoundry2, 105.6 on aifoundry3 and 126.6 on aifoundry1's card 1); tensor loads from DRAM cost 94.6 on zeros and 132.6 on random data. 68–70% of a DRAM read's power is on no metered rail, mostly on the DDR side: the manual's fit of the unmetered power puts 72.7, 72.7 and 81.6 pJ per DRAM byte on aifoundry2, aifoundry3 and aifoundry1's card 1 in the DDR PHY, the I/O rail and the DRAM chips. The mesh rail carries 18–21%, because the request crosses the chip twice, the SRAM rail 10–11% and the cores 0–2%.
- The DRAM row pattern does not change the energy. The manual's row walks, 1 KB tensor loads by 32 harts, one per shire, each over a region of its own, together far larger than the L3, cost 155.9 pJ/B when every access finds its row open, 158.5 when every visit to a bank opens a new row and 156.4 streaming, on random data; no pattern differs from the others at 99% on any of the three cards (§4.3).
More: the DRAM read and its rows
7. A timer bug worth knowing
Before anything above was trustworthy, the cycle counter needed fixing. 16% of back-to-back read pairs (320 by +138 and 314 by −118, of 4,000) differed from the expected 10 cycles (12–16% per launch on the three cards on 26 September, where too every pair differed by 10, 138 or −118). The raw values show what happens: a read whose low 7 bits are 0 to about 10 is 128 too small, because the carry into bit 7 lands about 11 cycles late (the window's end moved by a cycle between launches). A simulation of the open RTL (core-et's Erbium branch at b38a1a3: the same Minion core lineage in a later MCU-class configuration, not the taped-out ET-SoC-1) shows the mechanism, with a 12-cycle window in the simulation (on this card the window covered at least 0–10 in one launch and 0–9 in the other, on 19 September; in the three-card check no single window fitted every pair in 11 of the 15 launches of the raw-pair probe, all five on aifoundry3): twelve counters share one adder that folds overflows in round-robin, and a read ignores the pending overflow (hub §3). This isn't erratum 1.23 (RTLMIN-6496, two harts reading counters at once): only one hart read here.
The fix, and where it still misses
The firmware's four-reads workaround doesn't fix it. Adding 128 to any read with low bits 0–10 does, for all 4,000 pairs here, and in 4 of the
15 launches of the three-card check; in the other 11 some pairs stayed 128 off. Inside a running kernel it
catches most but not all of them (57 of 4,000 corrected timed no-ops read 138 or −118 here; 0–186 of 4,000 per launch
on the three cards, the most on aifoundry3); the seven lines about 128
cycles fast in §4 are such misses, picked out by taking the fastest of three loads. Any code that times short
intervals with hpmcounter3 on these cards (all three show the bug) needs the same correction, and should
still drop ±128 outliers (see fixcyc() in
workloads/memprobe/kernel/memprobe.c).
8. What the chip lets you see
Everything here runs as an ordinary user-mode kernel through the runtime, on the card's installed firmware. The table lists every instrument that turned out to work, at its finest granularity, and what stays out of reach.
The full instrument table
| What | Instrument | Finest granularity | Notes |
|---|---|---|---|
| Time | hpmcounter3, readable in U-mode | 1 cycle (1.67 ns), one load | Reads 128 short when its low 7 bits are about 0–10 (§7). Even corrected in the kernel, about 1 timed interval in 70 was still ±128 off on 19 September, and 0–5% per launch on the three cards on 26 September: drop those, or take the median of repeated loads rather than the fastest. The standard cycle CSR traps: the firmware lets user mode read only hpmcounter3-8 (mcounteren = 0x1F8). |
| Where a line lives | evict_va (CSR 0x89f), one line | Put one line in L2, L3 or DRAM | Level code 1/2/3 leaves the line in L2 / L3 / memory, and code 0 does nothing. Asynchronous: a load issued straight after queues behind it, and may find its DRAM row still open (§2). A dirty line's evict costs a write-back (§2). |
| Two loads in flight | two lds back to back | Second load's arrival | Each minion has 2 miss handlers. This is how row conflicts show up (§5). |
| Which memory shire served a line | syscall 10 (SYSCALL_PMC_MS_SAMPLE) | Exact read count per memory shire | 16 DRAM loads of a line add 16 (15 for 4 lines) to one shire's read counter and 0 to the other seven (one stray 1), for 64 of 64 lines (§4; and in every pass on three cards). Syscall 9 does the same for the shire caches' read/write counters. |
| Core events | hpmcounter4-8 | Fixed events only | Retired instructions, L1 misses and I-cache events, summed over a neighbourhood of 8 minions. Choosing other events needs M-mode. |
| Board power | DM_CMD_GET_MODULE_POWER | 10 mW, new value every 133 ms on this card; in the three-card check 126–139 ms on aifoundry2 and aifoundry1's card 1 and 223–224 ms on aifoundry3 when read every 0.1 s, 156–159 and 262–264 ms under a 10 Hz sampler, 187–189 and 320–322 ms at 20 Hz | The host can ask 60 times a second, but the service processor refreshes the value once per loop pass. |
| Rail power | Service processor stats trace (dev_mngt_service -t SPST:extract) | 3 rails: minion cores, SRAM, NoC · one record per 133 ms on this card | The PMIC's own running average (τ ≈ 1.13 s on aifoundry2, 1.24 s on aifoundry3 and 1.15 s on aifoundry1's card 1; hub §4.1), copied once per service-processor pass, 133 ms here (Power and temperature §2). There is no DDR rail: the rest is the board minus the three rails. |
| Out of reach | — | — | DRAM activate/precharge counts, shire-cache hit/miss and occupancy, NoC counters (all need an M-mode syscall to choose events), per-event energy and faster power sampling. §9 says what would reach them; the full list is in docs/research/counters-and-dram.md. |
9. What would split it further
Ideas for going further
- Clock domains. Rerunning §3–4 with the minion and NoC clocks set apart would split each stage's cycles by clock (so far only the L3's 110 is split, §3); it changes a shared card's state, so it waits for approval.
- Static against dynamic energy. The same frequency sweep at fixed voltage gives each rail's C·V²·f slope. The DVFS report has since approached static power another way, with a law for idle power in temperature fitted on this card: idle power rises 0.65 W per °C at 80 °C, and the leakage in it at 80 °C is 20–29 W (the fit does not pin it closer; one card, and the three-card check did not get enough long idle runs to confirm it). The energy manual has since fitted the same law to the other cards' idle runs of that check, with the leakage's e-folding held at this card's: 0.70 W per °C on aifoundry3 and 0.97 W per °C on aifoundry1's card 1 at 80 °C.
- DRAM activate energy, and cache hit/miss counts. Both need counters that only M-mode can program. A user-mode syscall that sets them (about 60 lines) has been written and works in the simulator (sys_emu); running it on the card needs a signed firmware image, and neither the signing tool nor a test key is available, so whether this card would boot a rebuilt image is untested (the hub's firmware caveat).
- Asking the team. What latency alone cannot split, such as the memory shire's ~63 cycles, the stages of the L3's 110 and the storage cell of each level, is asked in the hub's improvement ladder (rung 26, rung 39, rung 37 and the rest of rungs 37–44), and the interactive diagrams link each unknown part to its rung.
10. Reproduce
Reproduce this
# Rebuild everything from the committed raw data (no card needed), from the repo root
D=docs/reports/data/2026-09-19-memprobe-aifoundry2
python3 workloads/memprobe/analyze_power.py $D/power --json $D/power/summary.json
python3 workloads/memprobe/analyze.py --data $D --out $D/summary.json
# the three-card check (26 September): every card's kept passes, reduced by the same functions
V=docs/reports/data/2026-09-25-claims-v3
python3 workloads/memprobe/analyze.py --v3 $V/raw --v3-passes $V/results/mem.passes.json \
--out docs/reports/data/2026-09-26-memprobe-3cards/cards.json
# the page; §1's and §6's energies are the energy manual's, read from its manual.json (build_report.py's docstring lists the fields)
python3 workloads/memprobe/build_report.py $D/summary.json docs/reports/data/2026-09-23-energy-manual/manual.json \
docs/reports/2026-09-19-et-soc1-memory-anatomy.html
Card runs (each under timeout 10)
scripts/deploy-lab.sh aifoundry2 workloads/memprobe # or build locally on the lab machine
mkdir -p build/memprobe-data && cd build/memprobe-data
python3 ../../workloads/memprobe/gen_ops.py timer --out .
python3 ../../workloads/memprobe/gen_ops.py ladder --out . --lines 64
python3 ../../workloads/memprobe/gen_ops.py decomp --out . --lines 1500 --reps 3 --pre-delay 1000
python3 ../../workloads/memprobe/gen_ops.py msmap --out . --lines 64
python3 ../../workloads/memprobe/gen_ops.py bits --out . --base 0x8040000000 --trials 150
python3 ../../workloads/memprobe/gen_ops.py refresh --out . --name refresh_jit --n 19000 --jitter 3000 --start 0x1000
python3 ../../workloads/memprobe/gen_ops.py refresh --out . --n 24000 --start 0x1000
python3 ../../workloads/memprobe/gen_ops.py pagetimeout --out . --row-bit 13 --trials 60 --delays 0,100,200,400,700,1000,1500,2000,3000,5000,8000,12000,20000
for p in t_raw t_glitch ladder decomp msmap bits refresh refresh_jit pagetimeout; do timeout 10 ../memprobe/host/memprobe_host --program $p.ops; done # each run < 0.2 s of card time
for p in l1 l2 l3near l3far dram_seq dram_row; do python3 ../../workloads/memprobe/gen_ops.py table --pattern $p --out-file $p.tbl; done
cd ../.. # back to the repo root
python3 workloads/memprobe/run_power.py --host-bin build/memprobe/host/memprobe_host --tables build/memprobe-data --out build/memprobe-power --seconds 6 --gap 4
mkdir -p build/memprobe-data/power && python3 workloads/memprobe/analyze_power.py build/memprobe-power --json build/memprobe-data/power/summary.json --csv build/memprobe-data/power/sp_stats.csv
python3 workloads/memprobe/analyze.py --data build/memprobe-data --out build/memprobe-data/summary.json
python3 workloads/memprobe/build_report.py build/memprobe-data/summary.json docs/reports/data/2026-09-23-energy-manual/manual.json docs/reports/2026-09-19-et-soc1-memory-anatomy.html
Raw results (u32 timings with their labels, the power log, the merged rail trace) are in
docs/reports/data/2026-09-19-memprobe-aifoundry2/. The runs used arena base 0x8040000000,
1 GB aligned, so arena offsets are physical address bits; so did every pass of the three-card check, on each card. Every card run was its own timeout 10
process, and the latency probes held the card for under 0.2 s each.
Versions: published 19 September 2026; revised 24 September 2026 (the DRAM split, the row and energy comparisons) and 25 September 2026 (what closes a DRAM row, the model against the measurements, the order of the sections, and the charts). 25 September (version 3): the page says that every measurement is aifoundry2's and gives the runs behind each energy; the energy figure at the top is now the energy manual's, from both cards; §2 no longer singles out memory shires 3, 5 and 7; the memory shire's share of the 91 cycles is an upper bound; and the idle law is quoted as a slope and a range. 26 September (version 4): the latencies re-checked on three cards under the pre-registered plan (the note at the top); each card's medians beside 19 September's in §2–4; the counter window, the locked loop, the row conflict (about +38 cycles, not +40) and the evict ladder's open-row split corrected; the board-power period given per card; the energy manual's figures (the note on energy, the DRAM figure at the top, §6 and the energy chart) are its version-3 values from three cards. The record is docs/reports/data/2026-09-25-claims-v3. 27 September: the DRAM chip's share of a load is about 28 cycles (a 64-byte line is two bursts; was 25–28) and the memory shire's at most ~63; the explorer's row-conflict and refresh waits are the three-card values (38 and 208 cycles, were 40 and 210); the memory-shire leg charted per card; detail folded into collapsible sections. 27 September, later: why every fourth load of the locked loop is slow, charted with a model that takes four numbers from the refresh series (the slow load's extra time is the refresh period minus four open turns, within a cycle in every series; the page had said that none of this session's slow loads landed inside the refresh); the bit chart's row-conflict label reads the measured +38.5 cycles (was rounded to +40); the no-fence loads' median excess over the §4 model is 60 cycles (was 58); the board value's refresh in the three-card check includes its 20 Hz sampler (126–189 and 223–322 ms; was 126–159 and 223–264). 28 September: a link to the companion page, Anatomy of a memory access, interactively, under the title, and from §2–§6 and §9 to its diagram of each level and to the hub's asks Later that day, the review's cuts (§6's host-log note, §9's clock-domain idea, §2's cross-check; the DRAM figures said once, in the energy note at the top); the note's counts given by both rules. 28 September, later: the energy manual's three-card values replace the 19 September runs on the page itself: §6 gives every level on each card and splits each way of reading a level between the rails, card by card and with zeros or random data; the explorer's energy is the manual's level; the DRAM figure at the top and the regulators' delivery loss (10–19% by card, was 18–20%) are the three cards'. The 19 September runs are kept as a note on how they were measured, in §6.
Related reports
- Limits of observability, the hub — every instrument used here, what the power meters can and cannot see, and what would extend them.
- Anatomy of a memory access, interactively — this page's companion: each level (L1, L2, L3, scratchpad, DRAM) drawn and animated down to the transistors, every part with its source.
- Memory hierarchy — the pointer chase and the bandwidth of every level, measured the day before this page.
- On-chip communication — the shire map used here, and 20 ns per mesh hop measured between every pair of shires on aifoundry2.
- The energy manual, §4 — energy per byte at every level of the hierarchy, with ranges over repeated runs on three cards, and each path by rail: the source of the energies in §1 and §6, with the writes, gathers and scatters this page leaves out.
- Heat per millimetre — the energy of one mesh hop, split into bits that flip and ones carried.
- One hot line stops a shire — uses the
PA[10:6]home-shire mapping found here to home a contended DRAM line in a chosen shire. - Power and temperature — the power meters used in §6: what each one reports, how often, and how it averages.