M9 step 1: memory cheap wins (S9) #3

Open
opened 2026-10-02 00:16:05 +00:00 by jhgaylor · 0 comments
Owner

Blocks: #10

Split out of M9: S9 found the trigger already holds, but these cheap fixes come before any parking. Rerun the S9 benchmark afterwards.

From PLAN.md, section "M9: parking (scale)".


  • S9: measured 2026-10-01; the trigger already holds, but parking isn't the
    first fix.
    See spikes/s9-memory.
    • Idle cost: an idle, empty pane costs 3.1–3.3 MB of daemon RSS (50
      panes: 172 MB; 500: 1.6 GB, with 5 threads per pane). Shims and bash add
      about 1.7 MB more, so 500 idle shells cost about 2.5 GB in all.
    • Scrollback costs about 1.7 KB per row at 200 columns: 10k lines is
      33 MB per pane, and the 64 MiB cap is about 160 MB.
    • Memory isn't given back after panes close. glibc keeps it: 500 panes
      closed down to 1 still hold 191 MB.
  • Step 1: cheap wins, before any parking. Then rerun the S9 benchmark.
    • Drop Zig's 256 KiB per-thread signal stack in libghostty, which glibc
      gives every thread: 1.3 MB per pane. It's a one-line patch
      (ghostty-no-signal-stack.patch) to carry and send upstream.
    • Avoid libghostty's ReleaseSafe page fill: 1.5 MB per screen.
      ReleaseFast plus the patch measured 0.48 MB per idle pane, which already
      meets M9's done bar. But ReleaseFast drops safety checks on untrusted
      program output, so it's a decision (open); an upstream fix that avoids
      the fill would remove the trade.
    • Fix the malloc mmap threshold (mallopt at startup, or another
      allocator) and call malloc_trim after a pane closes. That saves about
      13 MB per pane at 10k lines, and memory comes back.
    • The output ring holds 4 MiB, not 2 MiB, because Ring::push extends
      before draining. Set it to 1 MiB (replay never uses more than
      MAX_REPLAY_BYTES) and drain first.
    • Lower the in-memory scrollback cap from 64 to 16 MiB (about 28 MB per
      pane). Full history is on disk.
    • Fold the reaper thread into the wait thread, and use a tiny shim
      binary instead of re-running the 21 MB daemon (about 0.3 GB at 500
      panes).
  • Step 2: parking, only if panes with real history still blow the budget.
    S9 found nothing that argues for PTY or client-buffer parking. Parking
    can't save the shells and shims either (0.75–0.87 GB at 500 panes).
  • Stalled clients: a client that stops reading cost about 28 MB in one
    burst, and up to about 64 MB per client from the 1,024-frame queue. Cap the
    queue in bytes, not frames.
  • CI benchmark: S9's bench.py at 50 idle panes (3.07–3.11 MB per pane
    over four runs) is the regression test M9 wants. Leave headroom.
**Blocks:** #10 Split out of M9: S9 found the trigger already holds, but these cheap fixes come before any parking. Rerun the S9 benchmark afterwards. _From PLAN.md, section "M9: parking (scale)"._ --- - **S9: measured 2026-10-01; the trigger already holds, but parking isn't the first fix.** See [spikes/s9-memory](https://git.inevitable.fyi/jhgaylor/illogical/src/branch/main/spikes/s9-memory/README.md). - **Idle cost:** an idle, empty pane costs 3.1–3.3 MB of daemon RSS (50 panes: 172 MB; 500: 1.6 GB, with 5 threads per pane). Shims and bash add about 1.7 MB more, so 500 idle shells cost about 2.5 GB in all. - **Scrollback** costs about 1.7 KB per row at 200 columns: 10k lines is 33 MB per pane, and the 64 MiB cap is about 160 MB. - **Memory isn't given back after panes close.** glibc keeps it: 500 panes closed down to 1 still hold 191 MB. - **Step 1: cheap wins, before any parking.** Then rerun the S9 benchmark. - **Drop Zig's 256 KiB per-thread signal stack in libghostty**, which glibc gives every thread: 1.3 MB per pane. It's a one-line patch (`ghostty-no-signal-stack.patch`) to carry and send upstream. - **Avoid libghostty's ReleaseSafe page fill:** 1.5 MB per screen. ReleaseFast plus the patch measured 0.48 MB per idle pane, which already meets M9's done bar. But ReleaseFast drops safety checks on untrusted program output, so it's a decision (open); an upstream fix that avoids the fill would remove the trade. - **Fix the malloc mmap threshold** (`mallopt` at startup, or another allocator) and call `malloc_trim` after a pane closes. That saves about 13 MB per pane at 10k lines, and memory comes back. - **The output ring holds 4 MiB, not 2 MiB,** because `Ring::push` extends before draining. Set it to 1 MiB (replay never uses more than `MAX_REPLAY_BYTES`) and drain first. - **Lower the in-memory scrollback cap** from 64 to 16 MiB (about 28 MB per pane). Full history is on disk. - **Fold the reaper thread into the wait thread,** and use a tiny shim binary instead of re-running the 21 MB daemon (about 0.3 GB at 500 panes). - **Step 2: parking, only if panes with real history still blow the budget.** S9 found nothing that argues for PTY or client-buffer parking. Parking can't save the shells and shims either (0.75–0.87 GB at 500 panes). - **Stalled clients:** a client that stops reading cost about 28 MB in one burst, and up to about 64 MB per client from the 1,024-frame queue. Cap the queue in bytes, not frames. - **CI benchmark:** S9's `bench.py` at 50 idle panes (3.07–3.11 MB per pane over four runs) is the regression test M9 wants. Leave headroom.
Sign in to join this conversation.
No description provided.