# Octane Blender addon 31.10 — race/async investigation

Repo purpose: hold the **pristine vendor addon** plus local hardening work for the
OctaneServer `SIGSEGV` crash on camera move.

Baseline: `octane_blender_addon-31.10-stable.zip`
(`bin/linux/octane_blender.so`
sha256 `9dba1ce9732c3d8782de1452871271a7e70dcbb09c9a19e804dd7225360ff00e`,
md5 `fb7bb263108aa53c2eddd6e147744890`), byte-identical to the addon installed
in `~/.config/blender/5.2/`. (See Errata: the earlier note labelled the md5 as
a sha256.)

## What actually crashes (from core dumps)

Six OctaneServer cores on 2026-09-21, all `SIGSEGV (SEGV_MAPERR)`:

1. **Camera batch-update path** (5 cores, the reported bug):

```
#0  Octane::ApiItem::isNode() const              liboctane.so
#1  OctaneBlenderServer::batchUpdateConsumerLoop()   OctaneServer +0x1263dc
    rdi = 0x0     (this == nullptr)
=>  movzbl 0x38(%rdi),%eax      <- scalar load, not SIMD
```

Disassembly of the call site (`0x5263c3`–`0x5263d7`):

```
call  ApiProjectManager::rootNodeGraph()
mov   %rax,%rbx
mov   %rbx,%rdi          ; rdi = root node graph
call  ApiItem::isNode()  ; no null check
```

`rootNodeGraph()` returned NULL and the consumer loop dereferenced it
unconditionally.

2. **XML recursion path** (1 core, 11:25):

```
#0  liboctane.so + 0x1e5a775
#1  OctaneServer::RecurseItemXml(...tinyxml2...)
```

Same class of fault, different unchecked-null path.

## What is and is NOT fixable here

**Not fixable in this repo.** The null dereference is in OTOY's *compiled,
closed-source* server (`OctaneServer` + `liboctane.so`, both stripped
prebuilts, not owned by any package). The batch pipeline lives in native code:

- `OctaneBlenderClientBase::SendBlenderUpdateRequest`
- `OctaneBlenderClientBase::EnsureAllUpdatesProcessed`
- the `grpcOctaneServer::BatchUpdate*` protobuf types

The `Failed to send BatchUpdateRequest` message the user sees comes from the
**native module** (`octane_blender.so`), not from Python. No Python edit can add
a null check inside `batchUpdateConsumerLoop`.

**Partially fixable here.** The Python addon (this repo is Python-only; no
`.cpp`/`.h` shipped) drives *when* updates are dispatched. It can shrink the
window where the server's consumer thread runs against a not-yet-built root
node graph. That is mitigation, not a cure.

Notably the addon already has some protection:

- `OctaneBlender().try_lock_update_mutex()` / `unlock_update_mutex()`
- `min_viewport_update_interval` throttle (default `0.04s`)

and `send_batch_updates()` is a **no-op stub** (logs, checks `self.enabled`,
returns — no native call). `set_batch_state()` has **zero Python callers**.
So batch dispatch is entirely native-side.

## The hardening applied

`core/session.py` → `EngineSession.view_update`:

Add an explicit **initialization-handshake gate**. The server is only safe to
consume camera-driven updates after the first sync has completed *and* the
change manager has been flushed and `SceneState.INITIALIZED` acknowledged. The
stock code computed `is_first_update` but allowed subsequent updates to dispatch
immediately after the initializing sync returned, before the handshake was
confirmed.

Patch adds `_octane_init_handshake_done`, set only after
`update_change_manager(True, False, True)` + `set_scene_state(INITIALIZED)`
succeed, and defers (rather than drops) any update until then. Deferred updates
are re-queued via `engine.tag_update()`, so no scene change is lost.

### Honest expectation

- Reduces frequency/duration of the vulnerable window.
- Does **not** guarantee the crash is gone — the server still dereferences a
  possibly-null root node graph in `batchUpdateConsumerLoop` and in
  `RecurseItemXml`.
- Real fix required from OTOY.

### Status update — 2026-09-21 22:00 re-verification

The hardening patch is **installed and live**, and the crash **still occurs**.

The gate cannot close this window, by construction:

1. `if not is_first_update and not self._octane_init_handshake_done:` — the
   **first** sync dispatches *unguarded*, which is the moment the root node
   graph is least likely to exist.
2. `_octane_init_defer_limit = 120` bounds the deferral: after 120 skipped
   updates the dispatch proceeds regardless.
3. The consumer loop is native OTOY code. Python only chooses *when* updates are
   handed over; it cannot add a null check or order the server's internals.

So "reduces frequency/duration of the vulnerable window" is the whole of it.

Verified in this pass:

- `~/.config/blender/5.2/scripts/addons/octane/core/session.py`
  sha256 `2aaaba296da1e14ccd3b157710331a98a42812345b62ea9cb794049829029821` —
  byte-identical to this repo's patched file, installed `2026-09-21 11:40:33`.
- Crashes at **18:22:20, 21:39:20 and 21:41:08** all post-date that install.

Fresh `gdb` on core `38926` reproduces the documented frame exactly:

```
#0  Octane::ApiItem::isNode() const                 liboctane.so
#1  OctaneBlenderServer::batchUpdateConsumerLoop()  OctaneServer +0x1263dc
    rdi 0x0            rbx 0x0
    si_addr 0x38
=> 0x00007f43a3d7f8f0 <_ZNK6Octane7ApiItem6isNodeEv>:  movzbl 0x38(%rdi),%eax
```

Kernel `dmesg`, this boot, pins both instants (local time):

```
Sep 21 21:39:20 rossa kernel: OctaneServer[7280]: segfault at 0  ip 00007f2c7185a775 ... in liboctane.so[1e5a775,...]
Sep 21 21:41:08 rossa kernel: OctaneServer[39101]: segfault at 38 ip 00007f43a3d7f8f0 ... in liboctane.so[d7f8f0,...]
```

`at 0` (`liboctane.so+0x1e5a775`) is the `RecurseItemXml` path; `at 38`
(`+0xd7f8f0`) is the `ApiItem::isNode()` / `batchUpdateConsumerLoop` path. Both
`SEGV_MAPERR`. Note the pair alternates between the two unchecked-null sites,
which is what you expect from a shared root cause rather than one bad call site.

## Server-side log evidence (2026-09-21)

Octane's own log: `/var/tmp/octane_log.txt` (33,198,005 bytes; appended across
sessions because `logAppendToFile` is on). **The bracketed prefix is UTC; the
message body is local (+05).**

Immediately before the 21:39:20 crash the log's last activity is a burst of
`ITEM_VALUE_CHANGED` records on camera/film pins:

```
[16:39:19.943]  ITEM_VALUE_CHANGED 1690 : 15558 NT_INT Int value
[16:39:19.943]    pinName dolly_zoom_enable
[16:39:19.943]  ITEM_VALUE_CHANGED 1691 : 15559 NT_FLOAT Float value
[16:39:19.943]    pinName dolly_zoom
[16:40:03.322]  Started logging on 21.09.26 21:40:03      <- restart
```

Findings:

- The crash instant (**21:39:20**) has **no error line**. The process died
  mid-update; logging stops at `[16:39:19.943]` and nothing follows.
- The server restarted at **21:40:03** (banner), announced
  `Server listening on 0.0.0.0:50051` at **21:40:09.412**, and crashed again at
  **21:41:08** — a **~59 s lifetime**. The file's final line is that listen
  banner; the second crash also logged no error.
- The log carries **6,458** `connected to null` and **4,896** `missing pin`
  lines, e.g. `ERROR : RecurseItemXml() : missing pin premultiplied_alpha`. Not
  proof of the fatal update, but the same "prerequisite not ready" class as the
  null root node graph, and it correlates with the `RecurseItemXml` core.

Session banners this day (local): 11:41:37, 12:36:46, 18:23:09, 19:30:10,
20:53:19, 21:40:03. Every session reports `OctaneRender Prime 2026.4
(15030000)`.

## Retained cores (as of 2026-09-21 22:00)

`coredumpctl list OctaneServer`:

| Local time | PID | Signal | Core size |
|---|---|---|---|
| 2026-09-21 18:22:20 | 85861 | SIGSEGV | 998.6M |
| 2026-09-21 21:39:20 | 6772 | SIGSEGV | 882.4M |
| 2026-09-21 21:41:08 | 38926 | SIGSEGV | 173.8M |

Only **three** OctaneServer cores survive now; the earlier "six" count came from
an older boot and those are gone. `coredumpctl` holds **4.2 GB** total (15 cores
including Blender), and `/` is at **91%** — prune before the next capture
(`sudo coredumpctl prune`).

Also: **OctaneServer is not running**, and `/tmp/octane_server_1000.lock` is
**stale**, still holding the dead PID `38926`. Clear it before restarting.

The 21:41:08 core (38926) was dumped and inspected with `gdb`; temporary copies
under `/tmp` were deleted afterwards (the store is centralised in
`/var/lib/systemd/coredump`).

## Interposition experiment (LD_PRELOAD guard) — negative result

Tested 2026-09-21 22:00–22:30. Code in `tools/octane-guard/`.

**Why it looked promising.** The crashing calls go through OctaneServer's PLT
(`call _ZNK6Octane7ApiItem6isNodeEv@plt` at `0x5263d7`), the symbols are
`GLOBAL DEFAULT` in `liboctane.so`'s `.dynsym`, and **neither binary has
`BIND_NOW``** — so a preloaded definition wins at GOT resolution. Confirmed with
a PLT-linked probe: calling the real `isNode()` with `NULL` gives `SIGSEGV`
(exit 139) bare, and survives with the shim loaded.

**v1 — wrap `isNode()` only.** Live run: the shim fired
(`isNode: this==NULL -> safe default`) and the server died anyway at the **next**
unchecked accessor:

```
#0  Octane::ApiItem::toNode()            liboctane.so +0xd7f990
#1  OctaneServer::RecurseItemXml(...)    OctaneServer +0x137834
#2  OctaneBlenderServer::batchUpdateConsumerLoop()
    rdi 0x0   si_addr 0x38
```

The null is not one bad call; it is an `ApiItem*` threaded through a whole
accessor chain.

**v2 — wrap the accessor chain.** Added `isNode`, `isLinker`,
`is{Input,Output}Linker`, `toNode`/`toGraph` (const + non-const),
`uniqueId`/`version`/`outType`, `name`, `hasOwner`. Self-test passes for all with
a null `this`. Result in session: **no crash, but the viewport froze.**

Guard counts, 308 each:

```
308 _ZNK6Octane7ApiItem6isNodeEv
308 _ZN6Octane7ApiItem6toNodeEv
308 _ZN6Octane7ApiItem7toGraphEv
  0 uniqueId / name / version
```

`uniqueId`/`name`/`version` never firing shows the null item is dereferenced
right at the `isNode`/`toNode`/`toGraph` trio — the code bails out before
reaching the later accessors. The null's **origin was not proven**;
`ApiProjectManager::rootNodeGraph()` was only the leading hypothesis, and in v3
it never returned NULL at all. Each suppressed accessor makes
`batchUpdateConsumerLoop()` **drop that update** — 308 dropped updates is a
frozen viewport. The crash had been converted into a silent failure, which is
worse.

**Pass-through verification** (rules out the scariest explanation, that the
wrapper was breaking *all* calls): non-null inputs reach the real liboctane
implementation identically with and without the shim
(`isNode=1`, `toNode=self`, `uniqueId=0xab`). The shim only intervened on null.
Make target: `make verify`.

**v3 — retry `rootNodeGraph()` instead of suppressing accessors.** Bounded
250 ms wait-retry, returning the real pointer when it appears. Result: it
returned **non-NULL on the first call** (`first non-NULL: 0x7f9d61e3a8b0`) and
the shim fired **0 times** — functionally inert. The operator then reported
**missing 3D models in the scene**.

**Verdict: the shim does not touch Blender data, and in v3 it did not run at
all.** It was only ever `LD_PRELOAD`ed into `OctaneServer`; `/proc/6949/maps`
shows zero guard mappings inside `blender`, and a shim cannot modify a `.blend`.
The missing geometry is therefore either (a) residue from v2's 308 dropped
updates, or (b) the underlying sync defect, which is also what the `missing live
child` bursts point at. Not resolved.

**Final test (22:32–22:35), and why it proves nothing about the shim.** Shim
reloaded (v3) and Blender **restarted** (PID 6949 -> 125302), giving a fresh
session. The operator then moved the camera: **no crash**, server healthy.

But the guard log shows the shim fired **zero times** in that whole test:

```
[octane-guard] rootNodeGraph: first non-NULL: 0x7f7301e3a860
rootNodeGraph events:  1
RECOVERED:             0
STILL NULL:            0
accessor suppressions: 0
```

The `rootNodeGraph` wrapper firing once proves interposition is genuinely live
in that process, so zero accessor fires means **no null ever reached the guarded
functions** — the crash path simply was not executed. The crash therefore did
not happen because its precondition was absent, **not** because the shim
defended against it.

The variable that actually changed is the **fresh Blender session**. Combined
with the earlier facts (the stock server crashed on its own at 22:31:35, and the
shim was never loaded into Blender), this is consistent with the crash
correlating with **Blender-side session state** rather than with anything the
shim does. One camera-move sample is not conclusive.

**Consequence for the decision.** Keeping the shim loaded is a safety net whose
value is unproven for this bug, and it carries a real cost: when a null *does*
occur it converts a loud crash into a silent dropped update (the frozen viewport
seen in v2). Because `OCTANE_GUARD_VERBOSE=1` logs every suppression, that
failure mode is at least recorded in `/tmp/octane-guard-err.log` rather than
being invisible.

**Second test (22:32–23:00) — shim removed at operator request.** Shim was
reloaded (v3) with a fresh Blender session (PID 125302). Result: **no crash**,
but the operator reported the viewport **stopped rendering after ~940 samples**
and asked for the shim to be removed.

Shim log for that entire run was **3 lines**: the startup banner twice plus one
`rootNodeGraph: first non-NULL` line. **0 accessor suppressions, 0
`rootNodeGraph` NULLs, 0 `RECOVERED`.** The shim therefore never intercepted a
single call, so it **cannot have caused the ~940-sample stall** — it was
functionally inert for the whole session.

Removed on request: `octane_guard.so` deleted (rebuildable with `make`), no
process had it mapped, and there are no shell/profile/`environment.d`
references. (`OCTANE_GUARD_WAIT_MS`/`VERBOSE` were only ever session-local
exports in the launching shell.)

**The ~940-sample stall is a separate, unexplained issue.** It coincides with
`renderPassInfoMaxSamples` being changed mid-session (`ITEM_VALUE_CHANGED …
renderPassInfoMaxSamples` at 17:47, 17:49, 17:50, 17:58), which is worth
following up, but no causal link has been established.

**Conclusion: interposition is not a viable fix for this bug.** It contains the
crash but replaces it with dropped updates / stale scene state, and it addresses
symptoms rather than the race. Reverted to stock. The shim is not installed
anywhere persistent; the source is retained under `tools/octane-guard/` only as
a record of the negative result.

## Recommended actions (highest confidence first)

1. **Install the matching 31.9 pair** (server + addon together) to test whether
   31.10 is a regression. Artefacts on disk (corrected path — see Errata):
   `~/wolf/downloads/octane_blender_addon-31.9-stable.zip`,
   `~/wolf/downloads/OctanePrime_for_BlenderOctaneAddon_Linux_31.9-stable.zip`.
   This is now the single highest-value experiment, because it is the only one
   that separates "31.10 regression" from "race present in all 31.x".
2. **Start the server before Blender**; let the scene finish its first sync
   before moving the camera. Given the ~59 s second-crash lifetime above, also
   let a freshly started server settle before reconnecting Blender.
3. Apply this hardening patch and re-test. **Done and re-tested — patch is
   installed and the crash reproduces.** Keep it for the reduced window, but do
   not treat it as a fix.
4. **Report to OTOY** with both backtraces.

## Errata

Corrections to the original revision of this document:

1. **Native module hash.** The baseline was written as
   `sha256 fb7bb263108aa53c2eddd6e147744890`. That is 32 hex characters — an
   **md5**, not a sha256. The md5 matches the installed module exactly; the
   real sha256 is
   `9dba1ce9732c3d8782de1452871271a7e70dcbb09c9a19e804dd7225360ff00e`.
   The module is byte-identical to `octane_blender_addon-31.10-stable.zip`, so
   the native bridge is **pristine and unmodified** — the fault is not in it.
2. **31.9 artefact path.** The zips live in `~/wolf/downloads/`, not
   `~/Downloads/` (which holds only the 31.10 addon zip and the 31.10 Prime
   installer).
3. **Core count.** "Six OctaneServer cores" was accurate for the earlier boot;
   only three remain retained. See the table above.
4. **Crash times.** Prefer the kernel instants (21:39:20 / 21:41:08) when
   correlating with `/var/tmp/octane_log.txt`; `coredumpctl` records the same
   events a little later at 21:39:43 / 21:41:12, which is when the dump was
   written rather than when the fault occurred.

## Measurement notes for the next capture

When reproducing, capture all of the following so the next pass does not have to
re-derive them:

- `journalctl -k -b | grep -i octane` — the kernel's own segfault line gives the
  fault address (`at 0` vs `at 38`) and which unchecked-null site fired.
- `coredumpctl info <pid>` plus a `gdb` register check
  (`info registers rdi rbx`) to confirm `rdi == 0` rather than a corrupted
  pointer.
- `grep -nE "\[<utc-hh:mm>" /var/tmp/octane_log.txt` — recall the prefix is
  **UTC** while bodies are local, so subtract the +05 offset when correlating.
- Whether the crash follows a **server restart**; the 21:41:08 core came ~59 s
  after the 21:40:03 banner, which suggests reconnect-after-restart is a
  distinct high-risk trigger worth isolating.

## Reporting template for OTOY

- OS: CachyOS (Arch), kernel 7.2.4-3-cachyos
- GPU: NVIDIA RTX 5060 Ti 16GB
- Driver: 615.71.09
- Octane: 2026.4 / 31.10-stable (Prime), build 15030000
- Host: Blender 5.2.2 LTS
- Crash: SIGSEGV SEGV_MAPERR, `this == nullptr` in
  `ApiItem::isNode()` from `batchUpdateConsumerLoop()`; second core in
  `RecurseItemXml()`
- Trigger: camera movement in the viewport after scene load
