Introduction
Kóoch is a GPU-driven game engine written in Rust, with an editor.
The rendering path is a Nanite-style GPU-driven meshlet pipeline: the hot loop runs in compute on the GPU, and the CPU only coordinates. The ECS stays on the CPU — that is a deliberate, settled decision, not a stage on the way to something else. What goes to the GPU is graphics, effects and their derivatives; physics will follow if and when rapier3d gains GPU support.
This is a personal, experimental project — not a stable production tool. The documentation reflects that: it explains what is, not what should be. Where something is missing or broken, this book says so and links the issue.
Audience
Two readers, with overlapping needs:
- Engine users — anyone writing a game on top of Kóoch. Start with Your First Project, then The Editor.
- Engine contributors — anyone touching
crates/*. Start with Crate Graph and the Decisions Log.
Status
Early development. The pieces work together end to end — window, ECS, scene serialisation, meshlet renderer, sky, physics, editor — but the feature surface is narrow on purpose and APIs break freely.
What works today:
- An 18-crate Rust workspace plus a facade, edition 2024.
- GPU-driven meshlet rendering with a LOD chain, plus glTF mesh loading. A frame is a list of views, so the editor’s viewport and the game’s camera render from one stage.
- Cook-Torrance lighting driven by the light components — see Lighting. New as of #441; before it, the renderer painted the world-space normal as colour and a scene with lights looked exactly like one without.
- Rigid-body physics on rapier3d: colliders, joints, collision events, sensors, materials, and custom gravity fields that sum.
- Procedural sky and volumetric clouds.
- Scene serialisation (
.scene, RON) driven by reflection, with more than one scene loadable at once. - Input as data: an action is an asset, with composites and processors, editable in a panel.
- An editor: viewport, hierarchy, Inspector, Console, asset browser, drag-and-drop, dockable layout, undo/redo, project Hub, a Game panel beside the View panel, and Play/Stop that snapshots and restores the authored world. Hovering a field shows its doc comment, units included.
- A project’s own components and systems, written in Rust, loaded into the editor as a
dylib.
What does not work yet, stated plainly:
- No hot reload. Seeing a code change means rebuilding and reopening the editor (#648).
- No build button.
cargo buildis yours to run (#158). - Reflection is shallow. No
Vec<T>, noHashMap, no user enums in components (#649). - No shadows. Lit with no shadows is where the renderer is: better than a normal painted as colour, and it cannot tell you where anything is touching (#476).
- No global illumination, which is why punctual light defaults are larger than physics says they should be (#450). The Lighting page explains the trade rather than hiding it.
Stack
| Layer | Crate / Library |
|---|---|
| GPU | wgpu 29 (Vulkan / DX12 / Metal) |
| Windowing | winit 0.30 |
| Math | glam 0.33 |
| Physics | rapier3d 0.34 |
| Audio | kira 0.9 |
| Input | gilrs 0.11 (gamepad), winit (keyboard/mouse) |
| Gameplay code | Plain Rust — native plugin via kooch_plugin_api |
| Editor UI | egui 0.35 + egui_dock 0.20 |
| Mesh | gltf 1.4 |
| Serialisation | serde 1 + ron 0.8 |
License
All Rights Reserved. Copyright (C) 2025-2026 Matías Galarza (“Lobinux”, lobinuxsoft).
The repository is public so the work can be read and so the project can use branch protection
and Pages. That is not a licence to use it: see
LICENSE.md.
How to read this book
User Guide and Scripting are the public API surface. Architecture covers internals. Reference holds the long-form material: the Decisions Log is a chronological record of architectural choices, why they were made, and what was traded away.
Two files in the repository outrank this book when they disagree:
docs/MEMORY.md
is canonical on decisions, and
docs/ROADMAP.md
is canonical on order.
Getting Started
You need a recent Rust toolchain (edition 2024) and this repository cloned somewhere.
cargo run -p kooch_editor
That opens the Hub, where projects are created and opened. Everything else follows from there:
- Your First Project — the whole loop end to end: a component, a system, and Play. Start here.
- Creating a Project — what the scaffold generates and why each file exists.
- The Editor — the panels, how Play works, and what is not built yet.
The rest of this page is the handful of things that are easy to trip over and do not belong to any one of those.
One engine per machine, shared by every project
The editor materialises the engine once per version in
~/.local/share/kooch/<version>/engine
and every project’s Cargo.toml points at it:
kooch = { path = "/home/you/.local/share/kooch/0.1.0/engine" }
Nothing is copied into the project. Two projects on the same engine version share one directory; two versions coexist, so a project pinned to an older engine keeps building after the editor updates.
⚠️ That path is absolute and $HOME differs per user, so a project that
changes machines names a directory that is not there. The editor owns
that line — it owns the directory it names — and rewrites it when a
project opens. Nothing to do by hand.
KOOCH_ENGINE_HOME overrides the base, for CI and for portable installs
that must not write to the user’s data directory.
It costs no build time. A project always compiled the engine from source; this only changes where the source is.
When that directory is replaced
Every materialised engine records which source tree it came from, in a
.kooch-engine-stamp beside it — the version, plus a digest of every
file. An editor compares its own source against that stamp, and says
so rather than acting on it:
Engine 0.1.0 — same version, different source than this editor ships
[ Install ] [ Keep ]
Install replaces the directory, leaving one copy behind, never two. Keep leaves it alone. Nothing is replaced under a project that was about to be built, which is what used to happen with only a log line to show for it.
🔴 A missing engine is installed without asking: there is nothing to keep, and a project that cannot build at all is not a choice worth offering.
🔴 What Keep cannot promise: engines are named major.minor.patch and
replaced in place, so keeping one holds until something else installs
over it — updating from another project, for instance. Two engines with
the same version have nowhere separate to live.
⚠️ Installing is refused while a build is running. Renaming the directory cargo is reading produces an error about a missing file in a crate nobody touched.
What is on this machine
Settings lists every installed engine, marks the one this editor ships and the one the open project uses, and removes the rest. Those two cannot be removed: both are named by a manifest, and deleting one leaves it pointing at nothing.
New versions are not created from there. The version is the engine’s
own major.minor.patch, so a new directory appears when an editor
shipping that version opens a project.
🔴 Without it, a new editor never updated the engine. The directory
is named after the engine version, that version used to be 0.1.0 for
every development build, and the old check only asked whether
Cargo.toml, crates and src existed — which is true of every copy of
the engine ever made. So a freshly installed editor found the directory,
called it current, and every project on the machine went on compiling
against weeks-old source with nothing said.
The version moves on every pull request now, automatically:
.github/workflows/version.yml bumps [workspace.package] in the PR’s
own branch — major for a ! or a BREAKING CHANGE:, minor for feat:,
patch for everything else. A number that sits still is what made all
three mechanisms that depend on it — this directory’s name, the
BuildStamp a project compares itself against, and the pipeline cache —
blind at once. Label a PR no-version-bump to opt out.
⚠️ One directory per version, and they are not cleaned up automatically. A version a week now means a directory a week; Settings is where the unused ones are removed.
⚠️ The build right after a replacement is a full rebuild, since every
engine source file is now newer than the project’s target/. That cost
is why installing is a question rather than something that happens while
you are opening a project to look at a scene.
A version this editor does not ship is never touched: that directory is what a pinned project builds against, and differing from the source in hand is the reason it exists rather than a reason to overwrite it.
Checking a copy that went wrong
The comparison above catches a stale engine, not a damaged one: deleting a file from a copy does not change what the copy claims to be.
KOOCH_VERIFY_ENGINE=1 kooch_editor
re-reads the whole tree, compares it against its own stamp, and re-copies when they differ. Off by default because it reads 8 MB every time a project opens.
⚠️ Rust is still required to build a project. Gameplay is native Rust compiled into the game, so the toolchain is not optional the way it is in an engine whose gameplay is a script.
Why the source is on disk at all
Because Rust has no stable ABI. A precompiled rlib links only against
the exact compiler and the exact dependency versions that built it, and
cargo does not model binary dependencies — which is why no Rust engine
ships binaries, Bevy included. The only route to “binary, no source” is
an extern "C" API in the shape of Godot’s GDExtension, and it costs the
typed ECS.
So the engine’s source is protected the way Unreal protects theirs: by licence, not by hiding it.
The licence is not optional
LICENSE.md is vendored with the engine, and the facade compiles it in:
#![allow(unused)]
fn main() {
pub const LICENSE: &str = include_str!("../LICENSE.md");
}
A game links the engine as an rlib, so that text is inside every
shipped executable. It is not a file someone has to remember to copy;
removing it means not building.
Packaging the editor
cargo build --release -p kooch_editor
cargo run --release --features editor --example package_editor -- dist/
dist/
kooch_editor the binary
engine/ 7.7 MB — the source it materialises for projects
.kooch-engine-stamp which tree this is, so an install can tell
whether it is newer than what is on the machine
assets/ what the editor itself renders with
engine_vendor::vendor_source looks in three places, in order:
KOOCH_ENGINE_SOURCE, engine/ next to the executable, and the engine
root — which only resolves when running from the engine’s own tree.
⚠️ package_editor refuses a binary older than the source. It once
shipped an editor built before this feature existed, and the AppImage
made from it wrote its own mount point into a project — a directory that
stops existing when the app closes.
⚠️ It packages for the platform it runs on. An editor for Windows
means running it on Windows, the same conclusion Bevy’s release workflow
reaches: metis is vendored C, which makes cross-compiling more than a
target flag.
Developing the engine itself
When the editor runs out of the engine’s own target/, project creation
points the manifest at the live clone and materialises nothing —
otherwise every engine change would need a re-materialise before the game
could see it. The check is where the executable is, not where the
source is.
Loading a scene
The boot scene is resolved in this order:
SceneBootstrapPlugin::with_scene(path), if yourmain.rssets one explicitly.--scene <path>on the command line — absolute, or relative to the working directory.scenes/default.scenebeside the executable.- The same, relative to the working directory.
So cargo run from the project root just works: the default path resolves because
the working directory is the project. A different level is
cargo run -- --scene scenes/Level1.scene.
🔴 Why the executable comes first
A shipped game is opened by double-clicking it, and that leaves the working directory
wherever the desktop felt like — your home, or /. Resolved against the cwd alone, a
released game starts with an empty scene and no error: the file was not missing from the
package, it was never looked for in the package.
So a packaged game keeps its content beside the binary:
dist/
mygame the executable
scenes/ default.scene, and the rest
assets/ everything the scenes reference, by GUID (.meta included)
The cwd stays as the fallback because that is what a plain cargo run inside a project
relies on — there the executable lives in target/debug/ and has no scenes/ beside it.
⚠️ The .meta sidecars are not optional. A scene references its assets by GUID, and the
GUID lives in the .meta next to each file. A copy that filters by extension and leaves them
behind produces a game that loads its scene and renders nothing.
Component registration runs before the scene loads
SceneBootstrapPlugin loads at Stage::First, which runs after every Stage::Startup
system has completed on the first frame. Component registration is a Startup system, so the
registry is fully populated by the time the scene is deserialised.
Flip that order and you get unknown component type: … and a scene that does not load. The
generated registrations.rs already puts registration in Startup; this matters only if you
register something by hand.
A scene with no camera renders black
The editor’s own camera is filtered out of a saved scene, so a scene needs to spawn its own
PerspectiveCamera. The default scene template includes one. Build a scene from scratch
without one and you get the clear-to-black fallback.
This is deliberate, not a bug. Injecting the editor camera as a temporary play camera — what Unity and Unreal do — is a possible future change, not current behaviour.
Running the game
cargo run
DefaultPlugins is the group that makes this a game rather than a collection of crates:
| Plugin | Role |
|---|---|
CorePlugin | Time, the AppExit event |
EcsPlugin | Storage, SceneManager, built-in components, transform propagation |
WindowPlugin | Winit window and GPU surface |
RenderPlugin | The mesh and sky pipelines |
SceneBootstrapPlugin | Loads the boot scene at startup |
Your project’s own main.rs is what runs, so any plugin you add there is picked up.
The Editor
The editor is where a project is authored: entities are spawned, components are attached and tuned, and the result is played back without leaving the window.
It is one program that runs in two arrangements, and the difference matters more than it looks.
Two arrangements, one editor
Embedded — the project’s own binary opens the editor:
cargo run # inside your project directory
The editor runs inside the project, so the project’s components are simply linked in. This is the simplest arrangement and needs no socket, no second process, and nothing to go wrong between them.
Standalone — the editor opens, and you point it at a project:
cargo run -p kooch_editor # from the engine repo
# then: Open Project
Here the editor is a separate program that does not have your project’s types compiled in. It gets them two ways at once:
- It loads the project’s
dylibto learn what components exist and what fields they have, which is what fills the Add Component menu and the Inspector. - It launches the project as a headless host (
--remote) and drives it over a local socket. That host is where the world actually lives, so a system you wrote runs in its process while you watch the result in the editor’s viewport.
Why the split exists. Rust has no stable ABI, so a pre-built editor cannot simply link a project compiled separately. The
dylibgets around that by requiring the same compiler for both — fine for code the editor itself built. The remote host covers the rest: running your systems against a live world. Seedocs/MEMORY.mdfor the full reasoning.
The Hub — the window you get from cargo run -p kooch_editor — is where projects are created
and opened.

The panels
| Panel | What it is for |
|---|---|
| View | The 3D viewport. Selection, transform gizmos, the physics debug overlay, and the meshlet debug-view dropdown. Owns the editor camera. |
| Game | What the game’s own camera sees — no gizmos, no selection outlines. A sibling tab of View, and only rendered while its tab is visible. Play does not switch you here or take the editor camera away; the two views coexist because a frame is a list of views. |
| World | The entity hierarchy of every loaded scene. Selecting here selects in View. |
| Inspector | The selected entity’s components and their fields. Where authoring happens. Hover a field name and its doc comment appears as a tooltip — units included, which is how you find out that a directional light’s intensity is in lux and a point light’s is in lumens. |
| Components | Every component type the engine and your project registered. |
| Archetypes | Which combinations of components actually exist, and how many entities are in each. A debugging view of how the ECS stored your scene. |
| Asset Browser | The project’s assets and the engine’s, as two roots. Right-click a .scene to make it the one the project opens with; that scene carries a ▶ in the tree and its name is in the accent colour. |
| Input Map | Edits a .inputaction asset: bindings, the five composites, processors. An action is an asset, not an entry in a map — see Writing a System. |
| Console | Structured logs from the editor and the launched project, filterable. Text is selectable and copyable. |
| Performance | Frame timings, and per-stage counters where they exist. |
Editing shortcuts
| Chord | What it does |
|---|---|
| Ctrl+Z / Ctrl+Y | Undo / redo the last edit. The Edit menu names the step it would take — Undo Duplicate Entity — so you can see what you are about to reverse. |
| Ctrl+D | Duplicate the selection where it stands. |
| Ctrl+C / Ctrl+V | Copy the selected entities into the editor’s clipboard and paste them as new ones, named Player Copy. What is held is the values, so a copy still pastes after the original is deleted. |
Ctrl+D, Ctrl+C and Ctrl+V act on the entity selection, so they are live only while the World panel or the View has focus. None of the four fire while you are typing in a field: Ctrl+C in the Console copies a log line, and in the Inspector it copies text. Each command is also in the Edit menu and in the World panel’s toolbar, with its chord written beside it — a greyed-out Paste means the clipboard is empty.
Undo follows the document, not the panel
Ctrl+Z undoes an edit to the thing you are looking at. The editor holds several documents open at once, and each keeps its own history:
| What you are editing | What Ctrl+Z reaches |
|---|---|
| The scene (World, View, or the Inspector on an entity) | the project’s world |
| A prefab open in the Inspector | that prefab’s document — one history per prefab |
| A material or an import setting | that asset |
| The input map in its panel | that map |
| The Console, the Asset Browser, the Build panel | nothing at all |
The Edit menu names both the step and the document — Undo Set intensity (this prefab) — so you can see which history you are about to move before you move it.
One stack for the whole editor is the Unity and Unreal model; it fits an editor that holds one
document open, and this one does not. A history per panel would be worse: the Inspector edits
whatever is selected, so its stack would hold edits to three different things. Godot 4 reaches
the same answer — histories keyed by scene, a global one for what belongs to no scene, and a
separate REMOTE_HISTORY for a live-edited remote world, which is exactly the split here.
A continuous edit is one step. Typing Player into a name field emits an edit per
keystroke and dragging a slider emits one per frame; edits to the same field collapse into a
single history entry, closed when you release the mouse or leave the field. Without that the
undo works perfectly and reads as broken — six Ctrl+Z to undo one rename.
Asset files are deliberately not undoable. Renaming, deleting and importing are not in any history: between the operation and the Ctrl+Z there is a filesystem watcher, an importer and whatever else is running, so an undo would be a promise the editor cannot keep. Unity leaves its Project window out of the undo stack for the same reason.
With a project open, undo travels to the project. The editor’s world is a mirror of one the project owns, so a Ctrl+Z is sent as the inverse of the edit that was made — the project applies it, and the mirror catches up on the next refresh. Undoing a despawn brings the whole subtree back with its values, under new entity handles: it is a rebuild, not a resurrection. Loading a scene or closing the project clears the history, because the world it describes is gone.
Play
Play does not rebuild anything and does not open a second window.
Pressing Play snapshots the authored world, flips the Playing gate so gameplay systems
start running, and simulates in the editor’s own viewport. Stop lowers the gate and restores
the snapshot, so the world goes back exactly as authored — you do not lose your scene by
testing it.
#![allow(unused)]
fn main() {
// kooch_remote::handlers::set_playing, in essence
if playing {
resources.insert(PlaySnapshot(WorldSnapshot::capture(resources)));
Playing::set(resources, true);
}
// Stop: the gate goes down *before* the restore, so no system
// observes a half-rebuilt world.
Playing::set(resources, false);
if let Some(snapshot) = resources.remove::<PlaySnapshot>() {
snapshot.0.restore(resources);
}
}
This is why your project’s systems register with run_systems: false while authoring: they
are registered either way and skipped per frame, so Play can flip them on live rather than
recompiling.
Known rough edge. A locally-opened project’s Play button still has an older path that shells out to
cargo run, which builds the project and opens a second window — minutes of nothing, and no snapshot. Tracked in #633. It does now run the game binary, which is what a player would get (#558).
Launch environment
Settings → Launch environment is a line of whitespace-separated
KEY=VALUE that the Play button hands the game, stored against the open
project’s path in the editor’s own config.
It exists because every knob this engine can be measured with is a
KOOCH_* variable — the frame they exist for is a game launched outside
the editor — and Play’s child process inherits the editor’s environment
and nothing else. Without the field, handing a game one variable meant
relaunching the editor with it set.
KOOCH_SHADING_PAD=4 KOOCH_FRAME_METRICS=log
🔴 Stored in editor_config.ron, not in project.kooch. A launch
option is a measurement, and a measurement committed to a repository is a
wrong configuration every collaborator inherits. Per project rather than
one global line, because “it silently applied to the other project too”
is how a capture ends up measuring something nobody asked for.
⚠️ Three variables are the editor’s and override anything typed here:
KOOCH_ENGINE_ROOT and KOOCH_PROJECT_ROOT, which the editor knows and
a text field does not, and KOOCH_LOG_FORMAT=json, without which the
Console cannot parse the game’s output at all — every line would arrive
as one opaque string that has lost the level and target it filters on.
RUST_LOG is the opposite: a default, so a line naming it wins.
No quoting. A value with a space in it would need a shell’s rules, and
these variables are single words; a token that is not a KEY=VALUE pair
is dropped with a warning rather than in silence, and the rest of the
line still applies.
What the editor does not do yet
Honest list, so nothing below is mistaken for a bug in your setup:
- No build button. Changing Rust code means running
cargo buildyourself (#158). - No reload. The project’s
dylibloads once, when the project opens. Seeing a code change means reopening the editor (#648). - No New Scene. Scenes have to exist on disk already (#619).
- Exposure and ambient light have no panel. Both are engine
Resourceswith sane defaults and no way to change them from the editor, so a scene that reads too bright or too flat cannot be corrected without editing code. See Lighting. - Reconnecting discards unsaved changes silently. Relaunching the host reloads the scene from disk; anything not saved is gone, and nothing warns first.
Textures
A texture reaches the GPU through three files: the image itself, the
.meta sidecar beside it, and the material that names its GUID. This
page is about the middle one.
Import settings
The sidecar carries the asset’s identity and how it is imported:
guid = "7b17f815-fcd2-4ede-867a-25d9ec31c792"
asset_type = "kooch_render::texture::asset::Image"
[import]
mipmaps = false
Select the texture in the Asset Browser and the Inspector edits the same
table — the checkbox writes this file. An absent [import] table means
the engine’s defaults, not “everything off”, so a texture imported
before the setting existed keeps behaving the way it did.
A malformed table logs a warning and falls back to the defaults. A settings file is never allowed to be the reason an asset does not appear.
mipmaps — default on
Pre-filtered half-size copies, all the way down to 1×1, sampled as a
surface tilts away from the camera. Without them a 1024-pixel grid on a
floor picks a different texel every frame and boils; the effect gets
worse in exact proportion to render_scale, because a smaller frame
covers the same texture with fewer samples.
Turn it off for textures read at their own scale, where the smaller copies are memory spent to make a 1:1 sample blurrier at glancing angles:
- a UI atlas drawn pixel-for-pixel,
- a lookup table whose neighbouring texels are unrelated values — an average of two of them is not a value the table has,
- a gradient ramp read by index.
Everything seen in perspective wants them on, which is why on is the default and off is the thing you say out loud.
🔴 The chain is built in linear light, not by averaging the bytes. Half black and half white is 0.5 of the light, which is 188 written back as sRGB — not 128. The engine gets this right by making the downsample a render pass, so the hardware’s sRGB decode and encode do the transfer function. It matters: averaging encoded bytes makes every distant surface darker than the one beside it, and the seam moves with the camera, so it reads as a lighting bug.
Changing the setting re-imports the texture immediately. That is not free plumbing: a mip chain is levels allocated when the texture is created and no API adds one afterwards, so the editor has to evict the uploaded copy and let the next frame put it back.
Tiling
How densely a texture sits on a surface is the material’s decision,
not the mesh’s: Tiling in the Inspector, uv_scale in the .ron. A
floor twenty units across wants 20, 20 from a grid whose square is
meant to read as one unit. Offset slides the texture; on a tiling
texture whole numbers change nothing, which is the point.
Scaling the mesh’s UVs is not the same thing — the mesh is shared, so it would change every object using it.
🔴 Tiling makes the uv move faster between neighbouring pixels, and the mip is selected from exactly that. The engine scales the derivatives with the coordinate; if it did not, a texture tiled 20× would sample about four levels too sharp and alias — the thing the chain exists to prevent, on the surfaces that asked for tiling.
Sharpness: bias and anisotropy
Two settings decide how sharp a texture reads, and they fix different problems.
Mip bias — automatic, and only with a temporal technique
A frame rendered at render_scale 50 % samples every texture for half
the pixels, so the detail the upscaler exists to reconstruct was never
rasterised. The engine compensates with the bias FSR documents:
mipBias = log2(render / display) - 1.0
The log2 term buys back the resolution the frame does not have; the
extra -1 is there because the jitter resolves sub-pixel detail —
with a history to accumulate into, a sharper level comes out correct
instead of shimmering.
🔴 Which is why it only applies when a temporal technique is on. With
upscale: None there is no history, and a sharper mip would be aliasing
on purpose. Nothing to configure: it follows the scale and the
technique.
Measured through the mip debug view: at half scale a surface samples one level sharper than native, because the reduced resolution costs a level on its own and the bias pays it back and spends one more.
Anisotropy — a setting, and the one that fixes floors
anisotropy in .rendersettings, 1 (off) to 16.
A surface at a grazing angle covers a footprint that is long and thin, and an ordinary filter has a single level for it: it takes the long axis, picks a level that would not alias there, and blurs the short axis by the same amount. That is why a tiled floor softens towards the horizon while a wall facing the camera stays sharp — and no amount of mip bias fixes it, because the level is right for one axis and wrong for the other.
Anisotropic filtering takes several samples along the long axis instead of one coarse one. On a grazing floor with a tiled checker, 16× kept 1.8× the detail of no anisotropy.
⚠️ It costs bandwidth, not arithmetic: more fetches on exactly the surfaces that already cover the most pixels. On a handheld measured as bandwidth-bound that is the expensive kind, which is why the default is off and the number is chosen by looking at a floor and at a capture.
The prototype textures
The engine ships Kenney’s Prototype Textures (CC0) under
assets/textures/prototype/, six colours of grid, checker and labelled
reference geometry, with a material beside each one.

⚠️ The index is not the same pattern across colours. dark_texture_01
and green_texture_01 are different images — the pack numbers each
colour independently. Pick from the sheet above, not from the number.
They are worth reaching for beyond blocking out a level: a scene of untextured surfaces has no high-frequency detail, and a mip chain, a LOD bias and a sharpening pass all act on exactly that. A renderer feature judged in a white room is a feature judged against nothing.
Shipping a Game
A build turns a project into a folder someone else can run: an executable, and the assets it needs, and nothing else. This page is what the Build panel does and how to make it produce something that starts on a machine that is not yours.
Presets
A .buildpreset is one way of building the project — “Linux release”,
“Windows”, “handheld”. It is an ordinary asset: it lives in assets/,
it is created from the Asset Browser’s context menu, and it is edited in
the Inspector like anything else. The Build panel only holds the
list, the button, and cargo’s output.
Presets belong in version control. They are configuration, and a project usually has more than one. The Build panel’s list is where you pick which one to build; no preset is “the” preset.
What a build produces
build/
linux/
My Game.x86_64 the executable, named for its platform
assets.kpack the scenes and everything they reference
project.kooch which scene the game opens with
windows/
My Game.exe
assets.kpack
project.kooch
libstdc++-6.dll ) mingw's C++ runtime — the game does not
libgcc_s_seh-1.dll ) start on Windows without these three
libwinpthread-1.dll )
Each platform gets a folder of its own under the preset’s output directory, so a preset that builds both does not have the second overwrite the first.
The extension follows the platform — .exe on Windows, .x86_64 on
Linux, the same convention Unity and Godot use. A folder holding both is
then unambiguous.
Which scene it opens with
Right-click a .scene in the Asset Browser → Set as Main Scene. That
scene is marked with a ▶ and an accent-coloured name from then on, and it is the one both
Play and a built game start from.
It is stored in project.kooch, which travels beside the executable —
that is the only reason a shipped game can know. Nothing else in the
package says which of five scenes is the first one.
🔴 This did not work before. main_scene existed in the manifest and
nothing read it: a game opened assets/scenes/default.scene whatever
the field said, so a project whose starting scene had any other name
shipped a build that started somewhere else — or started empty, with no
error anywhere.
A project with no project.kooch beside the binary still falls back to
assets/scenes/default.scene, which is what every build did until now.
What travels, and what does not
Only what the game reaches — and reaches is followed all the way down. Packaging starts from the project’s files, collects the GUIDs they reference, then follows those assets to the GUIDs they reference, and repeats until nothing new appears:
level.scene → floor.ron → grid.png
A scene names a material, the material names a texture, and all three travel. Each asset is collected once however many things point at it, and a cycle — two prefabs naming each other — terminates rather than hanging the build.
An asset nobody reaches stays behind, and so does every .buildpreset —
a build does not ship the instructions for making itself.
🔴 This used to stop at the first step: scenes and prefabs were read and nothing else was. The material shipped and its texture did not, and because a missing GUID is silent the game started and drew the 1×1 white fallback — a textured surface that looks like somebody authored it flat.
Assets only your code names
A GUID built in Rust is not something the packager can find by reading
files: a scene loaded by path, an asset chosen from a table, an id
assembled from a string. Declare those in project.kooch:
build: (
include: ["assets/meshes/suzanne.glb"],
),
Paths are resolved against the project first and the engine second. Each one is a root of the same walk, so declaring a material brings the textures it names — you never have to list what is underneath.
A declared path that matches no file is reported in the build log and does not stop the build.
Source does not travel either. A shipped game is a compiled binary; the
src/ folder is what produced it, not part of it.
This is why anything imported through the Asset Browser lands under
assets/ whatever folder was selected. The split is what makes “what
does the game need” answerable at all:
| Folder | Holds | Ships |
|---|---|---|
assets/ | scenes, meshes, textures, materials | the parts that are used |
src/ | components, systems, main.rs | compiled in, not copied |
.kooch/ | the pack key, local state | never |
The asset pack
With pack_assets on — the default — everything lands in a single
encrypted assets.kpack rather than as loose files. It is compressed
with zstd and sealed with AES-256-GCM, including the index, so the file
does not even reveal the names of what is inside it.
Scenes go in the pack too. A scene is the structure of the whole game; leaving it in plain text beside an encrypted pack would protect the textures and publish the design.
Turn pack_assets off while working out why a build behaves differently
from the editor — then the files are right there to look at.
⚠️ The key has to be inside the binary for the binary to read the pack. This raises the cost of taking your assets; it does not make it impossible, and nothing does. A game hands its meshes to the GPU in the clear because that is what drawing them means.
The key
Each project gets its own, generated once and kept at
.kooch/pack.key. It is not in version control — a repository carrying
it has published it, and history keeps it published after the file is
deleted. That is the same line Godot draws between export_presets.cfg
and its encryption key.
One key per project, so breaking one says nothing about the next.
For CI, set KOOCH_PACK_KEY to the key’s hex and nothing is written into
the checkout. Keep it in the secret store, not the repository.
The two modes
A preset’s mode is what the build is for. There are two, and both
are optimised.
| Release | Profiling | |
|---|---|---|
| Optimisations | full — LTO, one codegen unit | the same |
| Profiler | absent from the binary | compiled in |
| Open port | none | 0.0.0.0:8585 |
| Give it to people | yes | never (#558) |
Profiling streams every frame to the editor’s Profiler panel, CPU scopes and per-pass GPU timings alike — the only way to find out where a frame goes on the hardware the game has to run on. See Profiling.
🔴 Release is not “the profiler switched off”. With the feature absent, every scope in the engine expands to nothing at compile time and there is no socket to open. Nothing can be turned back on at runtime.
⚠️ There is deliberately no debug mode. A build compiled without optimisations runs several times slower — the editor’s own debug build measured 14.31 ms a frame against 4.94 ms for its release build — so profiling one tells you about that build and not about your game. The handheld’s entire budget is 13.9 ms.
Keep the two as separate presets — “handheld, profiled” beside “handheld” — so the ordinary build cannot acquire a socket because somebody left a dropdown on the wrong entry.
Running on another machine: min_glibc
A game built on an up-to-date desktop often refuses to start on a Steam Deck or a handheld, with a message about a missing symbol version. glibc is forward compatible and not backward: a binary linked against 2.43 does not run against 2.42, and the error says nothing about what to do.
Set min_glibc to the oldest version the build has to run on and it
links against that instead:
| Value | Runs on |
|---|---|
| empty | this machine and anything newer — the default |
2.28 | Debian 10, RHEL 8, and everything since — what Godot’s Linux exports target |
2.31 | Ubuntu 20.04 and newer |
Leave it empty while iterating locally; set it before handing the build to anyone.
It applies to the Linux half of a preset and no other: Windows has no glibc, and the floor is simply not carried there. One preset can set a floor and build both.
What it needs
Two tools, neither of which requires root — which matters on an immutable
distribution, where there is no dnf install to reach for:
cargo install cargo-zigbuild
# and zig itself, one tarball from https://ziglang.org/download/
# tar xf zig-*.tar.xz -C ~/.local/opt
# ln -s ~/.local/opt/zig-*/zig ~/.local/bin/zig
The build checks for both before compiling and names what is missing. A missing toolchain that surfaces ten minutes in, as a linker error, is the thing that check exists to prevent.
The field is ignored for targets that are not *-linux-gnu; there is no
glibc to have a floor.
Cancelling
The Build panel’s Cancel stops cargo. Nothing is packaged, so a half-built executable never reaches the output folder — packaging only runs when cargo exits clean.
Building for more than one platform
A preset has a checkbox per platform — Linux and Windows — and
ticking both builds both from one press, one after the other. They are
never built in parallel: two cargos on one machine fight over the same
target/ lock and interleave their output into a log nobody can read.
Each platform’s Rust target has to be installed. Every enabled
platform is checked before the first one starts compiling, with the
rustup target add line to run — so a missing Windows target is
reported up front rather than after Linux has spent ten minutes
building.
Windows also needs the mingw-w64 toolchain, and both halves of it:
metis is C and meshopt is C++. Having only the C compiler is a real
state to be in — Fedora ships mingw64-gcc and mingw64-gcc-c++ as
separate packages — and it fails inside a build script well after cargo
has accepted the target.
So both are checked up front, by the exact names cc-rs looks for, and
the refusal says what to install:
rpm-ostree install mingw64-gcc mingw64-gcc-c++ # Fedora, Bazzite
sudo apt install gcc-mingw-w64-x86-64 g++-mingw-w64-x86-64 # Debian, Ubuntu
sudo pacman -S mingw-w64-gcc # Arch
The C23 workaround metis requires is passed for you.
The three DLLs beside a Windows build
meshopt is C++, so the executable links mingw’s C++ standard library —
and on mingw that library is a DLL that ships with the compiler, not
with Windows. Leave it behind and the build runs on the machine that made
it and nowhere else, failing with a Windows dialog naming a file.
They travel automatically, stripped of their debug symbols on the way
(Fedora’s libstdc++-6.dll is 29.7 MB installed and 2.5 MB shipped), and
the build fails rather than producing a folder that looks complete
and holds a game that cannot start.
Do not delete them. They are three unexplained files beside a game, which is exactly the shape of a file somebody tidies away — and the game stops working the moment they go.
They are not linked statically because that does not work: mingw’s
libstdc++.a mixes static and dynamic symbols, so -static-libstdc++
and -static both leave the import in place. The problem is open
upstream in rust-lang/rust#65911.
A preset with no platform ticked builds nothing, and says so instead of guessing that you meant this machine.
Presets written before the checkboxes
They carried a target_triple instead. It is read once, on load, and
turned into the matching checkbox — an empty one meaning the machine the
editor is running on. A triple naming a platform with no checkbox opens
with none ticked and warns, rather than silently building something the
preset never asked for.
Your First Project
- 1. Open the Hub
- 2. Write a component
- 3. Write a system
- 4. Build
- 5. Use it
- What to read next
- If something did not work
A complete pass through the loop: create a project, write a component and a system, and watch them run. Roughly fifteen minutes, most of it the first compile.
1. Open the Hub
cargo run -p kooch_editor

Create a project, or open one you have. The first build of a new project compiles the engine too — several minutes, once.
A project made with an older editor is migrated on open. You do not have to do anything, but that first build will also be a full one.
2. Write a component
New Component from the editor, named Spinner, then open src/spinner.rs and fill it in:
#![allow(unused)]
fn main() {
use kooch::kooch_ecs::Reflect;
use kooch::kooch_ecs::component::Component;
/// Makes an entity rotate. Attach it and set the speed in the Inspector.
#[derive(Default, Reflect)]
#[reflect(category = "Gameplay")]
pub struct Spinner {
/// Degrees per second around the Y axis.
pub speed: f32,
}
impl Component for Spinner {}
}
speed is public, so the Inspector will draw a drag value for it. Nothing else is needed —
see Writing a Component for the attributes that change how it is drawn.
3. Write a system
This one touches Transform, whose fields are glam types. The prelude re-exports them, so
there is nothing to add to Cargo.toml:
Vec2,Vec3,Vec4,Quat,Mat3andMat4come throughkooch::prelude, and the wholeglamcrate is reachable askooch::glam. Adding your ownglamdependency is the one thing to avoid: aQuatfrom a different version is a different type, and the compiler error names two types spelled identically.
New System, named spin, then open src/spin.rs:
#![allow(unused)]
fn main() {
use kooch::kooch_ecs::Query;
use kooch::kooch_ecs::transform::Transform;
use kooch::prelude::*;
use crate::spinner::Spinner;
/// Rotates every entity that has a `Spinner`.
pub fn spin(resources: &mut Resources) {
let dt = resources
.get::<Time>()
.map(|t| t.delta_secs())
.unwrap_or(1.0 / 60.0);
let query = Query::<(&Spinner, &mut Transform)>::new(resources);
query.for_each(|(spinner, transform)| {
transform.rotation *= Quat::from_rotation_y(spinner.speed.to_radians() * dt);
});
}
}
The query matches only entities that have both components, so a Spinner on an entity
with no Transform is simply skipped rather than being an error.
4. Build
Registration already happened. The editor polls src/, found impl Component for Spinner and
pub fn spin(_: &mut Resources), and rewrote registrations.rs with both — within a second of
you saving, from whichever editor you saved in. The toolbar’s Resync button forces it, and
pulses when the generated file has moved ahead of your last build.
Then build. Today that means a terminal:
cargo build
and reopening the editor, because the project’s library is loaded once when the project opens. A build button and a live reload are #158 and #648; until they land, this step is manual and it is the slow part of the loop.
5. Use it
With the project reopened:
- Select an entity in World (or spawn one).
- Add Component → Gameplay →
Spinner. - Set
speedin the Inspector — try90. - Press Play.
It spins. Press Stop and the world returns exactly as you authored it — Play snapshots before it starts and restores on stop, so testing never costs you your scene.
What to read next
- Writing a Component — every field type the Inspector can draw, and the attributes that control it
- Writing a System — queries, stages, spawning
- Creating a Project — what each generated file is for
- The Editor — the panels, and what is not built yet
If something did not work
| Symptom | Cause |
|---|---|
| The component is not in the Add Component menu | The project has not been rebuilt and reopened since it was written — registering is not compiling |
| A field is not in the Inspector | It is private, has #[reflect(skip)], or is a type reflection does not support yet (#649) |
| The derive does not compile | A field’s type is not supported — Vec<T>, HashMap, your own enums. Mark it #[reflect(skip)] |
| The system never runs | It is in Update behind the Playing gate; press Play. Or its signature does not match pub fn f(_: &mut Resources) exactly — on one line — so the scanner missed it |
| The system runs in the wrong stage | Say which with #[system(PreUpdate)]; without it, every system lands in Update. See Writing a System |
| A child object or shadow lags one frame behind | The system writes a Transform in PostUpdate or later, after the engine already resolved GlobalTransform. Move it to Update |
| Play opens a second window and takes minutes | The old local-Play path (#633) |
Creating a Project
A project is an ordinary Cargo crate that depends on the engine. The editor scaffolds it, but nothing about it is magic — you can read every generated file, and most of them you will never touch.
What the editor generates
MyGame/
├── Cargo.toml
├── assets/
├── scenes/
└── src/
├── main.rs # generated, yours to edit
├── lib.rs # generated, editor-managed
└── registrations.rs # generated, editor-managed — do not edit
Cargo.toml
The one line worth understanding:
[lib]
crate-type = ["rlib", "dylib"]
Two artefacts from one crate. The dylib is what the standalone editor loads to learn
your component types without compiling them. The rlib beside it is what your binary links,
so the shipped game is an ordinary statically linked executable — no dynamic loading at
runtime.
The project declares its own features, and the default one is the game.
[features]
default = ["game"]
game = ["kooch/physics", "kooch/gravity", "kooch/camera", "kooch/audio"]
editor = [
"game",
"kooch/editor",
"kooch/remote",
"kooch/dynamic",
"kooch/physics-debug-render",
]
[dependencies]
kooch = { path = "…" }
Each one buys something specific. physics gives you rigid bodies — without it a
PhysicsBody is an inert component and nothing ever falls. gravity is the same story one
level up: a PointGravity that pulls on nothing. camera is the third: a VirtualCamera
that moves no camera.
The editor three are what authoring needs and a game does not: kooch/editor is the
embedded editor, kooch/remote is the socket the standalone editor drives your project
over, kooch/dynamic is the plugin API that lets it list your components without compiling
them, and physics-debug-render compiles the solver walk the physics overlay draws.
🔴 Why the game is the default, and why it matters
A shipped build must contain the game and nothing else. Bundling the editor ships the engine’s authoring surface next to your game — the tooling that authored it, in the same artefact — plus a file-dialog stack, an HTTP listener, and a reflected description of every type you registered.
editor is opt-in and the authoring binary asks for it with required-features, so the
guarantee belongs to the build, not to a cfg somebody has to get right. Check it
yourself:
cargo tree -e normal | grep kooch_editor_core # nothing
cargo tree -e normal --features editor | grep kooch_editor_core # there it is
Not “the linker drops it” — cargo never compiles it.
⚠️ Reflection stays in a game build, and that is not an oversight: a .scene is
deserialised by type name, so the game needs the registry to load its own scenes. What
leaves is the editor, the remote server and the plugin API.
main.rs — the game, and nothing else
use PROJECT_CRATE::registrations;
fn main() {
let mut app = App::new();
app.add_plugins(DefaultPlugins);
app.add_plugin(registrations::ProjectRegistrations { run_systems: true });
app.run();
}
No flags and no modes. Authoring lives in src/editor.rs, which is a second [[bin]]
gated behind the editor feature, so this file cannot reach it.
| Command | What you get |
|---|---|
cargo run | Your game |
cargo run --features editor --bin <crate>_editor | The editor, with your components |
cargo run --features editor --bin <crate>_editor -- --remote | A headless host for the standalone editor to drive |
--remote is headless on purpose: the editor draws that world in its own viewport, so a
window here would show the same scene twice.
The editor passes --features editor on every build it runs for you, so none of this is
something to remember while authoring — it matters the day you ship.
⚠️ Older projects are migrated when they open. One exception: a main.rs you edited is
left exactly as it is, with a warning, because a migration that silently deleted a line of
your gameplay would be worse than one that did nothing. Move your setup into the plain
App::new() form above and the release build is the game only.
registrations.rs — do not edit
The editor regenerates this file whenever you create or register a script. It scans src/
for two patterns and wires up what it finds:
impl Component for X→ a componentpub fn f(_: &mut Resources)→ a system
Detection is line-based rather than a full parse — enough for the generated templates and for typical hand-written code, but it does mean an unusual formatting of those signatures can go unnoticed. If a script you wrote does not show up, that is the first thing to check.
lib.rs — the editor’s entry point into your project
Also generated. It exports one plugin whose only job is to describe your components to a
standalone editor that loaded this dylib:
#![allow(unused)]
fn main() {
impl kooch::kooch_plugin_api::KoochPlugin for ProjectPlugin {
fn name(&self) -> &str { "my_game" }
fn build(&mut self, engine: &mut dyn kooch::kooch_plugin_api::Engine) {
registrations::declare_components(engine);
}
}
kooch::kooch_plugin_api::export_plugin!(ProjectPlugin);
}
Opening an older project
Projects made with earlier versions of the editor are migrated on open: the dylib crate
type, the dynamic feature, and the registrations wiring are all added if missing. You do
not have to do anything, but the first build afterwards will be a full one.
The compiler has to match
The dylib boundary carries Rust types directly rather than going through a C interface —
that is what makes the API pleasant. The price is that the project and the engine must be
built by the same rustc. A mismatch is refused with a clear message (the engine records
rustc -V -v at build time) rather than crashing, but it is refused.
In practice this is invisible, because the editor builds both. It becomes visible if you update your toolchain and rebuild only one side.
Writing a Component
- The smallest one that works
- What the Inspector can draw
- Attributes
- Pointing at another entity
- Registration
- What survives a save
A component is a plain struct that derives Reflect and implements Component. That is the
whole contract — the Inspector, scene serialisation and the Add Component menu all follow from
the derive.
The smallest one that works
#![allow(unused)]
fn main() {
use kooch::kooch_ecs::Reflect;
use kooch::kooch_ecs::component::Component;
/// How much damage this entity can still take.
#[derive(Default, Reflect)]
pub struct Health {
pub current: f32,
pub max: f32,
}
impl Component for Health {}
}
Create it from the editor, which drops this scaffold in src/ — or write the file yourself
in any editor. Either way the registrations regenerate on their own: the editor polls src/
and rewrites them within a second of a save.
⚠️ Regenerating is not rebuilding. The Inspector lists what your project’s last build
contained, so a component added a moment ago is in registrations.rs and in no binary yet.
The toolbar’s Resync button pulses while the two disagree.
Public fields show up in the Inspector automatically. No attribute is required to opt in; attributes exist to opt out, or to say something the type alone cannot.
What the Inspector can draw
Each field’s Rust type maps to a FieldKind, and the kind decides the widget:
| Rust type | Widget |
|---|---|
f32, f64 | Drag value |
u8…u64, i8…i64 | Drag value, clamped to the type |
bool | Checkbox |
String | Text field |
Vec2, Vec3, Vec4 | Component-wise drag values |
Quat | Euler angles, in degrees |
Mat4 | Decomposed to translation / rotation / lossy scale, read-only |
Option<Guid> + #[reflect(asset = "…")] | Typed asset picker |
Option<EntityRef> | Entity picker, and a drop target for a drag from the World panel |
Entity, Option<Entity> | Same widget, but see “Pointing at another entity” below |
A struct that also derives Reflect | Nested, drawn inline |
The maths types come from the prelude.
Vec3,QuatandMat4areglamtypes, andkooch::preludere-exports them so a project never declares its ownglamdependency — which is the point, since aQuatfrom a different version is a different type and the compiler error would name two types spelled identically.
Hover a field name in the Inspector to see its doc comment. The derive harvests
///straight off the field, so documenting a component for the next reader also documents it for whoever is authoring the scene — there is no second place to write it and no second place for it to go stale.
Anything outside that list — Vec<T>, HashMap<K, V>, your own enums — is not supported
yet. Recursive reflection for nested types and collections is
#649. Until it lands, a field of an
unsupported type needs #[reflect(skip)] or the derive will not compile.
Attributes
On the struct
#![allow(unused)]
fn main() {
#[derive(Default, Reflect)]
#[reflect(category = "Gameplay")] // groups it in the Add Component menu
#[reflect(inspector = "read_only")] // "hidden" | "read_only" | "editable" (default)
pub struct Health { … }
}
On a field
#![allow(unused)]
fn main() {
#[derive(Default, Reflect)]
pub struct Weapon {
/// Not shown, not serialised. Use for runtime caches and for types
/// the Inspector has no representation for.
#[reflect(skip)]
cached_target: Option<Entity>,
/// A typed asset picker instead of a raw Guid text field.
#[reflect(asset = "Mesh")]
pub projectile: Option<Guid>,
/// A dropdown of named values instead of a bare integer.
#[reflect(choices = FIRE_MODE_CHOICES)]
pub fire_mode: u32,
/// A row of checkboxes instead of a bitmask you compute in your head.
#[reflect(bits = DAMAGE_TYPE_BITS)]
pub damage_types: u32,
/// Only drawn when another field says it is relevant.
#[reflect(shown_when = BURST_ONLY)]
pub burst_count: u32,
/// A reference the picker will only let you point at an entity
/// carrying a `PhysicsBody`.
#[reflect(requires = "PhysicsBody")]
pub anchored_to: Option<EntityRef>,
}
}
choices, bits and shown_when take a path to a constant, not a string literal — so the
same table is used by the Inspector and by your code, and they cannot drift apart.
shown_when is what keeps a component with many mutually-exclusive fields readable: the
engine’s own Joint uses it so a hinge does not show you spring stiffness.
requires names a component, by its short name, that the target has to carry. The picker
filters by it and refuses a drop that fails it, saying why — a reference accepted but inert
is indistinguishable from a broken one.
Pointing at another entity
Use Option<EntityRef>.
#![allow(unused)]
fn main() {
use kooch_ecs::reflect::EntityRef;
#[derive(Default, Reflect)]
pub struct Turret {
pub target: Option<EntityRef>,
}
}
Three things assign it, and all three write the same value: your code
(turret.target = Some(EntityRef::live(entity))), the Inspector’s picker, and dragging an
entity from the World panel onto the field.
EntityRef is two states, because a reference means two different things depending on where
it lives:
Live— an index and a generation. What a running component holds, and whatEntityRef::entity()gives you back for a query or a lookup.Persistent— an identity that survives a reload. What a scene file holds.
You do not convert between them. Saving resolves live to persistent, loading resolves back,
and a reference whose target’s scene is not open stays persistent until it is — which is why
the field is Option<EntityRef> and not Option<Entity>. An Entity field has nowhere to
put an unresolved reference, so it loses the link instead of keeping it.
Entity and Option<Entity> still reflect, for a handle the engine resolves itself
(Parent is one). They refuse to store anything but a live reference.
Registration
You do not write it. The editor scans src/, finds impl Component for Health, and
regenerates registrations.rs with both halves:
#![allow(unused)]
fn main() {
// Registers the type with the running ECS — scene save/load and the Inspector.
registry.register_cpu_reflected::<Health>();
// Describes the type to a standalone editor that loaded this dylib.
declare_component::<Health>(engine);
}
The component’s name comes from std::any::type_name::<T>(), so there is exactly one name
for a type and no way for two sides to disagree about it.
What survives a save
A component is saved as its reflected fields. Two consequences worth knowing before you design a component:
#[reflect(skip)]fields are not saved. They are reconstructed by your code, or they are gone.- A reference to another entity is saved as an identity, not as a handle. The save path
resolves it and assigns the target a persistent id if it has none, which is why saving a
scene can modify the world. Nothing is asked of you beyond using
Option<EntityRef>; a handle reaching a file is refused by name, and the save fails rather than writing a reference that would load pointing at some other entity.
Writing a System
- The smallest one that works
- Reading and writing components
- Stages
- The
Playinggate - Spawning and despawning
- Registration
A system is a function. That is the entire type:
#![allow(unused)]
fn main() {
pub fn my_system(resources: &mut Resources) { … }
}
No trait to implement, no macro, no parameter-injection magic. Resources is the world, and a
system does whatever it wants with it.
The smallest one that works
#![allow(unused)]
fn main() {
use kooch::prelude::*;
/// Ticks every entity's regeneration.
pub fn regenerate_health(resources: &mut Resources) {
let _ = resources;
}
}
The editor’s New System command drops this scaffold in src/ and regenerates
registrations.rs, which picks it up by its signature.
Reading and writing components
Components are reached through Query, which is constructed from Resources and borrows what
it names:
#![allow(unused)]
fn main() {
use kooch::prelude::*;
use kooch::kooch_ecs::query::Query;
use crate::health::Health;
pub fn regenerate_health(resources: &mut Resources) {
let dt = resources
.get::<Time>()
.map(|t| t.delta_secs())
.unwrap_or(1.0 / 60.0);
let query = Query::<&mut Health>::new(resources);
query.for_each(|health| {
health.current = (health.current + 5.0 * dt).min(health.max);
});
}
}
A few shapes worth knowing:
#![allow(unused)]
fn main() {
Query::<&Health>::new(resources) // read one component
Query::<&mut Health>::new(resources) // write one component
Query::<(&Transform, &mut Health)>::new(res) // entities that have both
}
and on the query itself:
| Call | Use |
|---|---|
.iter() | An iterator, when you want to collect, sum, filter |
.for_each(|item| …) | The common case |
.for_each_entity(|entity, item| …) | When you need the entity id too, e.g. to look up an optional component |
.get(entity) | One specific entity, None if it does not match |
.is_empty() | Cheap early-out |
Conflicting borrows panic rather than corrupt. Holding &mut Health in two live queries
at once is caught by the access tracker. Scope a query with a block when you need to release
it before building the next one.
Stages
A system is registered into a stage, and stages run in a fixed order every frame:
| Stage | For |
|---|---|
Startup | Once, at launch. Load, allocate, seed. |
First | The very top of the frame. |
Input | Reading devices into intent. |
PreUpdate | Preparing what Update will need. |
Update | Your game logic — the default choice |
PostUpdate | After gameplay, and where Transform becomes GlobalTransform. |
GpuSync | Handing this frame’s data to the GPU. |
Gpu | Compute submitted with the frame’s encoder. |
Physics | Fixed timestep. May run several times a frame, or none. |
PostPhysics | Same timestep, after the solver. |
PreRender | Last chance before drawing. |
Render | Drawing. |
PostRender | After drawing. |
Last | The very end of the frame. |
If you do not have a reason, Update is the reason.
🔴 PostUpdate is the one that bites. It is where a local Transform is resolved into the
GlobalTransform that meshes, lights and cameras actually read. Write a transform before it
and the change lands this frame. Write it after — PostUpdate, Gpu, anywhere later — and
everything downstream renders one frame behind, forever, with no error and no log line. The
symptom is shadows or child objects that lag when the camera moves, which is not a bug anybody
traces back to a stage.
Physics runs on a fixed timestep, so a system in Physics or PostPhysics should use
Time::fixed_delta_secs() rather than delta_secs(). Using the wrong one is a bug that only
shows up when the frame rate changes. It may also run several times in one frame, or none at
all, so nothing that must happen once per frame belongs there.
The Playing gate
Your systems are registered whether or not the game is running, and skipped per frame while it is not. That is what lets the editor’s Play button start gameplay without a rebuild.
registrations.rs does it by wrapping each of your systems:
#![allow(unused)]
fn main() {
app.insert_resource(Playing(self.run_systems)); // false while authoring
app.add_system(Stage::Update, run_if_playing(my_system)); // skipped while the gate is down
}
What it means in practice: a system must not assume it runs every frame from startup. It may start running at any moment, against a world somebody has been editing by hand — and stop again when Stop restores the authored snapshot underneath it.
Spawning and despawning
Structural changes go through Commands, not through Query, because adding a component
moves an entity between archetypes and that cannot happen while a query is iterating it.
Commands::spawn needs &mut Resources itself, so it cannot be borrowed out of resources
while it is used. Take it out, use it, put it back:
#![allow(unused)]
fn main() {
use kooch::kooch_ecs::commands::Commands;
pub fn spawn_a_pickup(resources: &mut Resources) {
let Some(mut commands) = resources.remove::<Commands>() else { return };
let entity = commands
.spawn(resources)
.insert(Health { current: 100.0, max: 100.0 })
.id();
commands.apply(resources);
resources.insert(commands);
let _ = entity;
}
}
Two things that catch people:
- The id is allocated immediately; the components are not.
spawnreturns a validEntitystraight away, but the inserts are queued untilapply. Querying that entity beforeapplyfinds nothing on it. - Put
Commandsback.removetakes it out ofResources; anything running later that expects it there will not find it.
To despawn:
#![allow(unused)]
fn main() {
commands.entity(target).despawn();
}
Registration
You do not write it. The editor watches src/, finds
pub fn regenerate_health(_: &mut Resources), and regenerates registrations.rs:
#![allow(unused)]
fn main() {
app.add_system(Stage::Update, run_if_playing(health::regenerate_health));
}
It watches by polling, so saving from any editor is enough — including one that is not this one. There is no button to press.
Saying where it goes
Update behind the Playing gate is what a system gets when it says nothing. #[system(...)]
says otherwise:
#![allow(unused)]
fn main() {
#[system] // Update, gated by Play — the default
#[system(PreUpdate)] // a different stage, still gated
#[system(PostUpdate, always)] // and running while you author, too
}
always drops the Playing gate. Reach for it when the work has to happen in the editor as
well: a gizmo, an overlay, a streaming pump. It is a word rather than something inferred
because it is the one thing no amount of reading a function can tell you — a system that must
run while paused looks exactly like a gameplay one.
The attribute expands to nothing. It is read by the editor when it regenerates the file;
the compiler passes your function through untouched, and deleting the attribute never breaks a
build. What it does do is validate: a stage that is not one of the fourteen is a compile error
naming them, where a comment with the same typo would leave the system in Update forever and
never say so.
Registered is not compiled
⚠️ The generated file names your system within a second of you saving. The editor still runs the last build of your project. Those two disagree until you rebuild, and the toolbar’s Resync button pulses while they do.
This matters because the gap is invisible: the editor lists a project’s components and systems
out of its compiled dylib, so something added ten seconds ago exists in registrations.rs and
in no binary anywhere. The symptom is “I pressed Play and my system did not run”.
Crate Graph
Kóoch is a Cargo workspace of 20 crates: 19 under crates/ plus the
top-level kooch facade. The structure is intentionally fine-grained: each
subsystem lives in its own crate so that downstream crates only depend on
what they actually need. This keeps compile times low when iterating on a
single subsystem and makes the dependency surface auditable at a glance.
Layers at a glance
They sit in nine layers. Each layer may only depend on layers below it.
The layer of a crate is its longest path to a crate with no internal
dependencies. That is a property of Cargo.toml, not an opinion, so it is
read out of the workspace rather than assigned:
python3 .github/scripts/crate_layers.py # the two tables below
python3 .github/scripts/crate_layers.py --check # or just: did they drift?
🔴 This page had drifted, and the drift is the argument for that script.
It claimed eight layers over 18 crates, put kooch_input and kooch_camera
two and three layers below where they are, listed kooch_render in two
tables at once, and did not mention kooch_pack at all — while asserting
that the layers move “whether or not anyone updates this page”. Derived once
by hand is not derived.
| Layer | Crates | Role |
|---|---|---|
| L0 · foundation | kooch_plugin_api, kooch_ecs_macros, kooch_pack | No internal deps. Type vocabulary, proc-macros, the shipped-asset container. |
| L1 · core | kooch_core | App, Plugin, Schedule, Resources, GpuContext, the asset server. |
| L2 · primitives | kooch_ecs, kooch_window, kooch_audio | ECS, windowing, audio. Depend on kooch_core only. |
| L3 · domain | kooch_input, kooch_lighting, kooch_physics, kooch_remote, kooch_world | Built on the ECS. Input actions, lighting data, simulation, remote protocol, scene organisation. |
| L4 · built on domain | kooch_render, kooch_gravity | The renderer needs Inti’s shading model; gravity needs the solver. |
| L5 · built on the renderer | kooch_gizmos, kooch_camera | Gizmos submit geometry through the renderer; the camera asks the gravity field which way is up. |
| L6 · gizmo interaction | kooch_gizmos_handles | Draggable handles on top of gizmo drawing. |
| L7 · editor | kooch_editor_core | Editor logic as a library. Depends on 14 internal crates — the widest surface in the workspace. |
| L8 · binary + facade | kooch_editor, kooch | The editor main(), and the facade user projects depend on. |
Two of those placements are worth a sentence, because neither is where a reader would guess:
kooch_camerais L5, not L3, because aVirtualCameraaskskooch_gravity::gravity_upwhich way up is. On a planet, up is not+Y, and a camera rig that assumed it would roll over at the equator. Withoutkooch_gravitycompiled in, the mode falls back to world up.kooch_coredepends onkooch_pack, which is why the container crate is foundation rather than tooling: the asset server opens.kpackfiles, so the format has to sit under the crate that reads them.
Inter-layer flow
The arrow direction reads as “A depends on B”. Within a layer, crates are siblings.
flowchart TD
L8["L8 · binary + facade<br/>kooch_editor · kooch"]
L7["L7 · editor<br/>kooch_editor_core"]
L6["L6 · gizmo interaction<br/>kooch_gizmos_handles"]
L5["L5 · built on the renderer<br/>kooch_gizmos · kooch_camera"]
L4["L4 · built on domain<br/>kooch_render · kooch_gravity"]
L3["L3 · domain<br/>kooch_input · kooch_lighting · kooch_physics<br/>kooch_remote · kooch_world"]
L2["L2 · primitives<br/>kooch_ecs · kooch_window · kooch_audio"]
L1["L1 · core<br/>kooch_core"]
L0["L0 · foundation<br/>kooch_plugin_api · kooch_ecs_macros · kooch_pack"]
L8 --> L7
L7 --> L6
L7 --> L5
L7 --> L4
L7 --> L3
L7 --> L2
L6 --> L5
L5 --> L4
L5 --> L3
L4 --> L3
L3 --> L2
L2 --> L1
L1 --> L0
Detailed dependency table
Per-crate internal dependencies (external deps like wgpu, winit, etc.
are omitted).
| Crate | Depends on |
|---|---|
kooch_ecs_macros | — |
kooch_pack | — |
kooch_plugin_api | — |
kooch_core | kooch_pack, kooch_plugin_api |
kooch_audio | kooch_core |
kooch_window | kooch_core |
kooch_ecs | kooch_core, kooch_ecs_macros, kooch_plugin_api |
kooch_input | kooch_core, kooch_ecs |
kooch_lighting | kooch_core, kooch_ecs |
kooch_physics | kooch_core, kooch_ecs |
kooch_remote | kooch_core, kooch_ecs |
kooch_world | kooch_core, kooch_ecs |
kooch_gravity | kooch_core, kooch_ecs, kooch_physics |
kooch_render | kooch_core, kooch_ecs, kooch_lighting |
kooch_camera | kooch_core, kooch_ecs, kooch_gravity |
kooch_gizmos | kooch_core, kooch_ecs, kooch_render |
kooch_gizmos_handles | kooch_gizmos |
kooch_editor_core | kooch_camera, kooch_core, kooch_ecs, kooch_gizmos, kooch_gizmos_handles, kooch_gravity, kooch_input, kooch_lighting, kooch_pack, kooch_physics, kooch_remote, kooch_render, kooch_window, kooch_world |
kooch_editor | kooch_core, kooch_ecs, kooch_editor_core, kooch_render, kooch_window, kooch_world |
kooch | kooch_core, kooch_ecs (always); kooch_audio, kooch_camera, kooch_editor_core, kooch_gizmos, kooch_gravity, kooch_input, kooch_lighting, kooch_physics, kooch_plugin_api, kooch_remote, kooch_render, kooch_window, kooch_world (optional, feature-gated) |
⚠️ examples/example_plugin is a workspace member and is not one of the
20. It is a member so that cargo test builds it — a plugin that stops
compiling is a broken ABI, and finding that out from a user is late.
Crate roles
Foundation (L0)
Crates with no internal dependencies. They can be built in isolation and form the type vocabulary the rest of the engine uses.
| Crate | Role |
|---|---|
kooch_plugin_api | Stable ABI types for dynamic plugins (loaded via libloading), including the KoochPlugin trait a project implements. Lives below kooch_core so plugins compiled against an old engine can still be probed. |
kooch_ecs_macros | Procedural macros for the ECS: #[derive(Reflect)], #[derive(Component)]. Standalone proc-macro crate. |
kooch_pack | .kpack — the zstd + AES-256-GCM container a shipped game reads its assets out of. Foundation rather than tooling because kooch_core’s asset server opens them, so the format sits under the crate that reads it. ⚠️ The key ships inside the binary, so it is a deterrent and not protection — Godot’s own docs say the same about theirs. |
Core (L1)
| Crate | Role |
|---|---|
kooch_core | App, Plugin, PluginGroup, Stage, Schedule, Resources, Time, GpuContext, event system, asset server, pipeline cache, power profile detection, and scene_paths — the file names three crates have to agree on. The minimum any Kóoch binary needs. |
Primitives (L2)
Small purpose-built crates that depend on kooch_core and nothing else.
| Crate | Role |
|---|---|
kooch_ecs | The ECS itself: archetype storage, Entity/Component traits, Query, Reflect, SceneDocument/SceneManager, hierarchy, transforms, built-in components (Transform, Name, PerspectiveCamera, OrthographicCamera, Mesh, lights, sky, etc). |
kooch_window | Winit integration, WindowPlugin, surface configuration, raw event dispatch. |
kooch_audio | Kira-based audio playback. Sits here rather than beside the ECS crates because it does not depend on the ECS: there is still no AudioSource component (#63), so there is nothing to author. |
Domain (L3)
Crates built directly on the ECS.
| Crate | Role |
|---|---|
kooch_input | Gamepad / keyboard / mouse abstraction. On the ECS since an action became an asset (#55): a component points at an .inputaction by guid, so nothing in gameplay names an action. |
kooch_lighting | Inti — the shading model, the GPU light record, extraction, exposure and ambient. The light components live in kooch_ecs beside every other component; what lives here is everything that turns them into pixels. See Lighting. |
kooch_physics | Physics simulation. Rapier is the backend, behind kooch_physics’s own types. |
kooch_remote | The local-socket protocol that lets the standalone editor drive a running project’s ECS. |
kooch_world | Scene/world organisation, chunk streaming and activation. |
Built on domain (L4)
| Crate | Role |
|---|---|
kooch_render | The GPU work: meshlet pipeline (cull, visibility buffer, shading, Hi-Z), shadows, the froxel grid’s consumer side, RenderPlugin, materials, glTF loading. Above kooch_lighting because a renderer without a shading model paints normals — which is literally what this one did until #441. See Render Pipeline. |
kooch_gravity | Multi-gravity system (Mario Galaxy-style fields). Needs kooch_physics to apply forces. |
Built on the renderer (L5)
| Crate | Role |
|---|---|
kooch_gizmos | Immediate-mode gizmo drawing. Needs kooch_render to submit geometry. |
kooch_camera | Camera components and VirtualCamera (follow / look-at with damping). Needs kooch_gravity, not the renderer: a rig asks gravity_up which way up is, because on a planet up is not +Y and a camera that assumed it would roll over at the equator. |
Gizmo interaction (L6)
| Crate | Role |
|---|---|
kooch_gizmos_handles | Draggable translate / rotate / scale / plane handles, with snapping and Local/World modes. Split from kooch_gizmos because drawing a gizmo and interacting with one are different problems. |
Editor (L7)
| Crate | Role |
|---|---|
kooch_editor_core | All editor logic as a library: panels (hierarchy, inspector, viewport, console, asset browser, profiler), undo/redo, project state and manifest, scene and prefab save/load, play/stop, the launch screen, and the build presets that produce a .kpack. Used by the editor binary AND callable as a plugin from custom hosts. |
Binary and facade (L8)
| Crate | Role |
|---|---|
kooch_editor | The editor main(). Imports kooch_editor_core and runs an App with the editor plugins wired. |
kooch | Top-level facade that re-exports the others under one name and defines DefaultPlugins (Bevy-style PluginGroup). User project crates depend on kooch rather than picking subcrates directly. Cargo features (window, render, audio, editor, dynamic…) gate which sub-crates pull in. |
Layering rules
Important: Lower layers must not depend on higher layers. If you find yourself wanting
kooch_coreto know aboutkooch_render, you have an inversion. Common fixes: introduce a trait at the lower layer and implement it at the higher layer, or pass behavior in via a generic / closure.
The dependency graph is acyclic by construction — Cargo refuses a cycle, so
that half needs no enforcing. What is only enforced by review is whether a
new edge is the right edge: kooch_camera → kooch_gravity is a real
dependency and moved the crate two layers, and nothing objected.
crate_layers.py --check catches the documentation drifting from the graph.
It does not catch the graph drifting from the design.
Why so many crates?
Three reasons:
-
Compile-time isolation. Iterating on
kooch_renderdoes not recompilekooch_ecsorkooch_window. Cargo’s incremental compilation benefits from real crate boundaries far more than from module boundaries inside one giant crate. -
Feature gating. The facade can compose user-facing builds (a headless server build skips
kooch_renderandkooch_window; an editor build pulls inkooch_editor_core). Single-crate builds cannot do this ergonomically. -
Reasoning surface. Knowing that
kooch_worldcannot accidentally reach intokooch_renderbecause Cargo enforces it makes refactors safer than relying on lint rules.
Tradeoff: more Cargo.toml files to maintain, more pub use re-exports
when types need to cross crate boundaries, slightly higher first-build time.
For an engine in early development the compile-time win outweighs the
ergonomic cost.
Adding a new crate
- Create
crates/kooch_yourthing/withCargo.tomlextendingworkspace.packagefields andCargo.toml[workspace.dependencies]for external deps. - Add it to the
membersarray in the rootCargo.toml. - Add it under
[workspace.dependencies]so other crates can depend on it viaworkspace = true. - If it’s user-facing, re-export from
kooch::*and add a feature flag to the top-levelCargo.toml[features]section. - Document its role in this page, then run
python3 .github/scripts/crate_layers.py --check— a new crate usually moves somebody else’s layer, and that is the part nobody remembers.
Render Pipeline
Kóoch renders through a GPU-driven meshlet pipeline, Nanite-style: the CPU uploads a flat array of instances and dispatches, and every decision about what to draw — frustum, backface, occlusion, level of detail — is taken on the GPU by a compute shader reading that array.
The CPU never walks a scene graph deciding what is visible. That is the whole point, and it is what “GPU-driven” means here.
This page describes what the code does today. Where something is missing the page says so and links the issue.
A frame is a list of views
MeshletRenderStage owns one geometry pool and a SlotMap<ViewId, MeshletView>. Each view has its own render targets, its own cull state
and its own camera; the pool, the instance buffer and the pipelines are
shared.
That split is deliberate and was not always true. Cull state is per view by definition — what survives a frustum test depends on where the camera is — and sharing it across views produces an over-cull that only appears once a second view exists, or once shadow cascades do, where it reads as “the shadows are wrong” rather than as a shared-state bug.
Two views run today: the editor’s View panel and its Game panel.
Shadow cascades did not become views, which was the plan when this page was written. They record inside the stage instead, against the unjittered camera and with their own bounded projection, because a cascade shares the pool and the instance buffer but wants none of a view’s render targets. Virtual-shadow-map pages (#477) may still be the case that makes a view the right shape.
Each view records and submits its own command encoder. Several per-frame buffers are shared across views on exactly that basis: a write followed by a submit is ordered on the queue, so view B’s camera cannot reach view A’s pass.
The frame, pass by pass
Which path runs depends on one capability: 64-bit texture atomics. The
device either has TEXTURE_INT64_ATOMIC + SHADER_INT64 +
SHADER_INT64_ATOMIC_MIN_MAX or it does not.
The node labels below are the GPU scope names a capture prints, so a flamegraph and this diagram can be read side by side. Anything not named here does not have a timer on it.
flowchart TD
START([Frame begins]) --> EXTRACT[CPU: walk the ECS<br/>MeshRenderer + GlobalTransform → instances<br/>lights → Inti's GPU buffer]
EXTRACT --> UPLOAD[Upload instances, grow buffers to fit]
UPLOAD --> SHADOWS["shadows<br/>4 cascade culls + rasters,<br/>plus a cube face per shadowed point light"]
SHADOWS --> GRID["cluster grid<br/>the froxel light index — 4 passes, two of them draws"]
GRID --> R64{64-bit texture<br/>atomics?}
R64 -- yes --> A0["cull: one thread per instance-meshlet<br/>frustum · backface cone · LOD chain descent"]
subgraph FUSED["raster + shade — one fused scope, timed as a whole"]
direction TB
A1[Clear the R64 visibility buffer] --> A2["Raster: draw_indirect the survivors<br/>fragment does atomicMax(depth << 32 | ids)"]
A2 --> MV["motion vectors<br/>previous clip position, unjittered camera"]
MV --> SH{"compute<br/>shading?"}
SH -- yes --> A4["shade: compute — or (half rate)<br/>one dispatch → Inti, into an HDR target"]
SH -- "no, the default" --> A5["shade: fragment<br/>one fullscreen pass per material, depth-tested Equal"]
A4 --> UP["shade: upsample<br/>only when the rate is half"]
UP --> TAA["taa / sgsr2<br/>the temporal resolve — off by default, #481"]
TAA --> TM["tonemap<br/>HDR radiance → display-referred"]
TM --> RCAS["rcas<br/>sharpening, after the curve — off by default"]
end
A0 --> A1
R64 -- no --> B0["cull + raster A against last frame's Hi-Z"]
B0 --> B2["hi-z build: SPD pyramid"]
B2 --> B3["cull + raster B: what pass A occluded"]
B3 --> B5["shade: one compute dispatch → Inti<br/>no motion vectors, no TAA on this path"]
RCAS --> SKY["sky"]
A5 --> SKY
B5 --> SKY
SKY --> BLIT["blit the stage's colour over the sky"]
BLIT --> PRESENT([Present])
style A2 fill:#1e5f3a,stroke:#4dbe8f,color:#fff
style A4 fill:#5f3a1e,stroke:#be8f4d,color:#fff
style A5 fill:#5f3a1e,stroke:#be8f4d,color:#fff
style B5 fill:#5f3a1e,stroke:#be8f4d,color:#fff
style SHADOWS fill:#3a1e5f,stroke:#8f4dbe,color:#fff
style GRID fill:#3a1e5f,stroke:#8f4dbe,color:#fff
style EXTRACT fill:#1e3a5f,stroke:#4d8fbe,color:#fff
Two of those run before anything is drawn, and the order is not arbitrary: shading samples the shadow atlas and reads the froxel grid, so both have to be filled first. They are separately scoped because a shadow pass that costs four culls and four rasters was, until #785, hiding inside whatever number the frame reported.
The R64 path’s raster + shade is one fused scope covering
everything from the clear to the sharpening. Its children — motion
vectors, the shade dispatch, the upsample, the temporal resolve, the
tonemap, RCAS — are timed individually inside it, which is how a capture
answers which half of the fused pass is the cost.
Cull
One compute thread per (instance × meshlet). Each thread tests its own
meshlet and, if it survives, appends its (instance_id, meshlet_id) to a
visible_meshlets buffer with an atomic bump. The draw that follows is
draw_indirect off a count the GPU wrote — the CPU never learns how many
meshlets survived, and does not need to.
Tests, in order: frustum against the meshlet’s AABB, backface via its normal cone, and LOD chain descent — a meshlet is drawn when its own screen-projected error falls under the target and its parent’s does not.
🔴 The LOD selector read the projection scale from a single matrix element for a long time. That element is
f × (camera up · world up), so it is correct for a level camera, smaller for a tilted one, and zero at 90° of roll or looking straight down — which switched the selector off entirely. It now takes the norm of the row that producesclip.y. Any non-level view had been losing detail since continuous LOD shipped, degrading smoothly enough to read as “that is how the model looks”.
Visibility buffer
Instead of shading during rasterisation, the raster pass writes only which triangle covered this pixel. Shading happens afterwards, once per pixel, for the triangle that won.
R64 path. The fragment shader does one
textureAtomicMax((depth << 32) | ids) into an R64Uint storage
texture. Depth in the high bits means the atomic max resolves depth and
identity in a single operation — no depth buffer, no z-fighting between
coplanar meshlets, no ordering.
R32 path. Without 64-bit atomics the same idea runs in two passes
against a Hi-Z pyramid built with single-pass-downsample: pass A draws
what was visible last frame, the pyramid is rebuilt from that depth, and
pass B recovers whatever pass A wrongly occluded. Metal has no
atomic_uint64, so this path is not legacy — it is the Apple path.
Shading
Both paths reconstruct the surface the same way, through
surface_reconstruct.wgsl: perspective-correct barycentrics from the
triangle’s three world-space positions, giving world position, normal,
uv, tangent and analytical uv derivatives — the automatic ones are
wrong here, because neighbouring fragments in a 2×2 quad may come from
different triangles.
Only the visibility-buffer read differs between the paths. That was not true until #441: the R32 path averaged the triangle’s three vertex normals and never computed a world position at all, which was invisible while shading was a function of the normal alone and would have lit the centroid of every triangle the moment a point light needed a distance.
- R64, fragment — the default. One fullscreen pass per material,
depth-testing
Equalagainst a target holding each pixel’s material id. The depth test is the per-material cull, in hardware, with early-Z. Each pass binds its own textures. - R64, compute — opt-in. One dispatch for the whole screen, into an
HDR target. Turned on per project (
compute_shading) or per run (KOOCH_COMPUTE_SHADING), and it is what half-rate shading and the reduced-rate upsample require. - R32 shades with one compute dispatch. No texture sampling: a
compute shader has no implicit derivatives, and
textureSampleGradis a fragment-stage call. Scalars only.
🔴 “Per material” means per material in the PROJECT, not in the frame.
MaterialPipeline::shading_slotsis0..next_slotandsync_from_resourcesregisters everyMaterialtheAssetDatabaseknows about, so dropping an unused.roninto the project’s folder adds a full-screen sweep to every frame. A tile that owns none of a slot’s pixels does no reconstruction and writes nothing, but it does not leave for free: every thread still reads the R64 vbuf, chasesvisible_meshletsand theninstancesoff that read, and waits on three unconditional barriers.
KOOCH_SHADING_PAD=<n>appendsnsweeps whosematerial_idmatches no instance. The frame is bit-identical — every store inmaterial_pbr_compute.wgslis inside the branch that never fires — so the only thing an A/B across it measures is what an idle sweep costs.Measured on the OneXFly at 1920x1080, 2026-08-20: 178 µs a sweep (1.98 µs on a desktop 9070 XT at 1280x720). A sweep is a fixed dispatch cost plus per-pixel work, so quote the resolution with the number.
roll-a-ballhas three materials and pays 0.71 ms a frame — 22 % of its own shading pass. A game with twenty pays 3.7 ms. Use a pad in the hundreds when measuring: four extra sweeps sit under the device’s own run-to-run drift.
🔴 A
serdedefault is not a recommendation — it is what an old file silently becomes.compute_shading,shading_rateandtemporal_aaall default to what the engine already did — fragment path, full rate, no history — because an earlier version defaulted two of them to on and every existing project changed shading path and gained a temporal resolve in the same build. Two variables at once is not a change anybody can bisect, and the first report was “you broke the whole render”.
The window, and everything being live
window_mode sits beside vsync in the Presentation group: 0
windowed, 1 borderless, 2 fullscreen, 3 exclusive.
| mode | what it is | changes the output resolution? |
|---|---|---|
| Windowed | a decorated window at the project’s size | no |
| Borderless | the same size, no title bar — still a window | no |
| Fullscreen | the monitor, at the monitor’s current mode | no |
| Exclusive | asks the display to change mode | yes |
🔴 Exclusive does not work everywhere, and the engine says so rather
than letting it fail quietly. Windows and X11 implement it; winit
ignores it on Wayland and its own source says so twice —
warn!("Fullscreen::Exclusive is ignored on Wayland"), which leaves
the window exactly as it was and reads as the setting being broken. So
window_mode::effective degrades the request to fullscreen before it
reaches winit, with a warning naming both modes.
The resolution, and what the monitor reports
Two resources carry it, both live:
Resolution { width, height, refresh_mhz }— what the game asks for. In windowed and borderless it is the window’s inner size, applied throughrequest_inner_size; in exclusive it is the display mode to switch to.refresh_mhz: 0means “the best this size can do”.DisplayModes { modes, exclusive }— what the platform will actually do, published once the window exists. The list comes from the player’s monitor rather than from a constant, andexclusiveis false under Wayland. A game’s options menu is built from this: a resolution dropdown that changes nothing is worse than none.
🔴 In exclusive, the size has to match a mode exactly. Substituting a nearby resolution would change what the player sees without saying so, so the fallback is borderless fullscreen at the monitor’s own size, with a warning. Among modes of the right size the engine takes the one closest to the refresh asked for, or the highest when none was.
⚠️ request_inner_size is a request. Wayland answers with a
configure event rather than a return value, which the engine already
handles: WindowResized → GpuContext::resize → the render targets.
⚠️ The environment override KOOCH_WINDOW_MODE=windowed|borderless| fullscreen|exclusive is applied when the window is created; the
asset’s value lands a few frames later, because the settings asset
needs the asset server, which needs the GPU, which needs the window.
⚠️ None of this reaches a handheld under gamescope. The compositor hands the game a surface at the resolution it was configured for and scales the result; the knob that decides there is outside the process.
Every setting in the file is live, which is what makes a game’s own options menu possible:
| setting | how it lands |
|---|---|
| exposure, ambient, shadows, contact shadows | resources read per frame |
compute_shading, shading_rate, upscale, sharpening | applied per frame on the stage |
render_scale | next frame — render_frame_system calls resize once a frame and it early-returns unless the size or the scale moved |
anisotropy | rebuilds one sampler when the number changes |
vsync | reconfigures the surface when the mode differs |
window_mode | set_fullscreen / set_decorations when the window is not already like that |
Each of those is guarded on “is it already what was asked for”, because every one of them is a reallocation of something — a swapchain, a sampler, a set of render targets — and applying it unconditionally would rebuild it once a frame.
🟢 What a handheld ships with
Measured on the OneXFly at 10 W, settled, 1920x1080, 2026-08-20: 13.92 ms median, GPU 9.7 ms, against a 13.9 ms budget.
compute_shading: true, // the tiled path; half rate needs it
shading_rate: 2, // Half — one sample per 2x2 quad
upscale: 2, // SGSR 2
render_scale: 50, // Performance — 50 % (2x)
🔴 Cap the frame rate, and not only for the battery. At 1280x720 the same frame costs 3.9 ms of GPU capped at 72 fps and 13.2 ms uncapped on this part. Capped, the GPU is idle 68 % of the time, so it never reaches its power cap and holds ~1210 MHz; uncapped it throttles and every pass takes three times longer. Rendering 144 frames to display 72 pays for the same work three times — and the cap fixes the pacing too: max frame 15.25 ms against 47.09.
⚠️ The capped run was also at a higher TDP than the uncapped one, so the
3.4× is the cap and the power together. gpu_busy_percent reading 32 %
argues the cap is most of it — a part idle two thirds of the time is not
power-limited — but the run that separates them has not been taken.
🔴 The upscaler is the largest single choice on that list. Same
scene, same build, same session: upscale: 3 (FSR 3.1) costs 11.355
ms and upscale: 2 (SGSR 2) costs 2.062 — a 23.36 ms frame against
a 13.92 ms one. FSR 3.1 is not broken; it is a desktop technique, and its
own dropdown entry says so.
DLSS (#536)
upscale: 4 is NVIDIA’s, and it is the only technique here a build can
lack. The other three are shaders this engine owns; DLSS is a neural
network shipped as a binary blob, reached by linking NVIDIA’s SDK.
Three things follow, and all three are visible from a project:
-
🔴 It is a compile-time feature.
dlss_wgpu’s build script linkslibnvsdk_ngxstatically and runs bindgen over NVIDIA’s headers, so a binary either linked it or did not. Turn it on by addingdlssto a build preset’s Extra cargo features; the editor suppliesDLSS_SDKandVULKAN_SDKand refuses to start cargo when either the SDK or the Vulkan headers are missing. The headers are a separate package from the loader — a machine that runs Vulkan games has no reason to carry them — so the editor’s startup check lists them alongside Rust and ALSA, in the same one-paste command.⚠️ What cargo actually receives is
kooch/dlss—dlsson its own names a feature of your crate, which you never declared, and cargo answers that with “the package does not contain this feature”. The editor rewrites the bare spelling for you, unless your ownCargo.tomldeclares adlssfeature, in which case it means yours and is left alone. -
🔴 It moves the whole build to Vulkan.
dlss_wgpuis Vulkan-only, and on Windows wgpu picks D3D12 by default. Enabling DLSS therefore moves every Windows player onto Vulkan, not only the ones with an NVIDIA card. That is a decision about the whole build, which is why it is a feature rather than a setting. -
🟢 Asking for it is always safe. A build without the feature, or a machine without an NVIDIA card, resolves with the engine’s own TAA at full resolution and says so once in the log. The
.rendersettingsfile is left alone:upscale: 4is what the project wants, and the machine that can honour it should still see it.
Getting the SDK. Settings → the DLSS button clones
NVIDIA/DLSS at the pinned tag after you
accept NVIDIA’s terms, into ~/.local/share/kooch/sdk/dlss/<version>/.
The engine never mirrors it — hosting a copy is the “stand-alone
product” the licence forbids.
Shipping it. A build with the feature gains two files beside the executable, copied by the packager rather than by you:
| file | why |
|---|---|
libnvidia-ngx-dlss.so.<ver> / nvngx_dlss.dll | NGX dlopens it from the application’s own directory. Nothing links it, so there is no rpath to get right |
DLSS_NOTICES.pdf | Section 9.5 of NVIDIA’s Programming Guide, which anyone distributing the blob must include. Copied because a licence file nobody remembered is a licence breach |
⚠️ Unmeasured. DLSS has a number on no device in this repository yet. It is a desktop option with an NVIDIA card; the handheld’s default stays SGSR 2 at 2.062 ms, and a vendor backend that does not beat ours by a number does not get to be a default.
Dropping the output resolution buys more than any of these, because
render_scale is a percentage of the output: a smaller window shrinks
the render target with it and everything render_scale does not touch —
the resolve’s output, the tonemap, the blit.
⚠️ These are a recommendation, not defaults. RenderSettings::default()
stays what the engine did before any of them existed, for the reason the
callout above gives: a serde default is what an old file silently
becomes.
Then Inti — Cook-Torrance driven by the scene’s lights.
After the shade: rate, history, and the tonemap
Three passes sit between Inti and the sky, and all three exist on the R64 path only.
- Half-rate shading.
KOOCH_SHADING_RATE=halfshades a quarter of the pixels andshade: upsampleputs them back on screen. The scope renames itself —shade: compute (half rate)— so a capture answers which rate produced this without anyone trusting a log line. The cheap half of the frame stays full-rate: the visibility buffer, the depth, the motion vectors. motion vectors. Each pixel’s previous clip position, from the camera’s unjittered matrix. The jittered one goes to everything else — cull, Hi-Z, raster, every reconstruction that reads the buffer the raster wrote — and the pair being separable is the whole reason the vectors are not wrong by a sub-pixel offset every frame.taa, and it is off by default. The resolve exists and works; turning it on is #481’s remaining half. Debug views bypass it, because averaging a false-colour legend across frames is not a legend any more.rcas, and it is off by default. Robust Contrast Adaptive Sharpening, one full-screen pass,sharpeningin.rendersettings. 🔴 It runs after the tonemap, unlike everything else in this list: RCAS is adaptive because it solves for the filter weight at which the signal would clip out of{0, 1}, and handed radiance in the hundreds that limiter stops limiting. When it runs, the tonemap resolves into its texture instead of into the window. Reconstruction is soft by construction — a resolve builds each output pixel from samples that landed near it — so at arender_scalebelow 100 this is not polish.tonemap. Shading writes HDR radiance into a linear target and the tonemap converts it at the end, because TAA has to run on linear radiance. The operator is concatenated from Inti rather than reimplemented — two copies of a curve that must agree to within one 255th is how a parity test starts failing for a reason nobody can find.
🔴 Everything that compresses range — the resolve, the tonemap, any firefly clamp — needs the exposure applied first. Radiance in this engine is in the hundreds, and
c / (max(c) + 1)on those numbers posterises into flat bands that read as a broken toon shader rather than as a missing divide.
Sky and composite
The sky is a fullscreen pass: procedural gradient plus volumetric clouds
(3D value noise FBM, Beer–Lambert transmittance, Henyey–Greenstein
phase, in-scattering toward the sun). It draws first, and the meshlet
stage’s colour is blitted over it — alpha = 0 is the background
sentinel, so pixels no meshlet covered keep the sky.
⚠️
GpuContextdeliberately selects a non-sRGB surface format, on the reasoning that “most renderers handle gamma correction in the shader”. Inti does. The sky pass does not. If the two disagree on brightness, that is the sky’s half of a decision taken long ago and never finished.
Debug views
MeshletDebugMode is a Resource the editor sets per frame; the shaders
branch on a single u32. Off is the production path.
| Mode | Shows |
|---|---|
MeshletIds / InstanceIds | Cluster boundaries; per-entity coverage |
TriangleDensity | Triangles drawn per pixel — calibrates target_error_pixels. Anything brighter than green is sub-pixel triangle territory |
Overdraw | Visibility-buffer atomic writes per pixel |
FrustumRejected / BackfaceRejected / HiZRejected | What each cull stage discarded |
CullPassthrough | Everything that survived every stage |
OnlyLod0 / OnlyRoots | The two extremes of the LOD chain, in isolation |
Normals | The world-space normal as colour |
ShadowCascades / ContactShadows | What each shadow mechanism saw — see Inti |
SingleLight | The selected light, alone, in grey, with its shadow |
LightsPerPixel | How many lights the pixel actually evaluated. Cost becomes a property of where the pixel is, which no pass timing can show — raster + shade is one number for the whole screen. 🔴 A flat maximum means the frame is not clustering: every light, every pixel |
PointShadowFactor | One point light’s cube map answering for itself — no BRDF, no cosine, no exposure, no second light. Magenta: no casting lamp. Blue: past its range. Grey ramp: the factor |
PointCubeFaces | The cube map itself, six faces in a 3×2 grid (+X, −X, +Y, −Y, +Z, −Z). Dark blue is nothing recorded, which is what an occluder culled out of the map looks like |
The last three exist because “the shadow is not there” is four faults wearing one pixel — no lamp near this point casts, the point is past the lamp’s reach, the cube says lit because the occluder never reached the map, or the cube says dark and the other lamps fill it back in. Four fixes, one colour in a shaded frame.
Normals deserves a note: until #441 it was the shading model. The
renderer computed normal * 0.5 + 0.5 and multiplied by albedo, which is
why a scene with lights and a scene without them rendered identically.
It survives as a debug view because it is a genuinely useful look at the
geometry — it just stopped being what you get by default.
The atomic-counter modes need TEXTURE_ATOMIC; the editor’s dropdown
hides what the adapter cannot run rather than offering a mode that
silently falls back.
The Inti-side views — Normals, ShadowCascades, ContactShadows,
SingleLight — are not compiled into the shader a game runs. They
live in inti_debug.wgsl, which only the editor’s second pipeline
concatenates; production takes INTI_DEBUG_STUB instead and the call
sites fold to if (false). The reasoning, and why an untaken branch is
not free, is in Inti.
Depth: reversed-Z, and no far plane
The camera’s projection is perspective_infinite_rh_reverse_z. Near maps
to ndc.z = 1; infinity approaches 0 without reaching it. Depth
attachments clear to 0.0 and compare Greater.
The property worth knowing, because half the renderer leans on it:
ndc.z == near / distance
Exactly. Any shader recovers metres from the depth buffer with one
divide and no extra uniform — which is why the contact-shadow march can
take thickness and length in world units and have them mean the same
thing in every scene. With a finite far plane it takes two coefficients
plumbed to every consumer, and the first one that forgets ships a
parameter documented in metres that does not measure metres.
Two things follow, and both are load-bearing:
- The far plane is gone from culling too. That row of the projection
degenerates to a zero-length normal;
extract_frustum_planesreturns[0,0,0,0]for it and the cull shader walks five planes. - Unprojecting uses the NEAR plane.
ndc.z = 0is infinity now and unprojects tow = 0. Anything that builds a ray from a cursor takesndc.z = 1— same ray through the eye, always finite.
The bounded perspective_rh_reverse_z survives for shadow cascades: a
slice of an unbounded frustum is unbounded. Rationale and the full list
of what this touched: ADR 0002.
Limits worth knowing
- 🔴 65 536 instances. The visibility buffer packs
(instance_id << 16) | meshlet_id. A chunk of vegetation exhausts this. Bevy removed their equivalent limit in 0.17 with BVH culling. - 🔴 Six bind groups, six used. The two-pass shading pipeline uses
every group
TARGET_MAX_BIND_GROUPSallows. Shadow maps have to go inside Inti’s group — which is where they belong anyway, since a shadow map without its light is not a thing any shader wants. Raising the target to 8 would work on desktop and drop a baseline Vulkan only guarantees at 4. - Skinned meshes cull against their bind pose (#453), so an animation that reaches outside the rest volume culls a character who is on screen.
- The R32 path has no motion vectors and no history, so no TAA and no temporal upscaling there. Jitter on that path is a wobble and nothing else, which is why it is not applied.
- TAA ships off (#481). The resolve is built and the vectors feed it; what is missing is the half that turns it on by default without softening a still image.
Not in the pipeline yet
Shadows, contact shadows and clustered shading used to be listed here. Cascades (#476) and contact shadows (#735) shipped, and the froxel grid (#780) runs every frame — the passes in the diagram above are what replaced those bullets. What is genuinely still absent:
- Virtual shadow maps (#477). Cascades cover the sun; a hundred shadowed point lights each want a cube map, and that is the wall VSM exists to move.
- Global illumination (#450) — surfel + voxel, not raytraced. Its absence is why punctual light defaults are larger than physics says they should be; see Lighting.
- Atmosphere (#250, #248) — correct from orbit, and tinting the sunlight.
- The post-processing stack
(#254) — AgX, SMAA,
vignette. The
tonemapandrcaspasses exist; the stack around them does not, and exposure is a setting rather than an auto-exposure loop.
Why there is no render graph
There was one — kooch_render::graph, 497 lines, cycle detection and
topological sort — and nothing ever instantiated it. The real
renderer was built beside it.
The decision not to revive it is not laziness. Bevy 0.19 deleted their
RenderGraph and replaced it with ECS schedules, because the graph ran
as an exclusive system and was single-threaded — the engine that made
the pattern canonical retired it. Kóoch already has the replacement half
written: kooch_core’s scheduler batches GPU systems into a shared
encoder. What it needs is before / after ordering, not a second
scheduler that looks official and is not
(#392).
Lighting — Inti
Inti is Kóoch’s lighting system, named for the Inca sun.
The name covers the whole thing, not one crate: extraction, the GPU light record, the shading model, shadows, clustering, light textures, and the global illumination that is still to come. When something says “the Inti path” it means the same way “the meshlet path” does.
The crate is still called kooch_lighting. Renaming a crate rewrites
every serialised type_name in every .scene and .prefab in every
project — silently, since nothing checks that a type it cannot find used
to exist.
Until #441,
kooch_lighting/src/lib.rswas nine lines: a doc comment promising point, spot, directional and area lights, volumetrics and bloom, and aninit()that logged. The three light components existed, the editor drew their gizmos, the Inspector edited them, the remote protocol mirrored them — and no render crate read one. You could place a light and nothing on screen would change.
How a light reaches a pixel
flowchart LR
C["DirectionalLight<br/>PointLight<br/>SpotLight<br/>+ GlobalTransform"] --> E[extract_lights<br/>pure, no GPU]
E --> B[GpuLights<br/>storage buffer,<br/>grows geometrically]
B --> G["the froxel grid<br/>four GPU passes,<br/>per view"]
G --> S["inti_shade()<br/>in both shading paths"]
B --> S
S --> T["inti_tonemap<br/>exposure → ACES → sRGB<br/>inline on the fragment path,<br/>its own pass on the compute one"]
style C fill:#1e3a5f,stroke:#4d8fbe,color:#fff
style G fill:#1e5f3a,stroke:#4dbe8f,color:#fff
style S fill:#5f3a1e,stroke:#be8f4d,color:#fff
The light component is the source. Not the sky’s sun direction — an
earlier draft of #441 drove shading from SkyRenderer.sun_direction,
which would have delivered a lit-looking scene and left the three light
components exactly as inert as they were.
A directional light’s direction comes from its transform’s -Z, never from a field. A light that ignores its own rotation is a second source of truth, and the editor’s gizmo already draws the arrow from the first one.
A light with no GlobalTransform is skipped, not defaulted: it has no
direction and no position, and putting it at the origin pointing down
would be an invention that renders.
The shading model
Cook-Torrance, ported from Bevy 0.19’s pbr_lighting.wgsl — read from
source, not reconstructed from memory. Which matters, because three of
their fixes are baked in from the start rather than rediscovered later:
| Term | What it is |
|---|---|
| D | GGX / Trowbridge-Reitz, in Filament’s reassociated form. The naïve expression loses catastrophic f32 precision at low roughness and the highlight breaks into visible blocks |
| V | Height-correlated Smith (Heitz 2014), returning G / (4·NoV·NoL) combined — so the specular term must not divide again |
| F | Plain Schlick, with f90 derived from f0. A near-black dielectric with f90 = 1 grows a white rim at grazing angles that no real material has |
| Diffuse | Burley / Disney. Lambert is flat; this brightens the grazing edge on rough surfaces the way cloth and unfinished wood do |
| Multiscatter | Single-scattering GGX loses energy on rough metals — they go grey. Compensated by the split-sum integral’s analytic fit |
Two decisions worth stating because they diverge from something:
- The diffuse is weighted by
(1 - F). Energy the specular layer reflected is energy the diffuse layer underneath never receives. Bevy’s forward path still adds the two lobes unweighted, the way Filament does; their path tracer does the layering. We took the path tracer’s form, because a mirror is where the difference shows and a mirror is not an edge case. - Point and spot lights have a
radius, so a lamp has a size and the highlight it leaves has one too. See below.
Light size — radius
PointLight::radius and SpotLight::radius are the radius of the
emitting sphere in world units. 0, the default, is the mathematical
point every light in the engine was before, so an existing scene renders
unchanged.
It is specular only. It does not soften the diffuse falloff, and it does not soften shadows — a soft shadow is a separate technique driven by the same number, not a consequence of this one.
The technique is Karis 2013’s representative point: instead of integrating the BRDF over the sphere, shade against the single point on the sphere closest to the mirror ray, and widen the roughness to account for the rest of it. Five things travel together and the highlight is wrong without any one of them:
| Piece | Why it is not optional |
|---|---|
| The representative direction | The highlight has to move, not only spread |
A separate N·L for the specular layer | The two layers now answer to different directions, so the cosine is applied per layer rather than factored out |
(a / a_prime)² normalization | Spreading a fixed amount of light must not add any — without it radius is a brightness knob |
mix(a, a_prime, 1-(1-a)⁴) | Feeding the widened roughness straight to the BRDF makes smooth materials read too rough and too dim. Bevy’s own comment names Linearly Transformed Cosines as the real fix and this as the tuned stand-in |
Sphere visibility, r²/d² | At a grazing angle part of the sphere is below the horizon and cannot light the surface at all |
⚠️ The clamp max(0.0001, dot(offset, R)) in the representative point is
a fix, not a guard against division by zero: the approximation is
plainly wrong for a surface inside or touching the light’s sphere, and
without the clamp such a surface shows a hard discontinuity. It is
carried from Bevy, who carry it for
bevyengine/bevy#13318.
A directional light is excluded by kind, not only by its radius: there is no distance to a light with no position for the approximation to correct. A sun’s angular size is a shadow problem, not this one.
Ambient
A hemisphere lerp between a sky colour and a ground colour, authored in the project’s settings asset. It is not cosmetic: with no ambient term a metal facing away from every light renders pure black — correct for the model, and indistinguishable from a bug to whoever is looking at it.
It is a placeholder for a value the scene should compute, not author.
Ambient light is the sky, and the sky is about to become something the
engine simulates: atmospheric scattering
(#250,
#248) already has to
know what colour the air is in every direction, and Bevy’s 0.18
atmosphere lights the scene rather than only being drawn behind it. Once
that exists, ambient_sky_color stops being two colours someone picked
and becomes the atmosphere sampled — different at noon, at sunset, at
altitude and in orbit, without anyone touching a field.
The authored values stay useful for a scene with no atmosphere: an interior, a space station, a stylised game that wants flat fill. They just stop being the default answer.
⚠️ Today it lerps on world up, which stops meaning anything on the far side of a planet. A known limit of the placeholder, and one the atmosphere fixes on its way past: a per-planet atmosphere knows which way up is, because up is what it is a sphere around.
Units, and why the defaults are not physical
Lights carry real photometric units:
| Light | Unit | Default |
|---|---|---|
DirectionalLight | lux (illuminance) | 10 000 — lux::AMBIENT_DAYLIGHT |
PointLight / SpotLight | lumens (luminous flux) | 32 000 — lumens::ROOM_LIGHT_NO_GI |
The directional default is a physical fact and matches Bevy’s. The punctual default is forty times a real 9 W bulb, and that deserves an explanation rather than a shrug.
An 800 lm bulb three metres away really does deliver about 7 lux. An office reads 320 lux, and the other 313 are bounces — light off the ceiling, the walls, the desk. Kóoch computes direct light only, so the physically honest number renders as almost nothing.
Bevy has the same gap and resolved it by defaulting PointLight to
VERY_LARGE_CINEMA_LIGHT, one million lumens, with the comment “capable
of registering brightly at Bevy’s default exposure level”. That is a
confession, not a unit. Kóoch’s fudge is named after the compromise —
ROOM_LIGHT_NO_GI — and its own doc comment says it goes back to a real
bulb the day global illumination lands.
kooch_ecs::light_consts holds the named values, so an author picks a
situation instead of guessing a magnitude: lux::OFFICE,
lux::OVERCAST_DAY, lux::DIRECT_SUNLIGHT, lumens::CANDLE,
lumens::CAR_HEADLIGHT.
Hover a field in the Inspector and its doc comment appears as a tooltip, units included.
intensityon a directional light says LUX; on a point light it says LUMENS. They are different units with different magnitudes and they used to look identical.
Exposure
Physical light units need an exposure step or every channel clips to white and the model looks broken rather than unexposed.
Exposure carries an EV100, and PhysicalCamera is the control worth
using: aperture, shutter and ISO. f/16, 1/125, ISO 100 says
something to anyone who has held a camera; EV100 = 9.7 says nothing
about which way is brighter or what a step is worth.
| Preset | Settings | EV100 |
|---|---|---|
PhysicalCamera::sunny() | f/16, 1/125 s, ISO 100 | ≈ 15 |
PhysicalCamera::default() | f/2.8, 1/125 s, ISO 100 | ≈ 9.9 |
PhysicalCamera::indoor() | f/1.0, 1/125 s, ISO 100 | ≈ 7 |
The default is a middle setting, not a real situation: bright enough that a default sun does not clip, dim enough that a punctual light is visible. It lands near Bevy’s 9.7 so a scene authored against their numbers reads the same here — and note that their 9.7 is not “sunny 16” despite being described that way; sunny 16 is EV 15, and they calibrated theirs against Blender’s implicit exposure.
Tone mapping is an ACES approximation (Narkowicz 2015), then the sRGB transfer function. Both are provisional and belong to #254, which owns the real tonemapper and the auto exposure that lets a sunlit surface and a planet’s night side coexist in one frame.
Where the author sets it — .rendersettings (#744)
Exposure and ambient used to be Resources with defaults and no way to
change them: #441 built the control and left it out of reach, which is
this engine’s recurring failure committed knowingly. They are now fields
of a project settings asset — aperture_f_stops, shutter_speed_s,
sensitivity_iso, ambient_sky_color, ambient_ground_color,
ambient_intensity — edited in the Inspector like any other asset.
It is an asset rather than a panel because the machinery already existed: a RON loader that registers itself, reflection for the generic editor, the save-and-refresh path of #728, and the asset browser as its home. The evidence that bespoke settings panels do not get built is that this setting had none for as long as it existed.
⚠️ Author settings, not player settings. This ships with the game and
belongs in version control. What the player picks — resolution, volume,
key bindings — is #736 and lives under ~/.config/. Merged, they would
put an artist’s exposure and a volume slider in the same commit.
🔴 The fields are flat rather than a nested PhysicalCamera, and
every serde default is what the engine already did — see the render
pipeline page on why an old file silently taking a new default is how
“you broke the whole render” gets reported.
The GPU record
GpuLight is 80 bytes, #[repr(C)], mirroring IntiLight in
inti_pbr.wgsl byte for byte. Nothing checks that correspondence at
compile time on either side of the boundary — a reordered field reads a
light’s range as its intensity and renders something plausible and
wrong — so a test pins the size.
Array of structs, not struct of arrays, against the engine’s usual
rule, and for a reason that survives scrutiny: every shader invocation
touching light i reads all of light i’s fields within a few
instructions. Splitting into parallel arrays turns one cache line into
six scattered fetches. SoA pays when a pass reads one field across many
records — which is what light culling does, positions and ranges
only, and that is exactly how the froxel grid takes its input: a separate
pair of arrays, not a reinterpretation of this one.
It was 64 bytes — one cache line — until radius arrived and there was
nowhere left to put it. Growing the record widens every light in the
scene, the sun included, which is a bandwidth decision on a handheld and
not a free slot; three padding scalars now ride along to the alignment
std430 demands, and they are where the next field goes before the cost
is paid again. Bevy’s equivalent record is 80 bytes for the same reason,
and they keep directional and rect lights in separate arrays rather than
widening one record for all of them.
Spot cones are stored pre-packed as the multiply-add the shader
evaluates, saturate(cos_angle · scale + offset) — one MAD per light per
fragment instead of a subtract and a divide. The authored half-angles are
recoverable from it.
Angles are half-angles, measured axis to edge, like Unreal rather
than Unity’s single full spotAngle. gizmos/lights.rs chose that
convention when it drew the cone and wrote down that the lighting work
would either honour it or draw a cone half the width it lights.
Shadows
Two techniques, and they compose rather than compete because each is worst where the other is best.
Cascades, for the scene (#476)
Four cascades fitted to slices of the view frustum, rendered into one
atlas, sampled with Castaño ’13 — nine bilinear taps — under a PCSS
penumbra that widens with the gap between blocker and receiver
(sun_softness is the tangent of the sun’s angular radius, not a width).
Only the first active directional light gets cascades. That is a
statement about cascades, not about punctual shadows: a spot casts
through its own perspective map (#777) and a point through a cube (#778),
both below, and both default to cast_shadows: true.
The cascade fit is the one thing in the renderer that still asks for a bounded projection — a slice of an unbounded frustum is unbounded. See ADR 0002.
Contact shadows, for the last few centimetres (#735)
A cascade is correct at range and worst exactly at contact: at the texel density it can afford, the few centimetres where an object meets the ground is where its shadow detaches or swims, and that is what makes things look like they float over a scene rather than stand in it.
A short ray marched through the depth buffer, from the shaded point towards each light that opted in. Screen-space, so it costs the same at any world scale — a ray a few pixels long is a few pixels long whether the object is a crate or a moon.
The march is bevy_raymarch.wgsl: Bevy 0.19’s bevy_pbr::raymarch
copied, licence header intact. Diff it against upstream rather than
reasoning about it. What is this engine’s lives in contact_shadow.wgsl
(bindings, the four view helpers the port imports) and
contact_shadow_apply.wgsl (the call, the debug probe, the lift below) —
the line between the two is a file boundary, not a judgement call.
🔴 The one thing that could not be copied. Bevy’s march is compiled
only behind #ifdef DEPTH_PREPASS, so their ray’s origin and their depth
buffer came out of the same rasteriser with the same matrix and agree to
the bit inside the origin’s own texel. This engine reconstructs the
origin from the visibility buffer by barycentrics — a second
arithmetic path to the same point — and inside that texel the comparison
is decided by the last bit, with the jitter picking which way. It renders
as salt and pepper across every lit surface.
The fix is a lift: the ray starts one depth texel off the surface
along the normal, divided by n·v. Both factors are derived, not
authored — the texel’s world size from view_proj[1][1] and the buffer
height, the distance from near / ndc.z. The n·v term is the
slope-scaled depth bias every shadow map uses: a depth texel is a
screen quantity, so it spans more surface the more oblique the surface
is, and the error inside it grows in the same proportion. Clamped at four
texels, because n·v → 0 at a silhouette and an unbounded lift would
throw the ray clear of the object.
Per light, opt-in (DirectionalLight::contact_shadows on by default,
punctual lights off): the cost scales with light count, and a scene
has one sun and can have fifty lamps. The whole feature turns off with
contact_shadow_steps = 0.
⚠️ Screen-space means an occluder off-screen or behind the camera does not exist; the ray is clipped to the frustum and reports no hit, so the shadow fades at the screen edge rather than popping.
⚠️ The seam is real and is not a bug. Next to a nine-tap penumbra, a contact shadow is one ray with a hard hit/miss answer, so the boundary between pixels that find an occluder and pixels that do not is a discontinuity in coverage. No curve applied to the hit smooths it — softening Bevy’s remap was tried, cost the shadow 30% of its strength for nothing, and was reverted. Only averaging fixes it, spatially or temporally, and temporally is #732. Bevy has the same seam and leaves it to their TAA.
Seeing what the shadow system did
Three debug views, because “no light reaches this” and “something shadows this” look identical in a shaded frame and have different fixes.
- Shadow cascades — Bevy’s flat per-cascade hue, dimmed by
inti_sample_cascade, the same call the shading pass makes. Magenta: nothing casts. Black: inside no cascade volume. - Contact shadows — red: the march hit on its first step, which is the surface occluding itself. Green: a real occluder. Blue: the ray was under two pixels long. Grey: marched and found nothing.
- Single light — one light, alone, in grey, with its shadow. The one that answers what a shadow looks like; the other two answer what the shadow system did.
A fourth answers a question about cost rather than about appearance: Lights per pixel, below, for how many lights a pixel pays for.
Single light
Select a light in the World panel and this view shades the scene with that light and nothing else: no other light, no albedo, no ambient.
Each exclusion answers a different confusion:
| Removed | Because otherwise |
|---|---|
| Every other light | A surface lit by the wrong one still looks lit |
| The material’s colour | A dark albedo and no light landing produce the same pixel |
| Ambient | A point in full shadow never renders black, which is the reading the view exists to make unambiguous |
Roughness is kept — the width of a highlight is information about the
light, not the paint — and metallic is forced off, because a metal
takes its F0 from the albedo this view removes, and a metal shaded white
is a mirror rather than that metal decoloured.
The shadow comes from calling inti_light_contribution, the same
function the shading loop sums per light. A view that recomputed the
maths its own way could disagree with the frame, and then it is one more
thing to debug instead of the thing that settles the question.
⚠️ A light that shows no shadow is not necessarily broken. A spot casts since #777 and a point since #778, but both are limited to four maps each: past the budget a light keeps lighting the scene and stops casting. “Casts nothing” and “the shadow broke” render identically, so the editor prints which one it is next to the selector —
shadow_noteinkooch_lighting.
Magenta means the selection has no slot in the light buffer, and the note below the selector says which reason:
| Note | What happened |
|---|---|
This light is inactive — tick active in the Inspector | The light exists and is switched off, so it never reached the buffer |
Select a light in the World panel | The selection is not a light |
Those two produce the same magenta, and the first smoke of this view
hit it: two lights in the scene were active: false, so selecting them
looked exactly like selecting a crate. A view built to stop two causes
from looking alike does not get to ship a third pair of its own.
Which light travels in IntiFrame.debug_light, an index into the light
buffer that occupies what used to be that struct’s tail padding. There is
no seventh bind group and Inti’s is full, so a view needing a binding of
its own was not going to ship.
Spot lights (#777)
A cascade is an orthographic slice of the camera’s frustum and needs fitting, splitting and stabilising. A spot needs none of it: the light is a frustum, so its shadow view is its own cone and there is one map with nothing to blend into.
It renders into a layer of the same array the cascades use, behind them,
and its record is the same GpuCascade — inti_shadow_coords already
divided by w, which an orthographic cascade does not need and a
perspective does. That means the bias, the blocker search, the Castano
filter and the border clamp are one implementation, not two.
| Decision | Why |
|---|---|
FOV is outer_angle * 2 | The cone’s edge is where the light stops. Fitting the half-angle clips the round pool into a square |
| Up vector chosen against the cone | A spot pointing straight down is the most ordinary way to author one, and the case a fixed world-up basis is degenerate for |
| Cone clamped near 90° | tan runs away before it gets there and fills the matrix with infinities, which spread into every depth the pass writes |
texel_world_size is an ANGLE per texel | 2·tan(outer)/size·√2, Bevy’s texel_size, with no range in it. The shader multiplies by each fragment’s own axial distance to the light. Baking range in biases everything as though it sat at the far end of the cone: a 100 m spot over objects five metres away offset them 17 cm and the shadow lifted off — found in the first smoke |
MAX_SPOT_SHADOWS = 4 | One layer each, 16 MiB at 2048². A fifth spot still lights the scene with no shadow; dropping the light would be a worse failure |
🔴 The cull needs its LOD selector configured, and a factor of zero is
not a neutral default. CullParams::new leaves
lod_error_to_pixel_factor at 0.0, which makes every meshlet’s
projected error work out to 0 px — always under the threshold, so the
selector keeps only roots. A sphere’s shadow then comes out as a
wedge. projection_scale_y’s doc comment has said so since a rotated
camera hit the same zero: “a sphere collapses to a blob and a cube to a
spike”. A spot uses with_lod (perspective), a cascade
with_orthographic_lod; both apply SHADOW_LOD_RELAXATION, because a
shadow is a silhouette and loses detail a lit surface keeps.
🔴 Slots are handed out inside extract_lights, in its walk order, and
shadow_casting_spots reads that same order back. Two walks that
disagreed would light one spot through another’s map — geometry from
elsewhere in the room, which reads as a broken shadow pass rather than as
a crossed index.
🔴 No sun is not no shadows. A scene lit by a torch and nothing else
still renders its maps. The cascades are then fitted to a stand-in
direction so the pass has something coherent to not draw, and
FrameShadows::cascades_enabled stays false — otherwise a directional
light that does not cast would sample them and be shadowed by a sun
that is not there.
Point lights (#778)
A spot light is a frustum, so its shadow is the cone itself. A point light is not a frustum at all — it lights every direction — so it gets six 90° faces that tile the sphere, and the shading model picks one by the direction to the fragment rather than by a matrix.
That is why GpuPointShadow is sixteen bytes and carries no
view_proj: sampling a cube map takes a direction, and the direction
is a subtraction the shader already does.
| Decision | Why |
|---|---|
| A separate texture from the cascade array | Different size, and one texture cannot have two. Sharing would mean six 2048² faces per light — 96 MiB each against 6 MiB at the size a lamp actually needs |
DEFAULT_CUBE_SIZE = 512 | 6 MiB per light at Depth32Float. Bevy’s is 1024; ours is smaller because these render on a handheld and the shadow of a lamp is wanted soft, not detailed |
MAX_POINT_SHADOWS = 4 | Memory, not the technique. The max_texture_array_layers ceiling of 256 would allow 42, and 42 × 6 MiB is a quarter of a gigabyte of depth |
| Six culls, shared across lights | A light’s six faces are what can overlap on the GPU. Lights then serialise, costing three barriers at the limit against eighteen more survivor arenas idle whenever nothing casts |
| Casting lights ranked by distance to the camera | Past the limit a light keeps lighting and stops casting. Which one loses its shadow must not be decided by when it was spawned |
The depth is one divide, and that is the whole reconstruction
Bevy sends the lower-right 2×2 of the face projection per light and
computes depth = zw.x / zw.y. Expanded with a standard perspective,
w collapses to the major axis and depth to m23/major − m22 — and
with the infinite reverse-Z projection this engine migrated to
(ADR 0002), m22 = 0 and m23 = near:
depth = near / major_axis_magnitude
One scalar replaces their four. depth_is_near_over_the_major_axis
checks it against the real matrix rather than trusting the algebra,
because being wrong here reads as a bias that cannot be tuned rather
than as a formula that is visibly wrong.
🔴 distance_to_light is max(|x|, |y|, |z|), never length(). The
faces align with the world axes and their frustum planes meet at 45°, so
the largest absolute component is the depth. The Euclidean distance
would scale the bias by up to √3 toward the corners of a face — the same
class of mistake as the axial-vs-radial spot bias above, which shipped
and had to be fixed.
🔴 Cube maps are left-handed and this engine is not. The Z faces are
stored swapped in FACE_DIRECTIONS and the sampling direction is
mirrored on Z. Correcting either half on its own puts the shadow of
everything in front of a lamp behind it.
The filter is a gaussian, not Castano
The cascades and spots use Castano’s thirteen — a 2D gaussian that leans on bilinear hardware to get nine taps out of four fetches. That trick does not exist for a cube map, so Bevy’s cube path (and ours) is eight explicit taps at the standard D3D MSAA positions, weighted by a gaussian whose coefficients sum to 1.
A cube map has no uv plane to offset a tap in either, so the offsets move across the tangent plane of the sampling direction, built per pixel by a branchless orthonormal basis (Duff et al. 2017).
The radius is INTI_POINT_FILTER_TEXELS shadow texels, where Bevy
uses a fixed 0.003 in direction units: expressing it in texels stops it
silently changing meaning if the face size ever moves.
⚠️ The offset is added to a direction vector whose length is already the distance to the light, so the texel size is converted to metres once, by the caller. Doing it again inside the filter made the radius grow with the distance squared — which does not look like a wider blur, it looks like a gradient smeared across the floor.
A surface can opt out of receiving them
MeshRenderer::receive_shadows clears a bit on the instance, and
inti_light_contribution skips the shadow fetch entirely — not a
cheaper fetch, none at all. It covers the cascades, the spot and point
maps and the contact-shadow march, because all four are shadows.
Worth it because that cost is per pixel and per casting light: the same product that makes lighting expensive at all (#780 attacks it from the other side, by shrinking the set of lights a pixel considers). A ground plane already in shade, backfaces, emissive surfaces and anything an author knows will never show a shadow are all paying for a sample they discard.
🔴 The field existed long before anything read it. Unticking it in
the Inspector changed nothing, which is worse than the feature being
absent — the UI made a promise the renderer did not keep. Same shape as
the cast_shadows checkbox on a point light before #778.
a_floor_that_receives_no_shadows_has_none is what keeps it honest, and
it was verified failing: the same 0.0791 with the flag and without it.
⚠️ Bevy checks two flags before its fetch, one on the mesh and one on
the light (pbr_functions.wesl). This is the mesh half. The light half
is cast_shadows on MeshRenderer, which is still read by nobody —
a mesh with it unticked casts anyway.
What a cube costs, and when it costs nothing
Six faces is the most expensive shadow the engine draws, and
cast_shadows defaults to true, so the work has to be avoided
rather than paid.
Culled against the camera’s frustum, then limited — in that order.
Limiting first would spend all four cubes on the nearest lights even when
they are behind the camera, and the one lamp whose shadow anybody can see
would get nothing. The test is the sphere of the light’s own range, not
its centre: a lamp just off the edge of the screen still shadows pixels
that are on it.
A cube is redrawn only when something it depends on changed. The key is the light’s identity, its position, and a hash of every instance in the frame. Epic measures a cached local shadow map at 0.05 ms against 0.4–0.8 ms invalidated on a PS5, and a lamp bolted to a wall in a room where nothing moves should pay the first number.
🔴 That key is deliberately coarse: a crate moving across the level invalidates a lamp that cannot see it. A cube redrawn for nothing costs a frame’s work; a cube not redrawn when it should have been is a shadow frozen in place — silent, and blamed on the light, the material and the camera long before the thing that skipped the work. Narrowing it means asking which instances a light’s range reaches, which is the cluster structure (#780) and not this.
⚠️ Two tests render twice, because the first frame of a cached cube is always drawn and a stale cube only appears on the second. Both were verified to fail against a cache that never invalidates.
A light past the budget says so, once per transition rather than per frame. It keeps lighting the scene without a shadow, which is the right failure — but an author staring at a lamp with no shadow could not tell that from a bug.
The debug views are not in the shader your game runs
inti_debug.wgsl is concatenated only by a pipeline that can show a
view. Production concatenates INTI_DEBUG_STUB, where
inti_debug_is_view returns a literal false and the two call sites
fold away.
This is a performance decision, not tidiness. A branch nothing takes is
still code the shader carries: register allocation is worst-case over the
whole entry point, so a cascade sample and a screen-space raymarch parked
behind if (debug_mode == …) still raise the VGPR count. VGPR count caps
how many waves stay in flight, and waves in flight is the whole of an
integrated GPU’s ability to hide memory latency — which is the budget
this engine is held to.
The debug pipeline is built through a OnceLock the first time a view is
selected, so a shipped game never compiles it either. Both variants are
validated by tests, because nothing else compiles the debug one until
somebody opens it.
Clustering — the froxel grid (#780)
Until this landed, inti_shade looped over every light in the scene
for every pixel on screen, and a lamp on the other side of the map was
evaluated — falloff, cone, shadow-map sample and all — in every fragment.
The cost was pixels × lights, multiplied, and it was measured as the
frame’s largest single term on the OneXFly: raster + shade scaled
worse than linearly with resolution, because every new pixel paid for
the whole light list again.
The fix is a spatial index. The view frustum is diced into a grid of cells — “froxels”, frustum voxels — and each cell is given the list of lights whose volume reaches it. A fragment looks up its own cell and walks that list.
The grid is not a light structure
Reflection probes, irradiance volumes and decals are bound to a region of space in exactly the same way, and each cell’s record reserves a range for all five types from the start. It is also the structure virtual shadow maps (#477) mark pages with, and the one volumetric fog (#731) integrates through. It gets built once. Growing a second grid for each of those is the failure mode this shape exists to avoid.
Four passes, and why one of them is a rasterizer
flowchart TD
Z["z-slice<br/>compute, one thread per light"] --> F["finalize<br/>clamp the draw args"]
F --> C["count<br/>raster, one fragment per cell-light pair"]
C --> A["allocate<br/>compute, prefix sum"]
A --> P["populate<br/>raster, the same source again"]
style C fill:#5f3a1e,stroke:#be8f4d,color:#fff
style P fill:#5f3a1e,stroke:#be8f4d,color:#fff
The two middle passes are draws, not dispatches, and that is the part most descriptions of clustering get wrong. The grid is WxHxD; the pass runs on a WxH viewport and draws each (light, slice) pair as a quad covering the cells that light can reach. One fragment invocation is then exactly one (cell, light) pair — scheduled by the hardware that exists to schedule quads. Colour writes are off; the output is storage buffers.
It runs twice because the lists are tightly packed: the counting pass is what makes the offsets computable, and the offsets are what the populate pass writes into. 🔴 Both runs must reach the same verdict for every pair. They are the same source compiled twice for that reason — a disagreement would overflow one cell’s run into its neighbour’s, and nothing downstream could detect it.
Slices are logarithmic
slice = ln(-view_z) * factor.x - factor.y + 1
Cells are distributed the way depth precision falls off rather than by
metres: thin near the camera, thick far away. Slice 0 is everything
nearer than ClusterSettings::first_slice.
⚠️ The grid needs a far plane and the camera does not have one. Kóoch
projects with an infinite reversed-Z frustum (ADR 0002). Bevy reads back
the furthest light the GPU saw and resizes next frame; that is a readback
in the hot path. Here it is a setting — ClusterSettings::far, 200 m by
default. A light further out lands in the last slice with everything else
behind it, so that cell holds more lights than it should. Nothing renders
wrong; it just stops saving work out there.
Directional lights are not in it
They reach every cell, so a cell listing them would say nothing. They are
the leading entries of the light buffer and the shader walks them
linearly — which is why ExtractedLights::directional_count is a prefix
and not a subset.
Seeing what a pixel pays — Lights per pixel (#817)
Clustering makes cost a property of where the pixel is. That is the
whole point of the grid, and it is also why no pass timing can explain a
slow frame any more: raster + shade is one number for the screen, and
on the OneXFly it grew from 5.27 ms to 34.92 ms between a still camera
and a moving one without saying which pixels did it.
The debug view paints the count each fragment actually walks, read at
the point where it is paid — the same point_count + spot_count that
bounds the loop in inti_clustered_lights, plus the directional lights,
which the grid does not cluster because they reach every cell.
| colour | meaning |
|---|---|
| black | no light reaches this pixel at all |
| blue | few |
| green | half the top of scale |
| red | the top of scale, or more |
The top of scale is a control, not a constant. It is the one number
that decides whether the picture says anything: at 16 a hundred-light
stress scene is flat red, and the same frame at 40 separates into
froxels. Raise it until the image stops being flat — that value is
roughly what the busiest froxel carries. It rides in the frame uniform
(LightsHot), so moving it costs no recompile.
Fixed during a comparison, though: two screenshots taken at different tops mean nothing next to each other.
🔴 A whole screen at full red means the frame is not clustering.
inti.clustered == 0 evaluates every light for every pixel, which is
what the frame cost before #780 and what a path with no camera matrices
still does. The view does not special-case it, because a scene that
quietly stopped clustering should look alarming.
A light in the cell is not a light on the pixel (#835)
The cell is conservative by design: cluster_raster.wgsl accepts any
light whose bounding sphere touches the cell’s AABB, because a cell that
excluded a light reaching one of its pixels would drop light from the
image. A cell is also much larger than a pixel — at 1080p the default
grid is 17x9x24, so one cell covers roughly 113x120 pixels and a slab of
depth besides.
Both of those are correct, and together they mean a fragment’s list contains lights that do not reach it. Measured in #820: the busiest cell carries ~40 lights where ~14 reach the pixel.
So the shading loop asks a second time, at the pixel, where the answer is already in a register:
let reach = max(max(s.irradiance.x, s.irradiance.y), s.irradiance.z) * n_dot_l;
if (reach <= 0.0) {
return vec3<f32>(0.0);
}
🔴 Zero here is exact, not small. inti_distance_attenuation windows
with saturate(1.0 - factor * factor), which reaches zero at the range
rather than approaching it — the same property that makes the editor’s
wire sphere the truth about where a light stops. So the cut returns the
value the rest of the function would have computed: both BRDF layers, the
shadow cube and the contact march, all multiplied by an irradiance of
zero. No pixel changes, which is what tests/light_reach.rs pins
byte-for-byte.
What it removes is the work in between, and that work is not small: at the default of 16 contact-shadow steps, the ~26 unreachable lights in a busy cell were spending 416 depth taps per pixel to produce nothing.
⚠️ This is not a fix for the cell being loose. The cell is supposed to be loose. Tightening the grid — more slices, smaller tiles — trades against the cost of building it, and the measurement that would justify it is a different one.
A cut at a threshold rather than at zero is the next question, and it is a different kind of change: visible, tunable, and needing its own measurement. That is what the section below is.
❌ What does not work: cutting the list by count (#826)
Recorded because it is the obvious next idea and it is wrong.
If a cell carries ~40 lights and ~14 reach the pixel, sampling k of them stochastically and weighting the result looks like the standard answer — RIS, in the shape ReSTIR made famous. It was built, measured and removed, and 1139 lines went with it.
It is incompatible with the grid’s continuity property. A light joins a cell exactly when its contribution reaches zero, which is what makes the boundary between two cells invisible. Choose k of n per froxel and the froxel becomes visible: neighbouring cells pick different subsets of the same lights, and the picture breaks into flickering blocks the size of a cell — 75×80 px at 1080p, repainted every frame.
The histogram says the same thing arithmetically. Over a hundred-light scene:
| lights kept | share of the froxel’s irradiance |
|---|---|
| top 2 | 60.8 % |
| top 6 | 95 % |
| top 8 | 93 % of the worst froxel |
There is no small k. Getting to 95 % takes six of the fourteen that reach the pixel, and the two the ordering drops are exactly the ones a neighbouring cell keeps.
⚠️ A related hypothesis died with it: that the workgroup memory the sampler used was costing occupancy. Removing it freed 9.88 KB of LDS and moved a settled handheld frame from 40.7 ms to 40.5 — inside the noise. The shading dispatch is not misconfigured. It is doing real work.
The direction that survives is not fewer lights, it is cheaper lights — which is the section below — and a smaller reach per light (#835 above).
What each light costs — specular_floor (#821)
Clustering bounds how many lights a pixel walks. It cannot make any of
them cheaper, and every one of them pays the full model: GGX D,
height-correlated Smith V, Schlick F, the multiscatter fit, and the
representative point when the light has a radius. With ~15 lights
reaching a pixel, that is fifteen of those.
A light whose irradiance at a point is a fraction of the frame’s
exposure leaves a highlight nobody can see, and pays the expensive half
to produce it. SpecularFloor is the irradiance below which it shades
diffuse only:
KOOCH_SPECULAR_FLOOR=2000 ./your-game
0.0 is the default and keeps every light on the full model, so a project that never sets it renders exactly as before.
⚠️ Fresnel is substituted at normal incidence rather than dropped. f
weights the diffuse layer — diffuse = (1 - f) · … — so a skipped
specular that also skipped f would brighten the surface. A missing
highlight is invisible; an over-lit dielectric is not.
🔴 The editor cannot measure this and the environment variable is not
a convenience. On a desktop GPU the whole raster pass is 0.12 ms and
switching every specular layer off moves it by 0.001 ms — there is no
bottleneck there to remove. The frame this exists for is a game on the
handheld, launched over SSH, with no editor in the process. A knob that
lives only in a panel cannot be swept on the machine whose numbers
decide anything, which is the same lesson KOOCH_CLUSTERING already
carried.
Turning it off
KOOCH_CLUSTERING=off, or ClusterSettings { enabled: false, .. } in
Resources. The image is identical and the cost is the linear walk this
replaced. It exists to be the A/B: same camera, same scene, one capture
each, is the only honest way to say what the grid bought.
What a page pool would hold — the census (#866)
Before there is a page pool there is a number, and #866 says so:
“the first task in this issue is a measurement, not an allocation”.
cargo run --example measure_shadow_pages -- <scene> is that
measurement. It walks the froxel grid, marks every page each cell would
need from each light that reaches it, and prints what the distinct pages
would cost.
cargo run --example measure_shadow_pages --features lighting -- \
../roll-a-ball/assets/scenes/many_lights.scene
The walk lives in kooch_render::shadow::pages and runs on the CPU on
purpose. The marking pass it previews belongs on the GPU — that is
#477 — but here it only counts, so it needs no device and can be a test.
It also becomes the oracle that pass is checked against, the position
ClusterGrid::z_slice already holds against cluster_z_slice in WGSL.
The configuration it measures, and where it comes from
Read off the UE 5.8 Virtual Shadow Maps documentation directly, not quoted from this project’s own issues:
| Unreal | in pages.rs | |
|---|---|---|
| virtual resolution | “16k x 16k pixels” | PageConfig::virtual_size |
| page | “tiles (or Pages) that are 128x128 each” | PageConfig::page |
| level selection | “appropriate mip levels are picked by projecting the size of the screen pixels into shadow map space” | level_for |
| spot | “a single 16k VSM with a mip chain rather than clipmaps” | CensusKind::Spot |
| point | “a cube map of 16k VSMs, one for each face” | CensusKind::Point |
| directional | “clipmap levels 6 through 22”, finest 64 cm from the camera, broadest ~40 km, every level at full 16k | ClipmapConfig::default |
| marking | “depth buffer analysis is used as the primary method of marking pages that are needed to render” | CensusFrame::surfaces |
| the budget | r.Shadow.Virtual.MaxPhysicalPages, 4096 by default; 6144 for open worlds; 8192 thrashes | POOL_PAGES |
🔴 The pool is one budget for the whole scene — every light, the sun
included, allocates out of it — and overflow is not graceful: Epic’s
page-pool overflow shows as checkerboard corruption or missing shadows.
At Depth32Float those 4096 pages are 256 MiB, which is more than
this engine’s 152 MiB of fixed allocations. The pool is not inherently
smaller. It is adaptive, and that is a different property.
What it measured, 2026-08-20
many_lights.scene at 1280x720 — a hundred point lights, a sun, a floor
and sixteen Suzannes — in exactly that configuration. The run reports two
walks of the same grid: volume marks every cell of the frustum,
surfaces marks only the cells a mesh passes through.
| cells | volume | surfaces | MiB | saved | |
|---|---|---|---|---|---|
| the sun | 131 | 15 770 | 118 | 7.4 | 133.6x |
| a hundred local lights | 131 | 8 386 | 6 798 | 424.9 | 1.2x |
| everything | 131 | 24 156 | 6 916 | 432.2 | 3.5x |
| the screen’s floor — one texel per pixel, perfectly packed | — | — | 57 | 3.6 | |
| today’s fixed allocations, for five casting lights | — | — | — | 152.0 | |
| Unreal’s default pool | — | — | 4 096 | 256.0 | |
| Unreal’s open-world pool | — | — | 6 144 | 384.0 |
🔴 The marking input is the decision, and the sun is where it shows. Marked from froxel volumes the sun’s clipmap residents 15 770 pages — 277x the theoretical floor. Marked from the cells that actually contain geometry it residents 118, about twice that floor. A froxel is a box of mostly empty air, and a page allocated for air is a page no shadow ever reads. Epic says the same thing in one sentence: “depth buffer analysis is used as the primary method of marking pages”.
So #866’s own opening move — read it off the froxel grid that already runs — is what the measurement refutes. The froxel grid answers which lights reach which region of space, which is the right input for shading and the wrong one for page allocation.
Two further sweeps are kept in the run because both were predictions about the volume walk and both refuted the mechanism they tested, which is what leaves that walk’s count standing as an area rather than an artefact of how the grid is diced: 32x thinner slices moved it 31 %, and over a 20x range of cell counts it moved 25 % — while the surface filter moves it 133x.
🔴 Neither page size nor virtual size is the decision. Across 64/128/256-texel pages the bill is flat within 2 % — 424 to 432 MiB — because a smaller page is simply more pages. Across 4k/8k/16k virtual maps residency is identical (6 916 pages each time), because the virtual size is only the chain’s ceiling and the level chosen for a cell is the one whose texels match the screen.
What it says about the engine as it stands
🎯 For the content that ships today — five casting lights — the pool is 14.1 MiB against 152. Eleven times less, same image. That is the whole promise of memory that follows the screen instead of the sum of every light type’s worst case, and it holds.
⚠️ But local lights barely benefit from better marking: 1.2x. A point
light’s range already bounds it to the cells near geometry, so there
is little air left to stop paying for. A hundred casting local lights
cost 424.9 MiB — 6 916 pages, which is past Epic’s open-world
recommendation of 6 144 and into the band they say thrashes. That the
census lands there is the best evidence its magnitude is right; it is
also the answer. The pool replaces a cap of four slots with a cap of
memory, and on a handheld that is still a cap. The next lever is the
density target, not the marking.
What the census does not model
- Coarse pages. Epic marks some low-resolution pages unconditionally “to ensure that at least low-resolution shadow data is available” for systems that sample at arbitrary locations — volumetric fog above all, which is #731 here. That is an additive constant this walk omits.
- Occlusion. A cell with geometry is an upper bound on a cell with visible geometry: a cell behind a wall still marks. The real depth-driven pass sits between the surfaces column and the floor.
- Off-screen casters. The walk covers the camera’s frustum, so it counts what has to be marked. Geometry off-screen still has to rasterise into those pages; marking and casting are separate questions and only the first is measured.
- Invalidation. Epic’s rules are harsh — “any light movement or rotation will invalidate all cached pages”, and moving geometry invalidates the pages its bounds overlap from the light’s view — and none of that is a residency question, so none of it is here.
The marking pass, on the GPU
The census is a model. KOOCH_PAGE_MARKING=1 runs the thing it models:
one compute dispatch over the depth buffer, in
kooch_render::shadow::pages::mark and page_mark.wgsl.
Where the controls are. Performance → Debug → Mark shadow pages, beside the froxel grid’s own A/B: a checkbox, the sampling rate, and the readout — pages, MiB, samples, sample/light pairs, and what share of Unreal’s 4096-page pool that is. The environment variables are only the defaults, for the comparison that gets made on a handheld over SSH against a build nobody wants to make twice:
KOOCH_PAGE_MARKING=1 kooch_editor # every pixel
KOOCH_PAGE_MARKING=1 KOOCH_PAGE_MARKING_RATE=4 kooch_editor
🔴 In the Performance panel and not in .rendersettings, deliberately.
#477 is explicit that nothing on the shadow side should grow a public
setting — one written into the project and therefore promised to every
project — before the pool’s shape is decided. This is a diagnostic the
editor drives, so it lives where the editor’s other diagnostics do.
🔴 The count is for EVERY light the grid holds, not the handful with a shadow slot today — and that is the measurement rather than an oversight. A virtual shadow map exists for many lights; the Chalmers paper is titled “Efficient Virtual Shadow Maps for Many Lights”. Counting only the four that fit today’s cube slots would be measuring the cap the feature is meant to remove.
⚠️ The resolution is part of the reading, not context around it. The editor renders two views at two sizes, so the panel shows two different numbers a frame apart, and a page count without its resolution is not a number — this project has already had to retract a table that mixed 1080p with 720p.
🎯 The cross-check that the light side is right: pairs divided by
samples is the grid’s own lights-per-pixel. Measured in the editor on
many_lights.scene, 993 608 pairs over 51 180 samples is 19.4 lights
per sample, against the ~20 per cell the froxel grid reports for
itself.
It also logs shadow pages marked with the same numbers whenever the
count changes — on change and not per frame, for the same reason the
point-shadow warning is a flag rather than a count.
🔴 The depth says WHERE a surface is; the froxel grid says WHICH lights reach it, and neither is sufficient. Marking from the grid’s cells alone claims pages for ground no surface occupies — 133x, measured above. Marking from depth alone would walk every light per pixel, which is the loop the grid exists to remove. Epic states the first half and the Chalmers papers the second.
⚠️ KOOCH_PAGE_MARKING_RATE is not free accuracy in either direction. A
coarser rate is fewer threads and a wider pixel footprint, so the
level chosen comes out coarser and the count lower. 1 is the honest
reading and the expensive one.
Seeing it — Paint pages over the scene
The count says how many; the view says where. It colours every pixel by the shadow page it reads:
- Hue is the level — where the frame spends detail. A band of colour is a level boundary.
- Brightness is the page identity, hashed, so neighbouring pages differ and the tiling is visible. A page covering a quarter of the screen is a page too coarse for it; a mosaic too fine to resolve is detail nobody sees.
The sun’s page wins where there is a sun. A pixel is lit by many lights, and painting the last one walked would make the view depend on the light list’s order.
🔴 Painting forces the sampling rate to 1. At any coarser rate the view is a grid of dots over an unpainted frame, which reads as a broken pass rather than as a coarse sample.
🔴 It paints the view’s FINAL colour target, not the radiance one,
and that is two problems solved at once. The radiance target lives
inside the R64 stage where this pass cannot reach it; the final one is
Rgba8Unorm at the view’s output size and already tonemapped, so
the palette needs no exposure divided out of it and nothing downstream
can overwrite it.
⚠️ The depth buffer is at the RENDER size and the target at the OUTPUT
size, and they differ whenever render_scale is below 100. One thread
then owns a block of output pixels rather than one, and it fills the
whole block — writing a single pixel would leave a grid of dots over an
unpainted frame.
🔴 It is an instrument, not a feature. Nothing reads what it writes,
and it is off unless asked for — a measurement that runs whether or not
anyone wanted it is a cost nobody attributed. Its job is to disagree
with the census: every arithmetic decision in the shader has a twin in
pages.rs, and if the two counts diverge, one of them is wrong.
⚠️ They are not expected to match exactly. The census marks per froxel cell and the pass marks per pixel, so the pass is the finer instrument and the census the cheaper one. What would be a finding is a divergence too large to explain by that — an order of magnitude, or a count that moves the wrong way when the scene changes.
How fine a page has to be — shadow_density (#929)
The census found exactly one knob that moves the bill, and it is not page size or virtual size. It is how many shadow texels a screen pixel is allowed to ask for.
The level chosen for a page is a log2, so the setting is a power of two
or it lies about what it did: a coarser texel is a level coarser in
both axes, and half the density is a quarter of the pages.
shadow_density | Texel per pixel | Pages |
|---|---|---|
| 100 | one | the measurement |
| 50 | half | a quarter |
| 25 | a quarter | a sixteenth |
⚠️ Below 100 a shadow’s edge is softer than the surface it falls on, which reads as blur rather than as a lower setting. 50 is where it starts to show.
The page pool and its table (#866)
Marking answers which pages this frame needs. The pool answers where each of them lives — and the two run in the same dispatch.
The allocation is free because marking already did the hard part
mark_bit returns whether the calling thread is the one that flipped a
page’s bit from 0 to 1. That is a unique thread per page, established
by an atomicOr that had to happen anyway. Claiming a physical slot
there is one atomicAdd on a rare branch: no second pass, and nothing
walks the virtual space.
The alternative — sweep the mark bitmap afterwards and allocate what is set — is a dispatch over the virtual space. See the next section for how large that is.
🔴 The table is FLAT, and the number that used to forbid it is dead
The lookup runs per pixel per light in the shading pass, and prior
art is unanimous that it must be one indexed read: Chalmers (“quite
fast because they only require a single texture lookup”), Stephano’s
sparse VSM (pageTable[ivec2(floor(uv * numPagesXY))]), UE 5.8
(CalcPageOffset is flat arithmetic over 21 845 entries per map). The
first table here hashed instead — open addressing with tombstones —
and the measurement that killed it, on many_lights at 1096 frames:
shading 10.4 ms against 0.884 ms for the entire shadow track, on a
walk of up to 5 chain levels × up to 32 probes, per pixel per light.
The hash had existed for a real reason. With 128-texel pages over a
16384 virtual map, a mip chain per cube face and a 17-level clipmap,
one light addressed 278 528 pages; a hundred lights and a sun,
28 409 856 — a flat u32 table was 108 MiB, 42 % of the pool it
would index, describing pages that are 99.99 % empty. Two decisions
shrank the space by a factor of ~58 and made flat affordable:
LOCAL_MAX_TEXELScaps a lamp’s chain three levels below the sun’s — a factor of 64 in the pages one lamp can address, and the texel it gives up at four metres is two millimetres.- The address space stops paying for the capped levels. A lamp’s
chain is addressed from
local_level_floorup, so its stride is 2 046 pages instead of 131 070, and the sun’s clipmap sits at the tail of the view’s span. 101 lights and a sun now address ~485 000 pages — a few MiB of table atPAGE_CELLwords per entry.
This is Epic’s own shape: UE5 stays flat by never handing a distant
light a full virtual space (VSM_MAX_SINGLE_PAGE_SHADOW_MAPS is 8192
maps of one entry each).
The entry index IS the page id
- The first word is
slot + 1, soPAGE_ABSENTis 0 and an empty table is a zeroed buffer. Eviction stores 0 — no tombstones, because nothing probes past an entry any more, and the sweep pass that kept the hash’s holes in check is deleted outright. - The insert is a plain store, not a compare-exchange: only the thread that flipped a page’s mark bit inserts, and marking already guaranteed there is exactly one.
- Light slots are padded (
padded_lights, steps of 64) so adding a light does not shift the sun’s region or the next view’s base — the layout, and every resident page with it, survives scene edits until the count crosses a step. - The reader is one load per level tried. The sun’s walk starts at its containment level and typically resolves on the first; a lamp’s starts at the floor and has at most five levels to try, each a single indexed load where the hash paid a probe run.
page_table.wgsl holds the id arithmetic and the atlas layout, and is
concatenated into every pass that touches the table, so the writer and
the reader cannot drift apart.
Overflow has a name here because it has none on screen
PoolCounts::overflow counts pages the frame needed and the pool could
not seat. They render unshadowed. Epic’s own pool overflow shows up as
checkerboard corruption or missing shadows — a failure nobody recognises
by sight — so the panel names it instead.
The panel also cross-checks claims against resident. Both count the
same 0→1 transitions by two different mechanisms, and a disagreement
means one of them is broken.
The camera is part of the key, and the pool is sliced
One editor frame draws the same world from two cameras. A clipmap is centred on its camera, so the same world position is a different page in each — and a table keyed without the camera hands the second one the pages the first marked. The measured symptom was exact: shadows in one viewport and none in the other.
So a page id carries the camera above everything else:
page = view * view_span + light * stride + <chain offset>
which is UE5’s VirtualShadowMapId written as a multiply instead of a
table per id — and with a flat table the multiply is the address.
Three things follow, and none of them is optional:
- The table is aged by a pass, not wiped by
clear_buffer.age_viewwalks only this camera’s contiguous run of entries and evicts what went unrequested pastmax_age. It has to be per camera because the raster is fused with the shading — a camera samples an atlas a frame old, so wiping the whole table at the top of a frame leaves whichever camera marks second reading what the first just erased. - The pool is sliced, not shared, and the atlas is an array with a
layer per camera. A layer is an attachment a camera clears on its
own; the alternatives — a scissor, a stencil, a clearing draw — all
partition one surface and all of them are a rule somebody has to keep.
The budget does not multiply: a layer is
pages / viewsrounded up to a square, so two viewports cost what one did. - The uniform has a slice per camera.
Queue::write_bufferis not ordered against the encoder, so writing one range twice in a frame hands both passes the second value — the engine shipped that bug once already. A camera writing its own range cannot be overwritten.
The seating plan — who gets a slot under pressure (#942)
Both of the gaps this section used to list are closed: pages persist
across frames (see Cached pages are effectively free below), and
allocation is no longer first-come. What replaced first-come is worth
stating precisely, because the failure it ended was measured: 7 674
pages wanted against a 1 024-page slice, a steady state of 1022 reused · 0 new · 0 evicted, and 6 652 requests starving forever —
whoever claimed a slot first kept it, because a resident page is
re-marked every frame and never ages. Moving the camera produced
requests that never landed; the shadows visibly lagged the view.
Three dispatches at the tail of the marking pass, in an order that IS the algorithm:
| Pass | Threads | What it decides |
|---|---|---|
plan_view | one | prefix-sums the frame’s demand histogram against the slice’s budget: the cutoff rank, the quota within it, and the spare the cache may keep |
preempt_view | one per table entry | evicts every resident the plan did not fund — and lets residents of the cutoff rank take the quota before any newcomer, so a page with content beats one without |
adopt_view | one per table entry | seats every marked page the plan funded; what is past the cutoff is denied and counted |
The rank is the level, coarsest first: the sun’s clipmap (ranks
0..17) ahead of every local light, and within any chain the coarse
levels ahead of the fine — so under pressure a consumer loses its
finest detail before it loses coverage, and the sun (the one consumer
every frame has) never loses to a lamp. The demand histogram is
recorded where the mark bit is won, which makes the plan one 32-bucket
loop rather than a sort.
Two consequences the counters pin:
- The arithmetic closes. The plan funds exactly
sliceseats and the preemption frees everything unfunded, sopage_allocnever misses:overflowstays zero and every shortfall is a denial with a rank on it — the panel says what was sacrificed, not just how much. - A move reseats in one frame. Stale pages stop being marked, the
plan stops funding them, and
preempt_viewhands their seats to the new view the same frame — notmax_ageframes later.
What #942 does not do is make 7 674 fit into 1 024 — that is the next section’s job.
The resolution bias — making the demand fit (#943)
When the plan reports pressure, a persistent per-view bias walks the marking coarser, one level per frame: the LOCAL lights pay first (up to four levels), the sun only when they have nothing left to give (up to two). Each level quarters that party’s page demand, and the readers need no change at all — both walk their chains from the fine end and take the first resident page, so a coarser marking is simply what they find. UE5 runs the same loop as its page-pool-overflow bias; Olsson caps every light by projected area (Eq. 1) for the same reason: when demand cannot fit, serve everyone coarser rather than turn 87 % of the requests away.
Unwinding is two-tracked, and the asymmetry is the hysteresis. Where the arithmetic can prove a finer marking fits (slack ≥ 3× the party’s demand), the bias steps down immediately. Where it cannot — coarse clipmap levels do not quadruple, so the ×4 estimate over-blocks — it tries a step once 16 quiet frames of patience run out, and the ordinary raise reverts a failed trial the next frame. The still-resident coarser pages catch the readers meanwhile, so a failed trial costs one frame of fallback, not one of missing shadow. The panel prints the standing bias; one that sits high is the pool saying it is too small for the scene.
The classic pass under the pages — a token, not a tenant (#945)
With virtual_shadows on, every reader branches to the pages: the
cascade draws are gated, the spot and cube lists are empty. What
remained was the memory — 64 MiB of atlas and 6 per cube, standing
for a reader that never comes. They cannot go to zero (the shading’s
bind group needs live views, and wgpu refuses a zero-layer texture), so
classic_shadow_alloc — pure, tested — sizes them to a token: the
atlas at its clamp floor, one sixteen-texel cube, under half a
megabyte. The release key is the whole allocation tuple rather than a
bare texel count, so an author whose cascades already sat at the floor
still swaps on toggle; the resize-release door that already existed
does the swap in both directions.
The coverage gate — a shadow nobody can resolve claims nothing (#944)
Before the bias has to price anything, the demand is shrunk at the
source: a local light whose whole range projects under
shadow_min_pixels of radius on screen (8 by default, 0 turns it off)
marks no pages at all. It still shades — the readers walk its chain,
find nothing resident and return lit — and it gets its shadow back the
frame the camera comes close enough to flip the comparison. The gate
errs toward casting: it measures the range sphere, not the lit part of
it, and the sun is never gated because it has no radius. Epic runs the
same rule as a pass (PruneLightGridCS) before anything marks; here it
is one comparison inside a loop that already holds every operand. The
panel counts what it turns away on the census line.
Rasterising into the pages — the depth raster (#866)
Four passes, and their shape is the feature.
| Pass | Threads | What it produces |
|---|---|---|
| Cull | per clipmap level for the sun; ONE hierarchical set of dispatches for every lamp (#939) | which meshlets survive — at that level’s texel density for the sun, at a perspective error metric from each light’s own position for lamps |
| Compact | one per table entry | the resident pages, dense and bucketed by level — the sun’s clipmap levels first, then one bucket per lamp slot |
| Expand | pages × survivors, dispatched indirectly | (page, meshlet) pairs |
| Draw | one draw_indirect over every pair | depth in the atlas |
One hierarchical cull for every lamp (#939)
A lamp does not borrow the sun’s survivor lists. Those are LODs picked for orthographic boxes centred on the camera: borrowed, a close lamp’s casters fell outside the fine levels’ box and its shadow vanished as the light approached, while a coarse bucket handed root meshlets and drew a sphere’s shadow as a faceted lump.
What replaced the borrowing is Olsson et al. 2014 (§3.4/§5.2) adapted to the meshlet pool — four dispatches shared by all lamps, once per frame (a lamp’s cull is view-independent, so the editor’s second camera reuses the first one’s survivors):
- Pairs — light sphere against instance bounds sphere, over
lights × instances. The hierarchy: instances a light cannot reach never enter the meshlet domain. - Args — sizes the meshlet-domain dispatches from the GPU-side pair count.
- Error — the group-coherent LOD reduction (#465), every lamp at
once: the arena is indexed
[slot × group_capacity + group], so sibling meshlets of one lamp still converge one slot and casters never tear at LOD seams. - Cull — group-coherent cut + range + backface cone, perspective
error measured from the light’s position. Survivors land in fixed
per-lamp slices of one shared arena (
LAMP_SURVIVORSeach, counts written uncapped so overflow is a number, not a silence), and the counts land directly in the raster’svisible_counts— no copy, no per-lamp bind group, no CPU loop.
One survivor list serves all six faces and every chain level of a lamp
because a perspective error metric already scales with distance. The cap
is LAMP_CULLS = 64; its honest ceiling is the group-error arena,
LAMP_CULLS × group_capacity × 4 B. Slots are buffer order — ranking
casting lights (the classic path’s assign_point_slots) is #939’s named
follow-up.
The page filter is a configurable box (#941)
The readers filter by hand — a page’s neighbour texel can belong to
another level or another light, so no hardware sampler applies — and
the footprint is the author’s: shadow_softness in RenderSettings
is a box width in shadow texels. 1 is comparison-bilinear, bit for
bit the cube path’s hardware look; wider widths are the Castano-class
box with frac-clipped 1D edge weights, positioned with sub-texel
precision, costing (width + 1)² loads per light per pixel — which
is why sharp is the default and the widest choice is labelled with its
bill. No blocker search: the penumbra is uniform, not
contact-hardening.
The receiver bound — casters behind everything draw nothing (#940)
Olsson §4’s PMCD variant, at page granularity instead of the paper’s
per-face bound. Every sample that marks a lamp’s page is a receiver,
and the marking atomicMaxes its radial distance from the light into
the page’s fifth table word (positive floats bitcast to ordered u32s —
that is what lets an atomic hold a distance). The compaction carries
the bound into the widened page_list — the expansion sits at the
eight-storage-buffer limit and cannot bind the table — and the lamp
expansion adds one rejection: a caster whose nearest point lies
beyond the page’s furthest receiver occludes nothing the frame shades.
The bound is radial rather than per-face depth, so it errs toward
keeping; zero means “no receiver recorded” and rejects nothing, which
is what keeps planted rigs and the sun’s slab test untouched. age_view
zeroes it each frame — receivers are a frame’s question. The panel’s
pair-tests line counts what it turns away, which is the number that
says whether the fifth word earns its memory.
Cached pages are effectively free (#477/#866)
The pool always persisted its slots; since this change it keeps the content. Every table entry carries a content stamp — the generation its atlas depth was drawn under — and the compaction skips any resident page whose stamp still matches: not listed, not expanded, not drawn. The whole-layer depth clear is gone; the pass loads the layer and wipes only the dirty pages’ rects with one quad each.
Three things turn a generation over, and nothing else redraws a page:
| Source | Granularity |
|---|---|
| The sun’s snapped centre stepping, its direction, or the eye moving along its axis (the depth origin rides the eye) | per clipmap level |
| A lamp’s position, direction, range, kind or cone changing | per lamp — UE5 invalidates the same way |
A caster moving — its old and new bounds arrive as spheres and cs_invalidate zeroes the stamps of every page they reach | per page for the sun; per lamp for local lights (the #866 refinement is per cell) |
A moved-caster list past its buffer, or a pair-list overflow observed
by the panel’s readback, bumps a scene generation folded into every
hash: everything redraws once, which is coarse and never stale. The
panel prints rastered · cached · side by side — UE5’s rule of thumb
is dirty under 5% of residents in a typical frame.
One render pass for the whole clipmap
The atlas is a single depth attachment and every page is a sub-rect of
it. page_clip places a page’s own clip space inside its rect, so 1681
pages are one begin_render_pass and one draw_indirect rather
than 1681 of each. The hardware depth test does winner-takes-all exactly
as it does for a cascade.
🔴 That is why this pipeline has a fragment shader where
shadow_depth.wgsl has none. A triangle wider than its page keeps
rasterising past the rect and into the neighbouring page, which belongs
to another level — a caster would appear in a shadow map it was never
meant to be in, at the wrong scale, and nothing about the result would
say why. A scissor would fix it and cannot: scissor is pass state and the
page changes per instance. So the fragment shader discards outside the
rect, and the cost is early-Z. That is the price of one pass instead of
one per page.
The pair list is the whole trick
A shadow page is a 128-texel view of the world and a scene has thousands of meshlets. Rasterising every meshlet into every page is the cost virtual shadow maps exist to avoid; rasterising a meshlet once, into the pages it actually touches, is what makes 1681 pages affordable. The expansion is where actually touches is decided, and it is one sphere against one box.
The pair carries the cull’s own packed (instance << 16 | meshlet), so
it is self-describing — the draw never learns which level produced it,
which is what lets every level share one indirect draw.
🔴 The sun only, and the seam is not arbitrary
A cull is per view, and a view is where the LOD is chosen. The sun’s clipmap is 17 views. A hundred local lights with six faces and an eight-level chain each are 4848, and the LOD selector is a two-pass reduction over the meshlet DAG that cannot simply be inlined per page.
Local pages are marked and allocated today, counted as local in the
panel, and not drawn. Drawing them needs the cull itself moved onto the
GPU as one multi-view dispatch. That is the next machine, not a bigger
version of this one — and reporting the count is what keeps a pool that
looks full from being read as a pool that is full for the reason someone
assumed.
What it costs
The pool defaults to 2048 pages = 128 MiB, half of Epic’s 4096, and
the comparison that decided it is this engine’s own: 152 MiB of fixed
shadow allocations stand today for four casting lights. The pool is
less than that, adapts to the frame, and caps by memory rather than by a
count of slots. KOOCH_SHADOW_POOL_PAGES raises it.
⚠️ Nothing is cached across frames. The table is emptied every frame, so a static shadow is re-rasterised every frame — and caching is the optimisation virtual shadow maps exist for. Epic measures a cached local shadow map at 0.05 ms against 0.4–0.8 ms invalidated. It needs an eviction policy and an invalidation rule, and neither can be designed before anything renders.
⚠️ Nothing samples the atlas yet. That is #477, and it is what turns this into shadows on screen.
Sampling a page — where the shadow finally appears (#477)
inti_shadow takes one branch: if the page uniform’s sun flag is set,
the sun’s shadow comes from the pool and the cascades are not consulted
at all.
🔴 It replaces the cascades rather than blending with them. Two techniques over one surface disagree at their own boundaries, and the disagreement reads as a seam that belongs to neither.
It walks levels instead of recomputing which one was marked
The marking pass chose a level from the screen’s pixel density, which needs the camera’s focal length and the render size. Reproducing that arithmetic in the shading pass would be a third copy of it, free to drift by a rounding step — and a level off by one is a lookup that misses, which reads as a shadow that disappears rather than as a shadow at the wrong scale.
So the reader starts at the coarsest level that could contain the point and walks outward, taking the first resident page. Any resident page containing the point holds correct depth, whatever level marked it: the stored value is a distance along the sun’s axis and does not depend on how finely the page was diced. Typically the first level tried hits, and each try is one indexed load.
That walk is also what absorbs the frame of latency below.
⚠️ The pages sampled this frame were filled the previous one
The raster and the shading here are one fused fragment shader, so there is no depth buffer to mark from until shading is over. Marking and the page raster therefore run at the end of the frame, and what the next frame samples is what this one left.
The failure mode is the right one: a page that appears suddenly — a fast
camera turn, an object entering frame — is lit for one frame rather
than wrong. inti_page_shadow returns lit when no page in the chain is
resident, because a point nobody marked is a point the frame never
looked at, and guessing dark there would put shadow where no data exists.
The filter is clamped inside the page, and that is not optional
Taps are textureLoad, never a sampler. A hardware filter has no way to
be told where a page ends, and the texels past that border belong to
another clipmap level: not a softer edge, a shadow from somewhere
else. The 2×2 taps are clamped to the page’s own rect.
That is also why the atlas is Depth32Float read as a plain texture
rather than through the comparison sampler the cascades use.
🔴 A setting only exists if the frame can reach it
virtual_shadows shipped inert, and the way it was found is worth
more than the fix.
The frame read the author’s asset — resources.get::<RenderSettings>()
— and took a fallback when it was absent. RenderSettings is never
inserted as a Resources value: apply publishes derived structs like
ShadowSettings instead. So the lookup returned None in every build,
the fallback turned the pages off, and the environment override sat
behind an early return that fired first.
It compiled, it ran, it logged nothing, and it never rendered a page.
What caught it was a profile: two handheld captures, one with the pages
forced on and one without, came back identical scope for scope —
Device::create_bind_group at 55.0 calls a frame in both. Seventeen
clipmap culls cannot cost nothing.
Two tests stand where it was. One checks that the field arrives; the
other, the_frame_never_asks_for_render_settings, walks the crate’s
source and fails on any get::<RenderSettings> outside the settings
module — the bug’s class, because the behavioural test only covers
the field somebody remembered.
Where the settings live
Everything that was a panel diagnostic in #866 is now a project setting,
in a Shadows: virtual pages group beside Shadows: sun cascades and
Shadows: contact — one run per technique, adjacent, so what belongs to
which is never in doubt.
| Setting | What it decides |
|---|---|
virtual_shadows | pages instead of cascades. Off by default — every scene in the project was authored against the cascades |
shadow_density | texels per screen pixel; the page count falls with its square |
shadow_pool_pages | the memory budget, 1024–6144 pages |
virtual_shadow_debug | paints the page each pixel reads |
🔴 The marking rate is gone. While marking was an instrument, a coarser rate traded accuracy for threads. It decides which pages exist now, so one sample in sixteen is fifteen pixels whose shadow was never rasterised. It is pinned at one per pixel.
KOOCH_PAGE_MARKING=1 survives as a force on top of the setting, not
as its default — the comparison it exists for is made on a handheld, over
SSH, against a build nobody wants to make twice.
Marking per cluster, and the second-order effect that decided it
Olsson §III names per-sample marking as the branch to replace: a page is
a property of the cluster a sample falls in, so marking it once per
cluster rather than once per sample is the same answer for a fraction of
the pairs. It shipped behind KOOCH_CLUSTER_MARKING because this pass
chooses which pages exist, and a wrong answer here is a missing
shadow rather than a slow frame — a defect nothing logs.
Measured on the OneXFly, many_lights (100 point lights), same camera,
64 °C against 66 °C:
| per pixel | per cluster | ||
|---|---|---|---|
page mark | 19.674 ms | 2.729 ms | −7.2× |
page depth | 38.639 ms | 27.862 ms | −28 % |
shadow pages | 59.352 ms | 31.622 ms | −1.9× |
| frame, median | 91.01 ms | 55.13 ms | 11.0 → 18.1 FPS |
🔴 The 7.2× is what the paper predicts; the 28% is what settled the default. Nothing touched the rasteriser, and it got a quarter cheaper anyway — because marking per cluster does not merely cost less, it asks for fewer pages, and a page never asked for evicts nobody and rasterises never. A pass that is 4 % of the frame cannot buy that on its own; it bought it by changing what the next pass was handed.
So the switch turned around: cluster marking is the default and
KOOCH_CLUSTER_MARKING=0 returns the per-pixel path. The escape hatch
stays for the reason it was built — reach for it when a shadow is
absent and the cause is not obvious.
⚠️ What this did not fix. page depth is still 27.9 ms, now 67 %
of the GPU frame, because the pool is still over-subscribed: the
census puts many_lights at 6916 resident pages against a 2048-page
pool, so the pool is evicted and refilled every frame and the content
cache saves nothing. That is a budget defect, not a marking one, and the
same census says where it lives — the sun wants 118 pages and saves
133.6× over a brute-force allocation, while the 100 lamps want 6798 and
save 1.2×. Virtual paging pays for coherence, and a hundred
scattered point lights have none to sell.
What Inti does not do yet
- Nothing but lights is clustered. The grid reserves a range per cell for reflection probes, irradiance volumes and decals, and none of those exist yet. The ranges are empty and cost nothing to walk.
- The index list grows a frame late. How long it needs to be is a property of how the scene is lit, which only the GPU knows, so it comes back asynchronously. A frame that overflows renders its later cells under-lit rather than reading past the end of the buffer, and the next frame has the bigger buffer.
- Four spot maps and four point cubes, both memory limits rather than limits of the technique. Past them a light keeps lighting and stops casting, ranked by distance to the camera.
- No environment map, no IBL, no area lights, no volumetrics, no bloom. The crate’s original doc comment promised the last three. It now promises what it has.
- No global illumination (#450), which is why the punctual default is forty times a real bulb. That number goes back to physics the day GI lands, and not before.
Frame Pacing
A game loop is supposed to spin. An editor showing a still image is not.
Until #656 the engine made no distinction: the winit handler asked for the next redraw at the end of every frame, unconditionally, so the loop fed itself forever. Vsync capped it at the refresh rate, which is the only reason it cost one core per process rather than all of them. Idle, with a project open and nothing happening, that measured at two pinned cores and 51.8 W on a 9800X3D — to display an image that was not changing.
The contract
Two types in kooch_core::frame_pacing:
FrameRequest— what this frame decided the next one needs. Systems raise it; the runner reads it once per frame and resets it to a baseline. Raising is monotonic within a frame: the most urgent request wins, so draw order can never talk a system out of a repaint it asked for.FrameWaker— a clonable handle any thread can use to break the loop out of a sleep. The wake is sticky, so one that lands between the end of a frame and the moment the runner commits to sleeping is not lost.
Three paces, in order of urgency:
| Pace | Means | ControlFlow |
|---|---|---|
Continuous | Something is animating or simulating | Poll + request_redraw |
After(d) | Something is on a timer | WaitUntil(now + d) |
Wait | Nothing to draw | Wait |
An app that inserts no FrameRequest keeps spinning. That is not an
oversight — it is what a shipped game wants, and it means the opt-in is
explicit at every call site that needs it.
Who asks for what
The editor (FrameRequest::new(FramePace::Wait)) takes its answer
from egui, which already computes one: run_ui returns a
repaint_delay per viewport, ZERO while something animates and
Duration::MAX when the UI has drawn everything it has. Two things egui
cannot see are folded in on top:
- Play — the viewport texture changes from another process, with no
widget to notice it through.
Continuousfor as long as Play lasts. - A live remote session — the project’s stdout arrives on a socket,
not as a window event, so a fully asleep editor would hold its Console
output until the user happened to move the mouse.
After(250 ms).
A frame that failed to present asks for another unconditionally: what is on screen is not what that frame drew.
A project under an editor (RemotePlugin) sleeps by default and is
woken by its own socket. Between edits nothing simulates, so a frame
nobody asked for is a core spent mirroring a still scene; Playing
raises the pace for as long as Play lasts. The listener thread parks on
a reply only the main thread can produce, which is why the wake is not
optional — without it, an editor asking a perfectly healthy project a
question would hang until something unrelated produced a frame.
Frames stopped being a clock
Anything that said “every N frames” was reading a clock that no longer ticks at a fixed rate. An idle editor draws roughly four frames a second, so “every thirtieth frame” went from half a second to seven and a half.
The remote snapshot pull was the one such cadence in the tree, and it is
now expressed as a Duration. Any new cadence should be too — frame
counts were always a stand-in for time, and they are no longer even a
good one.
Input never waits
While the loop is idle, an input event is the only thing that will
produce a frame, so every window event other than RedrawRequested asks
for one. A WaitUntil deadline expiring reports through
StartCause::ResumeTimeReached, and a cross-thread wake arrives as a
winit user event — the proxy rather than request_redraw, because the
proxy is the API documented to be callable from another thread.
The swapchain image is asked for last
The game runtime’s frame is two halves that need very different things. The meshlet stage — cull, raster, shading, shadows, TAA, tonemap — draws into textures the engine owns and submits its own command buffer. Only the sky and the blit write to the surface.
So the surface image is acquired between them, and not before both:
record + submit the scene ──► get_current_texture() ──► sky, blit, present
~34 ms of GPU work blocks on the compositor ~0.65 ms
🔴 Acquiring first costs a full frame of overlap, and that is what
this did until #837. get_current_texture blocks until the presentation
engine releases an image, so asking for it before recording puts the
whole CPU-side of the frame after the wait — and the GPU cannot start
this frame’s work until the compositor has let go of the last one.
Measured on the OneXFly: a median frame of 37.14 ms made of 34 ms of GPU
and 3.006 ms of recording, added together rather than overlapped.
Nothing about the image is needed to record the scene. The dependency was in the control flow, not in the data.
⚠️ The editor already did this correctly, which is why the two paths
look different: systems/present.rs tessellates the UI, uploads its
textures and updates its buffers before acquiring, and the viewport
passes run earlier still. Only the game runtime had the acquire on top.
Vsync is a setting, not only a variable
vsync lives in .rendersettings, beside the shading path and the
render scale, and Presentation carries it to the surface through the
same environment first, asset second rule the rest of quality uses:
KOOCH_PRESENT_MODE=novsync overrides the file for one run, vsync
overrides it back, and unset leaves the project’s choice alone.
🔴 It was an environment variable and nothing else until then, which meant a game built with Kóoch could not offer a vsync toggle — the one graphics option every game ships. Its own options menu would have had to ask the player to set an environment variable.
⚠️ vsync: false in a project file is a measurement configuration, not
a performance setting. It costs a GPU drawing frames nobody sees, and on
a handheld it costs battery for the same. It is off in exactly one case:
somebody is reading a frame time and needs it to show work rather than
the wait for the vblank.
GpuContext::set_vsync compares against the mode the surface already
has before it reconfigures, because configure rebuilds the swapchain —
applying the resource unconditionally would rebuild it once a frame.
There is no headless surface, so no test in the suite watches that
happen; what the tests pin is the precedence rule and the serde default,
which is what every .rendersettings already on disk silently becomes.
What this does not fix
The frame-time distribution is bimodal on the OneXFly — the same GPU work produces a 34.7 ms frame and a 69.4 ms one — and this change does not address that.
Two explanations were on the table and both are now refuted, by three 30-second captures of one binary with one variable changed each:
| latency 2 | latency 3 | novsync | |
|---|---|---|---|
| frame/GPU p80 | 1.98 | 1.99 | 1.94 |
| frame/GPU p90 | 2.49 | 2.50 | 2.21 |
vkAcquireNextImageKHR ms/frame | 35.209 | 33.646 | 37.162 |
A third swapchain image does not move the ratio by a hundredth. Neither does leaving FIFO.
🔴 An acquire of ~35 ms against a GPU of ~35 ms is not a defect.
Being GPU-bound means the CPU waits somewhere, and get_current_texture
is where. Reading that number as a symptom is a mistake this document
used to make. What is genuinely unexplained is only the tail: the
frames where the wait grows by 50 ms while our own GPU work grows by 2.
⚠️ A present mode is close to decorative when a compositor owns the
display. These captures run under gamescope, which composites on the
same GPU, on its own schedule, and is invisible to our scopes — they
time our passes and nothing else. novsync turns off our vsync, not
its. Whatever is left lives outside this process, and no environment
variable on this side is going to find it.
The next measurement is not another engine knob. It is gamescope’s own frame statistics, or a run without gamescope at all.
And what the tail was hiding
GPU: ~35 ms
budget: 13.9 ms
2.5x over, with a still camera and at half shading rate. If the tail
vanished entirely the frame would still miss by more than double, and
shade: compute (half rate) alone — 19.7 to 22.8 ms across the three
captures — costs more than the whole frame is allowed.
The tail is a mystery in 30% of frames. The shading is 60% of every one of them.
Profiling
Where a frame actually goes (#785).
It is not ours
The flamegraph, the timeline, the frame history, the scope statistics and
the file format are all puffin,
drawn by puffin_egui into the editor’s own egui. GPU-side timing is
wgpu-profiler. The project’s
dependency policy is to never implement what a maintained crate already
covers, and profilers are a solved problem.
What is ours is the panel’s capture controls, the scopes, and the decision about what a build may contain.
The scope macro is profiling::scope!, never puffin:: directly
profiling is a facade: one macro
at the call site, and the backend is a cargo feature (puffin, tracy,
optick, superluminal). wgpu-profiler is built on the same facade.
⚠️ With no backend feature enabled, profiling::scope! expands to
nothing. That is what lets a shipped game contain no instrumentation at
all rather than instrumentation that is switched off — see #558 on what a
release build may carry.
| Build | Profiler | Channel |
|---|---|---|
Editor, --features profiling | compiled in | in-panel |
Game, --features profiling | compiled in | puffin_http over TCP |
| Release | absent at compile time | none |
A happy accident of the facade: wgpu and egui already use
profiling, so their internal scopes land in the same flamegraph
without a line of work.
Recording is off until you ask
Opening the panel records nothing. Puffin’s own default for
are_scopes_on is false and it is left that way; the Record button
is the only thing that turns it on.
🔴 There are two transport-looking controls in the panel and they do different things:
| Control | What it does | Cost while it is like that |
|---|---|---|
⏺ Record / ⏹ Stop | starts and stops recording | stopped: one atomic load per scope |
▶ / ⏸ (puffin’s own) | freezes the view on one frame | recording continues underneath |
The flamegraph is not drawn while recording. Drawing it was measured at 10.97 ms of a 15.98 ms frame — the instrument was two thirds of the measurement. Record, do the thing, stop, then read.
Captures
Save capture writes a .puffin to ~/.local/share/kooch/captures/ —
outside the project, because a capture is a measurement of a session on a
machine and not an asset of the game. Load capture reads the newest one
back into the panel, so a capture is readable without installing the
standalone puffin_viewer.
cargo run -p kooch_editor_core --features profiling --example read_capture -- <file>
prints the same ranking in a terminal, which is what a capture pulled off
the handheld will need. Two flags for the two questions an average
cannot answer:
| flag | what it prints |
|---|---|
| (none) | the ranking, and the frame’s GPU tree with self times |
--slowest | the worst single frame, whole. A mean hides a stall by definition: one 700 ms frame in a thousand moves it by 0.7 ms. |
--split | the fastest quarter of frames against the slowest, by self time, plus each frame against the GPU work that produced it. |
--over-time | the capture in ten chronological buckets, and a warning when the last drifts more than 10 % from the first. |
🔴 --split is the one to reach for on a capture where the camera
moves, because that capture holds two populations and a mean over both
describes neither. The frame/GPU ratio it prints is what separates the
GPU is the wall from we are waiting for the GPU twice — see the second
handheld capture below.
🔴 --over-time is the one to reach for on a handheld, and it exists
because of what the third capture below found: the same binary, the same
scene and a camera that never moves gets 46 % slower two minutes in.
A median over that capture describes a machine that was never running.
🔴 A capture can be silently unreadable
A FrameView turns the scope ids inside a frame into names through a
collection it builds as frames arrive. Anything that replaces the view
— Clear history, Load capture — starts that collection empty, and the
frames that follow come back as scope#ScopeId(67). The capture does not
look broken. It is simply useless.
GlobalProfiler::emit_scope_snapshot() asks for every known scope to be
re-sent, and it is called for the first two seconds of every recording.
Not once, and not for two frames, because:
#![allow(unused)]
fn main() {
let propagate_full_delta = std::mem::take(&mut self.propagate_all_scope_details);
...
Err(Error::Empty) => return, // the frame had no scopes
}
puffin takes the flag before building the frame and returns early if that frame came out empty — carrying the request away. The frame that closes right after recording starts is exactly that empty frame. And a scope only registers the first time it runs, so anything occasional registers after a single snapshot has already gone out.
And a second cause, which is the one that kept biting
The snapshot above fixes receiving the names. Saving them is a
separate problem, and it is the one that produced
scope#ScopeId(67) in capture after capture:
#![allow(unused)]
fn main() {
// FrameView::write
write.write_all(b"PUF0")?;
for frame in self.all_uniq() { frame.write_into(None, write)?; }
}
FrameView::write does not serialise the scope collection. The names
reach a .puffin only by riding inside a frame that all_uniq still
yields — and a FrameView retains a frame through two independent
nets: the last max_recent = 1000, and the slowest max_slow = 256. The
frame carrying the names survives if it is in either.
🔴 Which makes the failure intermittent, not a threshold, and getting that wrong costs a session. The first test written for it asserted the names were lost past 1000 frames and passed while proving the opposite: a synthetic capture’s first frame is also its slowest, so the second net kept it. Whether a real capture comes back readable depends on whether its first frame happened to be slow — which is why 846 frames on 2026-08-16 were named and 1022 the same evening were not.
ProfilerPanel::keep_all_frames sets max_recent to usize::MAX on any
view that intends to save. The cost is RAM on the capturing machine,
which is a desktop, and puffin packs older frames as they age. A bigger
arbitrary number would just move the same silent cliff somewhere less
obvious.
⚠️ read_capture now says so out loud: a capture whose scopes are all
scope#… prints a warning naming this cause, instead of printing a
ranking of anonymous ids that looks like data.
🔴 The first handheld capture: the sky is 55 % of the frame
Two captures of the same game on the OneXFly, differing only in internal resolution:
| scope | 640×360 | 1920×1080 | scales |
|---|---|---|---|
| frame (median) | 13.90 ms | 71.64 ms | 5.2× |
sky | 6.11 | 39.60 | 6.5× |
raster + shade | 3.70 | 27.83 | 7.5× |
shadows | 0.98 | 1.28 | 1.3× |
blit | 0.14 | 1.05 | 7.7× |
cull | 0.046 | 0.042 | 1.0× |
Nine times the pixels, and everything that scales with them does: the
frame is fill-rate bound, and the sky alone owns more than half of
it. shadows and cull do not move — they are geometry — and
together they are under 1.4 ms. They are not the problem, and no amount
of optimising them would have shown up.
🔴 Even at 640×360 the sky costs 6.11 ms of a 13.9 ms budget. Nothing won elsewhere fits that in. This is what #771 predicted from shader arithmetic; it is now measured.
⚠️ One of the two captures came back as scope#ScopeId(137) — the
silently-unreadable case above. It was recovered by mapping ids against
the readable capture from the same binary, which works only because both
came from one session. Save a capture that has names.
🔴 The second handheld capture: a slow frame costs twice its GPU work
1165 frames of the same scene with the camera moving (#814). Split into the fastest quarter and the slowest, by self time:
| scope path | fast | slow | delta |
|---|---|---|---|
Render > Surface::get_current_texture > vkAcquireNextImageKHR | 11.14 ms | 72.71 ms | +61.57 |
GPU > raster + shade | 5.27 | 34.92 | +29.65 |
GPU > shadows | 0.40 | 0.73 | +0.34 |
GPU > blit | 0.31 | 0.50 | +0.19 |
GPU > cluster grid | 0.075 | 0.167 | +0.09 |
The engine’s own CPU work is under 1.5 ms of a 75 ms frame. Everything else is the GPU, or waiting for it. But how much waiting is the finding — each frame against the GPU work that produced it:
| decile | frame ms | GPU ms | frame/GPU |
|---|---|---|---|
| p20 | 13.88 | 5.35 | 2.60 |
| p50 | 30.71 | 26.63 | 1.15 |
| p60 | 35.41 | 33.64 | 1.05 |
| p80 | 72.54 | 35.64 | 2.04 |
| p90 | 85.48 | 39.94 | 2.14 |
The same GPU work produces two different frames. 33.5 ms of GPU
became a 34.7 ms frame 167 times and a 69.4 ms frame 80 times — the bad
outcome exactly double, which a GPU does not do and a swapchain does.
Under FIFO with two images the compositor holds one while the GPU draws
into the other, so the acquire waits out the compositor’s turn instead
of overlapping with it, and a frame that misses one vblank stays
serialised. KOOCH_FRAME_LATENCY=3 asks for a third image; it costs a
frame of input lag, so the default stays at 2 until a capture from the
device says which is worse.
⚠️ The tool that found this had been in the repo for weeks. The first
reading of this capture used a hand-written script that summed scopes
without descending the nesting, so a parent and its children came out as
siblings, vkAcquireNextImageKHR never appeared, and 51 % of the frame
looked unattributed. read_capture walks the tree properly and named it
on the first run. Read the capture with the tool that models parents.
🔴 The third handheld capture: the machine gets 46 % slower in two minutes
The measurement that was supposed to answer “is the profiler lying?” (#785) and answered something else instead.
First, the profiler is not lying. Three independent readings of the same run agreed to within a frame:
| source | reading |
|---|---|
read_capture, median frame | 27.80 ms = 36.0 fps |
| mangoapp’s overlay, read off the screen | 34–39 fps |
| 1022 frames over 30 s of wall clock | 34.1 fps |
⚠️ What disagreed was gamescope, and it was not measuring the game.
stats.pipe reports fps=144.02, which is the panel’s refresh rate, not
the application’s: that counter increments once per composition, and
the compositor composites whether or not the game produced a new frame.
The focus=steam field does not discriminate either. Use mangoapp as
the outside witness; it is an in-process Vulkan layer and counts presents.
Then the finding nobody was looking for. Same binary, same scene, a
camera that never moves, 8435 frames — read with --over-time:
| bucket | frame ms | power | clock |
|---|---|---|---|
| 1 (first ~40 s) | 27.8 | 12 W | ~1150 MHz |
| 3–10 (from ~2 min) | 40.4 – 41.1 | 10 W | ~850 MHz |
Nothing accumulates: RSS and VRAM are flat across the whole capture, and the GPU tree keeps the same shape — every pass grows by the same factor. The device simply leaves its boost state and settles, and the settled number is the real one: 40.7 ms, not 27.8.
🔴 Warm the handheld for two minutes before capturing. A capture taken cold reports a machine that only exists for forty seconds, and every budget judged against it is off by nearly half. The 13.9 ms target is 2.9× under 40.7, not 2× under 27.8.
Reading the numbers
🔴 The Table view is flat, and it opens sorted by call count. That
is why a capture of a 70 ms frame can look like it is made of
BindGroup::drop: 56 calls of 0.1 µs sort above one pass of 40 ms. It
aggregates by function across the whole frame and does not model
parents — its own text says it is for finding functions that are called
a lot. For “what is inside what”, use the Flamegraph, which is the
tree, or read_capture, which prints the same tree in a terminal.
puffin_egui has those two views and no third one.
⚠️ A scope lives to the end of its block. Declared mid-function
without braces, profiling::scope! swallows everything after it:
upload instances reported 1.900 ms of which 0.031 was the upload, with
the whole render path nested underneath, and raster + shade (fused)
was billed for Queue::submit. Both are braced now. A flat table cannot
show this — the self-time column in the tree is what makes it obvious.
- Self time excludes children. A parent can last 5 ms with 0.1 ms of self time; sort by self time for “what costs”, read the flamegraph for “who is responsible”.
- The table’s headers sort. It opens sorted by call count, which is the least useful column.
Surface::get_current_texturebeing the largest entry is not a problem — it is the wait for the compositor. The first release capture had it at 2.7 ms of a 4.9 ms frame, with the whole engine rendering a viewport in 0.69 ms.- ⚠️
frameappears twice per editor frame: the View and Game panels each render the scene, with their own cull and shadow passes.
Profiling the game, which is the point of all of this
Everything above measures the editor: this machine, plugged in, drawing a viewport. The number the graphics roadmap is judged against is a frame of a game on the OneXFly at 10 W, and no measurement taken here produces it.
cargo build --release --features profiling
The binary opens 0.0.0.0:8585 and streams every frame to whoever
connects. Nothing else to write: DefaultPlugins carries
ProfilingPlugin whenever the feature is on, so a game becomes
profilable without its main.rs changing.
Then, in the editor’s Profiler panel, switch the source from This
editor to A running game, type the handheld’s address and press
Connect. The address is remembered in editor_config.ron, because it is
a home-network address that is needed every session and wrong in a way
that looks like the profiler being broken.
puffin_viewer --url 192.168.0.36:8585 reads the same socket, if a
second application is preferable to a panel.
Capturing without a person watching the clock
cargo run -p kooch_editor_core --features profiling \
--example capture_remote -- 192.168.0.36:8585 out.puffin --seconds 300
The panel can do this, but it needs somebody watching for the right moment to press Save. That is fine for one capture and bad for what captures are actually for: an A/B needs two runs of the same route, and a click at the wrong moment silently makes them incomparable. This records a fixed window, keeps every frame, and writes the file.
It writes and does not analyse, deliberately. A scratchpad script that summed scopes without descending the tree once produced an issue built on a false premise; reading is left to the tool that models parents.
- 🔴
0.0.0.0, not127.0.0.1. Bound to loopback the game is reachable only from the handheld, which is the one machine that will not be running the viewer. The symptom is a connection that times out with nothing logged on either side. KOOCH_PROFILER_ADDR=0.0.0.0:9000moves it without a recompile, for when a build left running on the device is still holding the port.- Recording is on from the first frame here, unlike the editor panel. Nobody is going to press Record on a handheld over SSH.
- A port that will not open logs an error and the game keeps running. Killing the process someone wanted to measure is the worse answer.
- 🟢 The scope-name problem above does not apply to a late viewer:
the server keeps its own
ScopeCollectionand re-sends all of it to every client that connects, so one attached an hour in still gets names. - 🔴 It does apply to a late server.
scope_deltais a delta:new_framefills it fromnew_scopesand drains that list, so a server created after a scope first ran never learns its name and the viewer drawsscope#ScopeId(67)forever.ProfilingPluginruns before the first frame, and asks for a snapshot anyway so the guarantee does not depend on where it sits in the plugin list.
One frame boundary, and where it lives
Puffin builds a frame out of the scopes that closed between two
new_frame calls. Two boundaries in a frame produce a flamegraph of
half-frames; none produces a single frame that grows forever and never
renders.
The boundary is a system in Stage::Last, and it is a stage rather
than a line in the loop because there are two loops:
kooch_core::runner::default_runner for a headless app and
kooch_window’s winit loop for a windowed one. A stage runs under both.
⚠️ The editor marks its own boundary, in
kooch_editor_core/src/systems/render/ui.rs. The editor does not add
ProfilingPlugin; if it ever does, that call goes away in the same
commit or the flamegraph becomes half-frames.
Stages are named per stage on purpose
Schedule::run_pre_physics and friends expand run_staged! once per
stage instead of looping over an array. Puffin caches a scope’s id in a
static belonging to the call site and registers it under the first
name that site ever sees — one scope inside run_stage would file every
stage of every frame under Startup.
Two things the remote view does backwards, on purpose
- The flamegraph is drawn while frames arrive. The local view hides it while recording because drawing it cost 10.97 ms of a 15.98 ms frame. That cost lands on the machine drawing it, and the frames being measured are produced on the other one — the observer is finally outside the experiment.
- Clear reconnects instead of emptying the view. The collection that
turns scope ids back into names belongs to the process that recorded
them, which is on the handheld; there is no
emit_scope_snapshotto call from this side. Reconnecting resets the view and makes the server re-send every name.
GPU scopes — the half a CPU profiler cannot see
On the OneXFly the frame is GPU-bound at 96 % and the engine’s CPU
work is ~2 ms of it. Every CPU scope in this document can say the frame
is slow; none of them can say which pass spends it. GpuScopes
(kooch_core/src/gpu/profiler/) wraps wgpu-profiler to answer that,
and reports the results into puffin so they appear as a GPU thread
beside the CPU rows rather than in a second tool.
The passes it names, in the order they are recorded:
| Scope | Encoder | What it covers |
|---|---|---|
shadows | meshlet stage | four cascade culls + rasters, plus a cube face per point light |
cluster grid | meshlet stage | the froxel light index — four passes, two of them draws |
cull | meshlet stage | the scene-wide meshlet cull dispatch |
raster + shade | meshlet stage | the fused R64 pass, and the parent of the five below |
├ motion vectors | meshlet stage | previous clip position per pixel, from the unjittered camera |
├ shade: compute / (half rate) / shade: fragment | meshlet stage | the label names the path that ran, so a capture answers “which one produced this” without trusting a log line |
├ shade: upsample | meshlet stage | a sibling of the shading, not a child: the question is whether what half rate saves survives what putting it back on screen costs, and that is two numbers a capture can subtract |
├ taa | meshlet stage | the temporal resolve, when it is on |
└ tonemap | meshlet stage | HDR radiance to a display-referred image |
sky | game encoder | the raymarch #771 accuses |
blit | game encoder | the stage’s colour composited over the sky |
On the R32 path the middle rows are cull + raster A, hi-z build,
cull + raster B and shade instead.
Turned on by the same --features profiling as everything else. A build
without it carries a GpuScopes whose every method compiles to nothing,
so the render code has one shape rather than a cfg at each pass.
The API is begin / end, and both halves live on one encoder
wgpu_profiler::Scope borrows the encoder for the scope’s lifetime,
which leaves the code being measured with no encoder to record into.
begin returns a query and hands the encoder straight back.
🔴 A scope must close on the encoder that opened it. It pushes a
debug group, and wgpu rejects the encoder outright at finish() —
“A debug group was not popped before the encoder was finished”. A
profiling build would panic where a release build runs.
🔴 Nesting is by declared parent, not by call order. A scope opened
while another is open is not its child; begin_child is. Left to
begin, a pass and the pass containing it come back as siblings and
their times read as additive.
The GPU clock is not puffin’s clock
A GPU timestamp’s absolute value is undefined — wgpu-profiler says so
in as many words. Reported raw, the GPU track lands an arbitrary distance
from the CPU track and the viewer draws a frame stretched across the gap
with both ends too small to read. puffin_bridge translates each batch
so it ends at now.
⚠️ Durations and nesting are exact; the position on the axis is not. The results belong to a frame a few submits back, and wgpu exposes no calibrated timestamp to correlate the two clocks with. Read a GPU row for how long a pass took, never for what a CPU row was doing at that instant.
🔴 wgpu-profiler’s own puffin feature cannot be used here
It depends on puffin ^0.19.1, and this workspace patches puffin to 0.20.
Enabling it yields either an unresolvable lock or the two-GlobalProfiler
failure described at the bottom of this page. puffin_bridge.rs is its
src/puffin.rs adapted — 45 lines against API that is identical between
the two versions — and wgpu-profiler is taken with
default-features = false.
⚠️ TIMESTAMP_QUERY is three separate wgpu features. Scopes on an
encoder need TIMESTAMP_QUERY_INSIDE_ENCODERS specifically;
gpu/features.rs requests all three, conditionally, and an adapter
missing them yields scopes that measure nothing instead of a failed
submit.
In the editor, too
The editor builds its own render stage rather than going through
RenderPlugin, so it inserts its own GpuScopes at startup and closes
the frame in present_editor_frame. What it adds beyond the game’s
scopes:
editor ui— what egui costs on the GPU, kept apart from the viewport passes. “Why is the editor slow” and “how expensive is my scene” are different questions and now have different rows.- ⚠️
sky,cullandraster + shadeappear twice per frame — the View and Game viewports each render the scene, the same way the CPU scopeframedoes.
⚠️ It still measures a desktop viewport, plugged in. The budget is a frame on the handheld, and only a game build produces that.
What is not built yet
Scopes finer than a pass used to be listed here. They exist now — the
five children of raster + shade in the table above — and the answer
they gave is that the shading is the cost: on the settled handheld,
shade: compute (half rate) is 14.9 ms of a 28.1 ms raster + shade,
with the raster itself in the 4.4 ms the parent keeps as self time.
What is still missing:
- Scopes inside a shading dispatch. Which part of the BRDF
evaluation costs — the light loop, the shadow samples, the contact
shadow march — is not something a pass timer can separate. That is a
debug view’s job (
LightsPerPixel) or a shader permutation’s, not a scope’s.
⚠️ Four dependencies come from git, temporarily
puffin, puffin_egui, puffin_http and profiling are patched to git
revisions. Their released versions are one migration behind: puffin_egui
0.30.0 pins egui 0.33 against this workspace’s 0.35, and profiling
1.0.17 pins puffin 0.19 against the panel’s 0.20.
Two egui versions make the panel’s Ui a different type from ours. Two
puffins are worse: two separate GlobalProfilers, every scope
recording into one and the panel reading the other, and the symptom is an
empty flamegraph with nothing to blame.
They are [patch.crates-io] rather than plain git dependencies so the
whole graph resolves to one copy, and pinned by rev rather than branch.
Remove the section when the releases land.
Retired
Pages describing code that no longer exists.
They are kept because the reasoning in them was real and cost something to arrive at, and because a decision is easier to revisit when you can still read what it replaced. Nothing here describes the engine as it is.
SDF ray-marching and its BVH
The engine’s original rendering path was signed-distance-field ray-marching, accelerated by a
BVH that several consumers shared. Both crates — kooch_sdf and kooch_bvh — were deleted in
July 2026.
The technique died; the data did not. Signed distance fields remain the representation behind the voxel and dual-contouring work, where they are extracted to meshes that go through the same GPU-driven meshlet pipeline as everything else. What was retired is the renderer that marched them directly.
The current path is described in Render Pipeline.
BVH-Driven Ray Marching
This chapter documents how the SDF ray-marcher integrates with the
GPU LBVH builder shipped in kooch_bvh (issue #115 PR-3) to skip
evaluating primitives whose AABB does not contain the current
sample point. The integration is the subject of #115 PR-4.
Why a BVH at all
Without spatial culling, eval_scene(p) had to evaluate every SDF
primitive at every sphere-tracing step. A planet-scale scene with
~1 M primitives would burn ~1 M transform_point + sdf_* calls
per ray per step, and a ray takes up to 256 steps. The
arithmetic is unforgiving: even at 50 ns per primitive eval the
shader would not converge inside a frame budget.
A BVH lets each ray query “which primitives are near p” in
O(log N) and limits the per-step work to exactly the leaves the
ray currently overlaps.
Data flow
┌─────────────┐ ┌───────────────┐ ┌────────────────────┐
│ ECS │ ─► │ update_scene │ ─► │ BvhState (S4) │
│ (SDF │ │ (Vec<Aabb>, │ │ ├─ BvhGpuBuilder │
│ compos) │ │ Vec<LeafA…>) │ │ ├─ slot_a / slot_b│
└─────────────┘ └───────────────┘ │ └─ pending build │
└─────────┬──────────┘
│ kick_if_dirty
▼
┌────────────────────┐
│ Bvh::build_gpu │
│ (PR-3 — Morton + │
│ onesweep + Karras)│
└─────────┬──────────┘
│ poll_swap
▼
┌──────────────────────────────────────┐
│ slot[current_slot]: nodes + indices │
│ + leaf_aabbs (stable, not pending) │
└─────────┬────────────────────────────┘
│ bind to fragment shader
▼
┌──────────────────────────────────────┐
│ raymarch_main.wgsl::eval_scene_bvh │
│ per-step stack walk, per-role acc │
└──────────────────────────────────────┘
The two-slot pattern is the answer to the read-after-write hazard
on a single shared GPU buffer. While the renderer reads
slot_a.nodes_buffer for frame N, a new build can write into
slot_b for frame N+1 — the swap happens at poll_swap after the
build’s submission resolves. wgpu would otherwise insert a
synchronisation barrier to serialise the read against the write,
stalling the frame pipeline.
Traversal-driven CSG composition
Earlier drafts of PR-4 considered building a per-ray hit list of
primitive indices and then iterating the existing postfix CSG token
stream with a “skip if not in hit list” check. The parallel auditor
flagged this as marketing: it still iterates O(N) tokens per
sample, and an array<u32, 256> thread-local hit list spills 1 KiB
per ray into private memory — register-file death on RDNA 2 / 4.
The shipped design replaces the postfix token stream entirely. The BVH traversal is the evaluation loop. Each leaf carries its CSG role (ADD / INTERSECT / SUBTRACT) and per-instance smoothness, and the traversal accumulates per-role distances inline:
fn eval_scene_bvh(p: vec3<f32>) -> f32 {
var add_acc = ACC_UNION_IDENTITY; // +1e10
var int_acc = ACC_INTERSECT_IDENTITY; // -1e10
var sub_acc = ACC_UNION_IDENTITY; // +1e10
// ... stack walk, point-in-aabb cull, per-leaf eval + combine ...
var result = add_acc;
if scene_meta.has_intersects != 0u {
result = sdf_smooth_intersection(result, int_acc, scene_meta.k_int_scene);
}
if scene_meta.has_subs != 0u {
result = sdf_smooth_subtraction(result, sub_acc, scene_meta.k_sub_scene);
}
return result;
}
The “default tree” shape (smooth_subtract(smooth_intersect(adds, ints), subs)) is preserved — it is now expressed structurally by
the per-role accumulators + the fixed final combination. Per-role
k_max lives in SceneMeta; per-instance smoothness lives in
each LeafAabb.
Identity elements
| Role | Combinator | Identity value | Why |
|---|---|---|---|
| ADD | smooth_union | +∞ (1e10) | smooth_union(+inf, x, k) ≈ x |
| INTERSECT | smooth_intersection | -∞ (-1e10) | smooth_intersection(-inf, x, k) ≈ x |
| SUBTRACT | smooth_union | +∞ (1e10) | subs are unioned, then subtracted |
Picking 1e10 (rather than f32::INFINITY) is deliberate — keeps
the smooth-blend math NaN-free under all inputs.
Determinism
smooth_union and smooth_intersection are not strictly
associative in float32. The cull-vs-cull byte-identity regression
test (cull_vs_cull_byte_identical_n_*) requires the per-role
accumulator visit order to be a function of BVH topology only —
never of runtime ray geometry.
The traversal pushes left BEFORE right on its 32-deep stack, so
pop order is right-first and stable across frames. Do not switch
to a t-near-sorted children push without re-deriving the
determinism story from scratch — the regression test will catch it,
but the failure mode is “single-pixel-bit mismatches that confuse
post-processing”, not a clean panic.
AABB inflation
A primitive’s AABB is computed from its analytic shape, scaled by
the entity’s scale, rotated, translated, and inflated by its
role’s k_max. Smooth blends extend the support beyond the raw
geometry; without the inflation, a primitive whose surface lies
exactly on its AABB would have its smooth-union tail truncated at
the cull boundary.
Per-role inflation (rather than per-instance) keeps the AABB tight:
a primitive in a scene where every other ADD has k = 0.1
inflates by 0.1 even if its own smoothness = 0, because the
operator between them carries the larger k.
Performance scaling
Measured on a Ryzen 9 9800X3D + RX 9070 XT, Bazzite F43 / Mesa
radv (#115 PR-4 S11 bench output):
| N primitives | BVH (cs_main) | Fullscan (cs_fullscan) | Speedup |
|---|---|---|---|
| 1 024 | 3.79 ms | 3.55 ms | 0.94× |
| 10 240 | 5.93 ms | 6.88 ms | 1.16× |
| 65 000 | 4.57 ms | 14.56 ms | 3.19× |
Small-N is dominated by stack-walk overhead. The cross-over sits
between 1 k and 10 k in this configuration; from 65 k upwards the
cull is the dominant cost saver, exactly as the
O(N) → O(log N) goal predicts.
What lives where
crates/kooch_bvh/— the GPU LBVH builder (PR-3). Public API:Bvh::build_gpu,BvhGpuBuilder,BvhGpuBuild,GpuBvhHandle.crates/kooch_render/src/raymarch/bvh.rs—BvhState: double- buffered slots + dirty hash + kick / poll_swap lifecycle.crates/kooch_render/src/raymarch/aabb.rs—primitive_aabb: per-type local half-extents → world-space inflated AABB.crates/kooch_render/src/raymarch/instance.rs—LeafAabb(32 B std430),SceneMeta(64 B uniform), CSG role constants.crates/kooch_render/shaders/raymarch_main.wgsl—eval_scene_bvhtraversal + per-role accumulators + fixed final combination.
Out of scope (filed as follow-ups)
- Refit BVH (#115 checkbox 110) — incremental updates without a full rebuild. Useful for scenes with constant primitive count and small per-frame motion.
- OBB-exact AABBs (vs the current
abs(rot_matrix) · half_extentsenclosing OBB-AABB). Tighter cull at the cost of per-frame CPU work. - Workgroup-shared bitmap cull (vs the current per-thread stack walk) — blocked on tile-based shading, which itself is blocked on the G-Buffer (#132).
- Archetype-level dirty marker (vs the current
u64hash of primitive bytes + leaf metadata).
Multi-consumer BVH
This chapter documents the engine-shared GPU BVH that backs three
consumers in lockstep: the SDF raymarch culling from PR-4, the
physics broadphase from kooch_physics::broadphase, and the GPU
frustum cull behind the mesh pass. It is the subject of #115 PR-5,
the closing PR of issue #115.
The previous chapter (BVH-Driven Ray Marching)
introduced the GPU LBVH builder and the two-slot double-buffer that
sidesteps wgpu’s read-after-write hazard. PR-5 generalises that
state into kooch_bvh::SharedBvhState — a single resource the rest of
the engine binds against.
Why one structure for three consumers
Acceptance criterion 116 of #115 demands: “multiple systems use the same structure.” Independently, each consumer wants the same set of spatial queries — find leaves overlapping an AABB, a frustum, a ray. Building three private BVHs over the same scene each frame would triple the GPU build cost (and, more painfully, the per-frame readback the CPU mirror requires) without any algorithmic gain.
The shipped architecture builds one BVH per scene-dirty frame, lets
every consumer bind the same nodes / sorted_indices / leaf_aabbs
buffers, and hides the lifecycle behind a single resource. Per-
consumer side-payloads (raymarch’s per-instance smoothness, future
collider mass / restitution) live in the consuming crate and ride
the same kick → swap pulse via the
type-state BuildToken.
flowchart LR
Scene["ECS scene<br/>(SDF / collider / mesh)"] --> Hash["scene_hash<br/>(folds side payloads)"]
Hash --> Kick["SharedBvhState::kick_auto"]
Kick -->|build| Build["BvhGpuBuilder<br/>(Morton + Karras)"]
Kick -->|refit| Refit["refit_gpu<br/>(leaves + AABB only)"]
Build --> Slot[("slot[i].nodes<br/>slot[i].sorted_indices<br/>slot[i].leaf_aabbs")]
Refit --> Slot
Slot --> Raymarch["raymarch_main.wgsl<br/>traversal-driven CSG"]
Slot --> Broadphase["BroadphasePairs::collect<br/>(CPU mirror)"]
Slot --> Frustum["frustum_cull.wgsl<br/>→ DrawIndexedIndirectArgs[]"]
style Raymarch fill:#ff7f50,color:#000
style Broadphase fill:#90ee90,color:#000
style Frustum fill:#87ceeb,color:#000
The four-buffer slot
OutputSlot is the per-side double-buffer the orchestrator rotates.
Each side owns four parallel buffers, all sized for the current
primitive count and grown on demand:
| Buffer | Owner | Producer | Consumed by |
|---|---|---|---|
nodes | shared | GPU build / refit | every consumer (BVH walk) |
sorted_indices | shared | GPU sort | raymarch leaf payload lookup, refit topology |
leaf_aabbs | shared | CPU kick(...) | raymarch (gating), frustum cull |
| side payloads | private | per-consumer kick | raymarch fragment shader (RaymarchPayload[]) |
The first three live in kooch_bvh::shared::OutputSlot. Side payloads
live in the consuming crate (kooch_render::raymarch::bvh::PayloadSlot
holds RaymarchPayload[] at binding 5 of the raymarch pipeline). The
private buffers’ double-buffer mirrors the shared one — when
poll_swap flips current_slot, every parallel double-buffer must
flip alongside it or the renderer reads stale-paired data.
LeafAabb: per-leaf metadata + flag scheme
Every leaf carries 32 bytes of std430-clean metadata (mirrors the
WGSL LeafAabb byte-for-byte; offsets pinned by an offset_of!
test):
#![allow(unused)]
fn main() {
#[repr(C)]
pub struct LeafAabb {
pub aabb_min: [f32; 3],
pub flags: u32,
pub aabb_max: [f32; 3],
pub entity_id: u32,
}
}
The flags field is the multi-consumer contract — each consumer
filters by its own bit during traversal:
| Bit | Constant | Consumer |
|---|---|---|
| 0–1 | ROLE_RAYMARCH_* | Raymarch CSG role (ADD / INTERSECT / SUBTRACT) — only meaningful when IS_RAYMARCH is set. |
| 2 | IS_RAYMARCH | Leaf participates in the SDF raymarch traversal. |
| 3 | IS_COLLIDER | Physics broadphase (#42). |
| 4 | IS_VISIBLE_MESH | Frustum / occlusion culling (#91). |
| 5 | IS_LIGHT | Reserved for the light culling consumer (#27). Defined here so no future consumer accidentally claims the bit. |
| 6–31 | free | Future consumers. |
entity_id is the ECS entity index broadphase / frustum cull use
to return entity-keyed pair lists and visibility sets. Raymarch
ignores it.
Note: AABBs are inflated by the per-role smooth-blend
k_maxso the cull stays conservative under raymarch smooth blends. The S7 bench measured an envelope/tight pair-count ratio of 2.086 in a synthetic mixed scene (broadphase false-positives bench), justifying the per-role tighter AABBs follow-up filed at the close of #115.
Lifecycle: kick → poll_swap → bind
SharedBvhState::kick_auto is the production entry point per frame.
It picks between rebuild and refit using the should_refit
heuristic over the previously-mirrored leaf AABBs:
flowchart TD
Start["kick_auto(items, leaves, hash)"] --> Pending{"pending in flight?"}
Pending -- yes --> N1["return None"]
Pending -- no --> HashCheck{"hash == last?"}
HashCheck -- yes --> N2["return None"]
HashCheck -- no --> First{"cpu_mirror = None?"}
First -- yes --> Kick["kick → full rebuild"]
First -- no --> Card{"cardinality match?"}
Card -- no --> Kick
Card -- yes --> SR{"should_refit(prev, curr,<br/>0.25, 10.0)?"}
SR -- false --> Kick
SR -- true --> KR["kick_refit → fast path"]
Kick --> Token1["BuildToken<'_>"]
KR --> Token2["BuildToken<'_>"]
Token1 --> Attach["token.attach_payload(...)"]
Token2 --> Attach
Attach --> Drop["token drops"]
Drop --> Frame["next frame:<br/>poll_swap drains payloads"]
Suppression cases (pending in flight, hash unchanged) return None
before any state mutates — the consumer’s parallel buffers don’t
regrow either. This is the lesson from a footgun that earlier drafts
ate: see Type-state BuildToken.
Type-state BuildToken: enforcing the side-payload invariant
Before S3.5 the orchestrator returned bool from kick. Each
consumer maintained a pending_payload: Option<...> field that it
had to keep in lockstep with the orchestrator’s pending. That
invariant was implicit — and held only as long as a single consumer
played by the rules.
The footgun the type-state refactor closed was subtler than the “forgot to clear pending_payload on failure” scenario: the buffer-regrow on suppressed kick.
Note: Original behaviour:
kick_if_dirty(...)calledpayload_slot.ensure_capacity(n)before asking the orchestrator whether the kick was committed. If a previous kick was still pending, the second call would still grow the payload buffer — reallocating thewgpu::BufferArc — while the closure registered by the first kick still held a refcounted clone of the old buffer. On the eventual swap, the closure uploaded the captured payload into the orphaned buffer and the renderer kept reading the regrown one. Silent stale data, no panic, no warning.
SharedBvhState::kick and kick_refit now return
Option<BuildToken<'_>>. Some(token) is the only path that
exposes target_slot and n and admits an attach_payload(closure)
registration; None means the kick was suppressed and the consumer
mutates nothing. With ensure_capacity deferred until after the
token arrives, suppressed kicks no longer regrow buffers the
orchestrator will not write to. The invariant is type-enforced
instead of convention-enforced.
#![allow(unused)]
fn main() {
// production raymarch path (BvhState::kick_auto_if_dirty)
let scene_hash = Self::hash_scene(&items, &leaf_aabbs, &payloads);
let Some(mut token) = self.shared.kick_auto(
device, queue, items, leaf_aabbs, scene_hash, 0.25, 10.0,
) else {
return false; // suppressed → nothing to do
};
// from here, kick is guaranteed committed:
// token.target_slot() and token.n() are stable
// attach_payload runs on the matching swap, or drops on failure
attach_payload_upload(&mut self.payload_slots, device, &mut token, payloads);
}
On poll_swap success every attached uploader fires in registration
order with (queue, target_slot). On failure each uploader is
dropped without running, so the captured payload Vec and the
cloned buffer Arc are released cleanly — there is no
“who-clears-up-stale-pending” question.
CPU mirror: free byte-identical mirror from the build’s readback
The GPU build path always reads back the resolved nodes array and
the sorted_indices permutation — BvhGpuBuild::poll needs the
permutation to produce Bvh::leaves in Morton order. Pre-S4 the
orchestrator threw both away. S4 captures them in CpuMirror, owned
by SharedBvhState and refreshed on every successful build / refit.
CPU consumers (today: physics broadphase; tomorrow: debug tooling,
authoring traversals) walk the mirror with Bvh::for_each_aabb /
for_each_sphere / friends. No second build — the readback was
already paid for.
The byte-level invariant: cpu_bvh.nodes is bit-identical to the
GPU’s current_nodes() buffer after every swap, build or refit.
The Karras AABB union is element-wise (union(min, min) = min,
union(max, max) = max) and order-independent, so the GPU’s parallel
multi-dispatch propagation and the CPU’s post-order DFS produce the
same BvhNode array down to the bit pattern. The S7 sync goldens
in crates/kooch_bvh/src/shared/sync_tests.rs field-by-field
compare each BvhNode post-build and post-refit; any divergence
is a CPU↔GPU desync bug, not a precision issue.
Refit fast path
kick_refit rewrites only the leaves and re-propagates internal
AABBs over the existing topology. It skips Morton encoding,
the onesweep sort, and Karras’ internal-node construction
entirely. For a scene where centres did not move (or moved within
the heuristic threshold), this is the one-pass cost rather than the
full ~5-pass pipeline.
The refit is fence-only — a 4-byte staging copy at the end of
the encoder signals “submission completed”. No nodes readback per
frame; the production hot loop never pays the (2N-1)·32 B cost.
The CPU mirror updates in place via Bvh::refit_in_place, which
applies the new leaf AABBs through the stored sorted_indices
permutation and re-propagates internals on the CPU. O(N) work, no
GPU traffic.
should_refit(prev, curr, move_threshold_ratio, change_threshold_pct)
is the cheap predicate. Defaults from the PR-5 plan: 0.25 and
10.0 — refit is OK when fewer than 10 % of the AABBs moved their
centre by more than 25 % of their largest extent. Tighter values
land via the S7 bench results once a real workload tells us what
“moderate movement” means in practice.
Three consumers
Raymarch (PR-4)
The raymarch fragment shader binds nodes + sorted_indices +
leaf_aabbs + the raymarch-only RaymarchPayload[]. Each ray
walks the BVH in a 32-deep stack, gates leaves by IS_RAYMARCH,
reads the role bits and per-instance smoothness, and accumulates
per-role distances inline. The traversal is the evaluation loop —
there is no separate hit-list pass. Postfix CSG token streams from
#307 do not apply to this path.
Physics broadphase (S4 of #115 PR-5, #42)
kooch_physics::broadphase::BroadphasePairs::collect(&shared) walks
the CPU mirror, filters leaves by IS_COLLIDER, queries
Bvh::for_each_aabb for every collider, and returns canonical
(low, high) entity-id pairs deduplicated across the symmetric
query. CPU-first because narrowphase (#40) is still CPU; the
GPU broadphase path is filed for when narrowphase moves to the GPU
and the readback round-trip becomes the constraint instead of the
optimisation.
Frustum cull (S5 of #115 PR-5, #91)
kooch_render::frustum::FrustumCull dispatches a compute pass over
the GPU’s leaf_aabbs buffer and writes one
DrawIndexedIndirectArgs per leaf in original input order. Visible
leaves get instance_count = 1; culled or non-IS_VISIBLE_MESH
leaves get 0. The mesh pass consumes the buffer via
draw_indexed_indirect; the GPU command processor skips zero-
instance entries with no shader work. Zero CPU readback per
frame — the camera writes the frustum uniform once per change and
that is the only CPU→GPU traffic.
The shader is the per-leaf parallel positive-vertex slab test against the 6 frustum planes:
for (var i: u32 = 0u; i < 6u; i = i + 1u) {
let plane = frustum.planes[i];
let n = plane.xyz;
let pv = vec3<f32>(
select(aabb_min.x, aabb_max.x, n.x >= 0.0),
select(aabb_min.y, aabb_max.y, n.y >= 0.0),
select(aabb_min.z, aabb_max.z, n.z >= 0.0),
);
if (dot(n, pv) + plane.w < 0.0) { return false; } // cull
}
return true;
The 10k-cube AC test in
crates/kooch_render/src/frustum/tests/cull.rs runs the same
algorithm on the CPU and asserts byte-perfect agreement on every
one of the 10000 leaves — the GPU cull is the same computation
parallelised, not an approximation.
Module layout
| Path | Role |
|---|---|
kooch_bvh::shared::state::SharedBvhState | orchestrator + counters |
kooch_bvh::shared::pending::{BuildToken, SwapInfo, Pending} | type-state lifecycle handles |
kooch_bvh::shared::mirror::CpuMirror | CPU mirror struct + from_build / apply_refit |
kooch_bvh::shared::heuristic::{should_refit, kick_auto} | rebuild-vs-refit policy |
kooch_bvh::shared::slot::OutputSlot | per-slot stable buffer set |
kooch_bvh::leaf::LeafAabb | per-leaf metadata + flag bits |
kooch_bvh::Bvh::refit_in_place | CPU refit over an existing topology |
kooch_render::raymarch::bvh::BvhState | raymarch consumer wrapper |
kooch_render::frustum::FrustumCull | frustum cull GPU compute |
kooch_physics::broadphase::BroadphasePairs | CPU broadphase consumer |
kooch_render/shaders/frustum_cull.wgsl | the compute shader |
Out of scope (filed as follow-ups)
- Tighter per-role AABBs. The S7 bench measured a synthetic envelope/tight pair-count ratio of 2.086 — above the 1.5× action threshold. Filed as a priority issue at the close of #115.
- GPU broadphase. Stays CPU-first until narrowphase (#40) moves
to the GPU; the API surface (
BroadphasePairs::collect(&shared)) stays the same, only the body offrom_cpu_mirrorswaps for a compute dispatch. - Mesh pass GPU-driven integration. S5 ships the indirect-args
buffer; wiring the mesh pass to consume it via
draw_indexed_indirect(and to maintain a per-leaf mesh metadata buffer once entities ship distinct meshes) is filed separately — it depends on the engine’s mesh atlas design. - Archetype-level dirty marker. The
u64hash_scenefold is conservative; an ECS-side change-detection signal would let kick decisions skip the hash computation entirely.
Decisions Log
Chronological record of architectural decisions. Each entry captures what was decided, why, and what it cost (or commits us to). Reading this is faster than reading 12 PR descriptions to figure out why a function exists.
Format:
YYYY-MM-DD · Title (refs)
Decision: one-paragraph summary. Why: the constraint or insight that drove it. Consequence: what this commits the project to, or what it rules out.
Coordinate system
Permanent · Right-handed, -Z forward, Y up
Decision: Use the same convention as glTF, OpenGL, Vulkan, Blender, Maya, and
glam’s default view/projection matrices. -Z is forward, +Y is up, +X is right. Why: Going against the grain means flipping Z at every loader boundary (glTF importer, USD, FBX, exported camera transforms). Unity picked left-handed because of DirectX heritage; they pay the flip cost in every importer. We don’t want to. Consequence: Identity quaternion(0,0,0,1)faces -Z. Anyone coming from Unity (left-handed +Z forward) needs to mentally flip when authoring scenes.
SDF ray-marching as primary render path
Permanent · Sphere tracing, not rasterization
Decision: Primary render pipeline is SDF sphere tracing. Mesh rasterization is a secondary pass layered on top. Why: Want experimentation latitude (Mario Galaxy gravity, infinite procedural geometry, smooth blends) that rasterization can’t give cheaply. SDFs are also a clean GPU-resident data model that pairs well with the hybrid ECS. Consequence: Performance ceiling is lower than a modern PBR rasterizer at the same hardware budget. Hybrid mesh+SDF (Dreams style) is the long-term escape hatch but ~3000 LOC away.
Hierarchical coordinate scales (NOT floating origin)
2026-04 · Issue #50 (blocks #51, #52, #54, #90)
Decision: Universe (i64 sector + f64 offset) → Solar system (f64) → Planet (f32) → Surface (f32) → Camera-relative render (f32). Origin rebasing is a trigger when the player gets far from origin, not the sole mechanism. Why: Inspired by No Man’s Sky, Star Citizen, KSP. Pure floating origin works for one scale (Outer Wilds) but breaks down across astronomical ↔ surface transitions. Consequence: Multiple
Transformtypes, conversion at scale boundaries, but precision stays bounded inside each scale. Locks in design for camera-relative transforms (#51), sector boundaries (#52), world streaming (#54), and navigation (#90).
SDF tracer roadmap
2026-04-22 · Step count bump 128 → 256 (PR #227 closes #221)
Decision: Quick fix raised
max_stepsdefault from 128 to 256. Visible gaps in concave SDF necks reduced to imperceptible. Why: Two attempts at smarter tracers failed: Enhanced Sphere Tracing (PR #222, branch killed) and the iq closed-form ellipsoid (#229, killed) both produced gray patches at CSG seams. Brute force was the surgically minimal change that worked. Consequence: ~2× iteration cost on the GPU, predictable. Locked until Segment Tracing (Galin et al. 2020, issue #224) lands and lets us drop back to ~128 steps with correct Lipschitz bounds.
2026-04-22 · ESL is dead (#222 killed twice)
Decision: Do not retry Enhanced Sphere Tracing while the shader uses the
s_minworkaround for non-uniform scale. Two attempts on 2026-04-22 confirmed it cannot work. Project memory documents the dead ends. Why: Naive ESL with non-Lipschitz CSG produces poly-edged gray patches near silhouettes (calc_normal crosses gradient discontinuities at seams). Lowering omega from 1.6 to 1.3 makes it worse, not better. Hybrid (over-relax only in far field) doesn’t fix it either. Consequence: Segment Tracing (#224) is the only open path forward for tracer optimization.
Editor camera as ephemeral ECS entity
2026-04-22 ·
EditorCamera + EditorOnly + PerspectiveCamera + Transform(#199 → PR #219)Decision: The editor camera is a regular ECS entity, not a
Resource. It carries anEditorOnlymarker that theEphemeralComponentsfilter checks during scene serialization — the entity is invisible tofrom_ecsand todespawn_all. Why: The renderers already iterate cameras by priority. Resource approach would require a special code path in every renderer and a manual mode swap. Entity approach lets future editor-only entities (gizmos, grid, debug lights) ride the sameEditorOnlyfilter without new plumbing. Consequence: Play mode strips the editor camera. Scenes need their own non-ephemeral activePerspectiveCamerato render anything in play. UX feature “Play uses editor view” is a separate future issue.
2026-04-22 · Quat internally for cameras, Euler-cached for inspector
Decision: Cameras store
Quatinternally for orbit/fly rotations. The inspector caches Euler angles per-field to avoid gimbal lock and to keep each X/Y/Z field stable while the others are edited. Why: Inspector UX needs each axis editable independently of the others — aQuatround-trip mangles two axes when you edit the third. Cameras need continuous quaternion math for cinematics. Different contexts, both correct, do not unify. Consequence: Two rotation representations exist in the codebase. Documented; do not “simplify”.
2026-04-22 · Fly-mode pivot is camera position, NOT focus point (#199 PR #219)
Decision: In orbit mode, rotation pivots around
focus_point. In fly mode, it pivots around the camera itself;focus_pointis re-anchored tocamera_position + forward * distanceafter each rotation. Why: Bug found in manual testing — without this invariant, fly mode would drift laterally as you looked around (your “feet” moved when you turned your head). Consequence:EditorCameraControllercarries explicit logic for the two modes; do not refactor toward a unified pivot.
Render orchestration
2026-04-23 · Editor is the render orchestrator for offscreen (PR #235 closes #129)
Decision:
kooch_editor_core::systems::startupinstantiatesRayMarchRenderer + MeshPassRenderer + SkyRenderPassdirectly asResourcesandviewport::render::render_viewportruns the three passes in one encoder against the offscreenViewportTarget. TheRayMarchPluginis not used by the editor. Why: Doing this through aRenderGraphabstraction would have been ~500 LOC for a 3-pass pipeline. Plain procedural orchestration wins until there are 5+ passes. Consequence: When a fourth pass (post-process composite, G-Buffer, shadow map) is added, re-evaluate building aRenderGraph.
2026-04-25 ·
RenderPluginIS the game render path (PR #267 closes #260)Decision:
RenderPlugin(inkooch_render) is the play-binary orchestrator. Same 3-pass pipeline as the editor’srender_viewport, but writing to the swapchain surface instead of an offscreen texture.RayMarchPluginstays as the standalone demo path (raymarch_demo). Why: StubRenderPluginthat only cleared the screen was dead weight. The semantically right name for “the game’s render plugin” isRenderPlugin. No separateGameRenderPlugininvented. Consequence: Editor and play share one conceptual model with two orchestration callsites. A future regression in either path is immediately reproducible in the other.
2026-04-23 · Mesh pass: two pipelines, one target, one encoder (#129)
Decision: Raymarch pipeline runs first with
LoadOp::Clear, mesh pipeline runs second on the same target withLoadOp::Load. No shared shader, no unified material system — explicitly NOT unified. Why: Unifying the SDF shader and the mesh shader would have been a multi-week refactor for an MVP feature. Two pipelines is correct enough. Consequence: Material system per-pipeline grows independently until #130 PBR forces convergence.
2026-04-23 ·
Depth32Floatconstant +LessEqualfor sky (PR #237 closes #236)Decision: All passes share
VIEWPORT_DEPTH_FORMAT = Depth32Floatas a publickooch_renderconstant. Sky pipeline usesCompareFunction::LessEqual; mesh pipeline usesLess. Why: Sky writesfrag_depth = 1.0explicitly so meshes behind it can supersede. Depth clears to 1.0. WithLess,1.0 < 1.0is false and sky never draws — black viewport. Bug found in first manual test. Mesh keepsLessbecause no mesh is exactly at the far plane. Consequence: Future depth format change is one line inlib.rs. Documented in shader comments next to the comparison choice.
Sky / atmosphere
2026-04-23 ·
SkyRendererdoes NOT blend between multiple skies (PR #247 closes #246)Decision:
SkyRendereris a singleton-by-priority component. Highest-priority active wins; no crossfade composite pass. Day/night is animated within one shader, not by blending two materials. Why: Crossfading entire sky materials was scope creep. Unity, Unreal, and Bevy don’t do it natively either. Animated parameters within one material handle the real-world use case. Consequence: NoSkyCompositepass. If we ever need sky crossfade, that’s a new pass with explicit cost.
2026-04-23 ·
SkyRendererandAtmosphereVolumeare separate componentsDecision:
SkyRenderer= singleton ambient backdrop (deep space or default gradient).AtmosphereVolume= volumetric shell per-planet with scattering, N coexisting in the world. Why: Architecture ofstellar_deliveryand Unreal’sSkyAtmosphere. Singleton sky and per-planet atmosphere have different lifetimes, coordinate frames, and shader budgets. Forcing one component to do both invents complexity. Consequence: Two paths to maintain, both simpler than one overloaded path.AtmosphereVolumeships in a future PR (#248).
Scene management
2026-04-24 ·
SceneManageragnostic of component types (PR #266 closes #259)Decision:
SceneManagerlives inkooch_ecsand knows nothing about Camera, Sky, or any specific component. The default scene bootstrap (Camera + Sky entities written to disk on project create) lives inkooch_editor_core::project::ensure_default_scene. Why: Same split asEphemeralComponents: mechanism in core, policy in editor. LetsSceneManagerbe reused by headless tools that have a different “default scene” idea. Consequence:kooch_ecscannot be the place to teach the engine “every project starts with a Camera and a Sky.” That decision is the editor’s.
2026-04-25 · Scene bootstrap runs at
Stage::First, NOTStage::Startup(PR #267 closes #260)Decision:
SceneBootstrapPlugin::load_boot_sceneruns atStage::First, which fires once-per-frame after allStage::Startupsystems complete. TheBootSceneresource is consumed on first call so it’s effectively a one-shot. Why: Race detected in manual testing — if userregister_componentsand SceneBootstrap both ran atStage::Startup, scene deserialization happened beforePlayer(custom component) registered →unknown component type: Playererror.Stage::Firstguarantees a clean handshake. Consequence: Replicable pattern for any future plugin that depends on user-registered state. First frame waits one stage tick for the scene to appear; imperceptible at 60 FPS.
2026-04-25 · Play uses
cargo run --manifest-path, no exe-detection (PR #267 closes #260)Decision:
EditorAction::Playrunscargo run --manifest-path <project>/Cargo.toml -- --scene <abs>. The oldis_project_binaryflag andcurrent_exe.starts_with(target)guard are gone. Why: Cargo handles incremental build, caching, and run as one primitive. Custom exe detection only worked for the half of project launches that ran the binary directly; not for in-processOpenProjectflows. The new approach works for both. Consequence: First Play after a code change costs acargo build(~0.1–30s). Editor stays responsive (cargo runs as child). Async-build modal with cancel is a future UX issue, not architecture.
2026-04-25 · Project template is play-mode-only (PR #267 closes #260)
Decision: Generated
main.rsis ~10 lines:App::new() + DefaultPlugins + register_components. The dual editor/play branching the old template carried is gone — the editor is its own binary, never embedded in user crates. Why: Cleaner mental model, cleaner code. The “editor inside the user binary” pattern was a leftover from before the editor binary existed; it confused exe-detection and confused users. Consequence: Existing user projects need migration (one-line change inmain.rs+ Cargo.toml cleanup). New projects are clean.
wgpu strategy
2026-04-23 · Stay on wgpu 29 for 24 months minimum (PR #239 closes #238)
Decision: Do not migrate to ash / vulkano / dx12-rs. No active migration trigger. Hybrid wgpu + ash only if RT pipelines become a requirement and upstream issue
#8560(Metal pipelines design) stays unresolved past April 2028. Why: Bevy ships Solari (path-traced GI) on wgpu in September 2025. If they don’t migrate prematurely, we don’t either. The audit indocs/research/wgpu-capabilities.mdlists 5 concrete migration triggers; none are active. Consequence: No raw Vulkan / Metal escape hatch in user code. Specific blocked features (mesh shaders cross-backend, FSR 2 viable, 3D texture arrays, GPU memory reporting) work around or wait.
2026-04-24 ·
PipelineCachewithfallback: true(PR #257 closes #251)Decision: Enable
wgpu::PipelineCachekeyed on(adapter.name, driver_info, engine_version). Save onDrop for GpuContext; SIGKILL is tolerated. Why: 100–500 ms cold-start saving per pipeline. TheunsafeofDevice::create_pipeline_cacheis covered byfallback: true— driver rejects an invalid blob without UB. Hash key invalidates on driver upgrades. Consequence:~/.cache/kooch/pipeline_cache/<hash>.binfiles accumulate (they’re tiny). Deleting them is harmless; engine regenerates on next run.
2026-04-24 ·
PowerProfileenum lives inkooch_core::power(PR #258 closes #253)Decision:
PowerProfile::{Plugged, Balanced, Battery, Debug}as aResourceinkooch_core::power. Auto-detect on Linux via sysfs and$SteamDeckenv var. Override viaKOOCH_POWER_PROFILE. Why: The Steam Deck / OneXFly target makes battery awareness non-negotiable. Renderers will gate quality defaults (DoF, SSR, TAA off in Battery) per-feature in future PRs. Consequence:kooch_corecarries the policy enum but renderers do not yet read it. Integration is per-feature PR work, intentional.❌ Reverted 2026-08-20. Nothing ever read it. Four months on, the only consumer was the editor’s own menu drawing its current value — a panel with no backend, which is this project’s own named antipattern with the arrow reversed. Removed with the menu it lived in,
KOOCH_POWER_PROFILEincluded.What replaced it is better than what it promised: quality is decided per PROJECT in
.rendersettings—compute_shading,shading_rate,upscale,render_scale— with every value measured on the target rather than inferred from whether a cable is plugged in. A handheld preset that meets the 13.9 ms budget is written down in the roadmap, and it came from captures, not from a heuristic reading sysfs.⚠️ The lesson worth keeping: the consequence line above described the failure and shipped anyway. “The renderers do not yet read it” was true on the day and stayed true, and an enum nobody consumes costs a menu, an env var, a detection heuristic and four months of looking like a feature.
Inspector / editor UX
2026-04 ·
GlobalTransformtolerates shear, inspector warns (PR #217 closes #214)Decision:
GlobalTransformis a 4×4 matrix that can carry shear (non-uniform scale through a rotated parent), but the inspector does not attempt to decompose it. Instead it shows a warning icon and exposes alossy_scale()helper. Why: Decomposing shear is ambiguous (multiple TRS triplets reproduce the same matrix). Hiding the issue creates worse bugs downstream. Educating the user is honest. Consequence: Users authoring shear-causing parent chains see the warning. No automatic “fix” is offered.
2026-04-25 · Three-system editor architecture: Gizmos / Editor / UI Toolkit (research #276, doc
docs/research/editor-three-system-architecture.md)Decision: Editor evolves into three separate, pure-Rust, custom-built subsystems:
kooch_gizmos(visual gizmo API + visualizer registry, usable at runtime too) +kooch_gizmos_handles(interactive translate/rotate/scale, editor-only);kooch_editor_api(user editor extensions: inspectors, panels, actions, loaded via libloading from a usereditor/crate);kooch_ui(declarative HTML-like UI Toolkit:.kooch_uimarkup +.kooch_styleCSS subset + Rust behavior, retained-mode with fine-grained signals, coexists withegui). Why: Godot’s self-hosted monolith couples concerns; Unity’s separation (Gizmos / Handles / Editor scripts / UI Toolkit) lets each evolve independently and gives users one mental model per need. We follow Unity’s separation. External libraries —transform-gizmo, Slint, Dioxus — rejected: only cover narrow slices, none address user-extensibility for custom component visualizers, and we want the engine to be self-contained pure Rust with no FFI. Consequence: A multi-quarter commitment. Three implementation epics (one per subsystem) replace the original gizmo epic #198 as sub-epic of the Gizmos one. The currentkooch_render::gizmosmodule (PR #277) migrates intokooch_gizmosin phase 1. Thekooch_uitoolkit is the heaviest piece (multi-month) and runs in parallel with the others.
2026-07-25 · Keep
kooch_ecs; do not adoptbevy_ecs(decision #605)Decision:
kooch_ecsstays and improves in place.bevy_ecsis the reference to steal individual designs from, never a dependency. Why: #603 removed the GPU component storages that had justified a custom ECS, so the justification was re-derived from measurements rather than repeated. No technical blocker was found —bevy_ecsis genuinely standalone (65 crates, nobevy_app/bevy_render), the GPU-driven renderer touches the ECS throughQueryin four places, andbevy_reflectexpresses our custom field attributes. What decided it: 42 call sites reach into component storage directly against 3 that useQuery, which is work required in every path and whichkooch_ecscan already express;EntityAllocator::revivepreserves entity identity across Play/Stop, whichbevy_ecsrefuses by design while 177 sites outside the crate hold anEntityin a field; and 51 of the 80 affected files arekooch_editor_core, the one area where Bevy offers no upstream design to copy because it has no editor. Consequence: improvements are ordered by demonstrated pain, not by feature parity. Encapsulating the ECS behindQueryis the prerequisite for any future backend change — today the contact surface is 80 files. The schedule graph belongs inkooch_core, not the ECS: the ordering bugs it would fix live inapp.rs.
2026-07-25 · Entities are referenced by a persistent id, not a handle or an index (feat #607)
Decision: a component may hold an
Entityand have it survive a save. Identity is an opt-inPersistentId(EntityGuid); the wire form isEntityRef, which isLive(Entity)in memory andPersistent { scene, id }on disk. Ids are scene-local and remapped per instance.Parentbecomes an ordinary component andparent_indexis legacy-read-only. Why: reflection had no way to express “points at an entity”, so the scene format carried the parent link out of band. That worked for one component and could not scale: joints hold two entities, and an index into one document cannot address another scene at all. Assets had already solved the same problem by addressing through aGuid. Scene-local ids follow Unity (SceneLoadFlags.NewInstance) and Unreal (Level Instances), and are what allows one scene to be instantiated twice without both copies claiming the same identity. Consequence: saving a scene mutates the world, because whether an entity is referenced is only known once references are written —SceneDocument::from_ecstakes&mut Resources.Entitystill does not implementSerialize, so serialising a live reference is an error rather than a handle written to disk. A reference whose target is absent saves and loads as unset, which is the normal state for a reference into a non-resident cell under #566. Unblocks #560 and cross-scene references.
2026-07-25 · The world is the container; scenes are content loaded into it (feat #609)
Decision:
SceneManagerbecomes a registry of open scenes with one active, instead of a single current path whose load replaced the world. Scenes carry aGuid;SceneMemberrecords an entity’s authoring home and is derived on load rather than serialised. Saving writes only one scene’s entities; closing despawns only its own. Why: the model #566 settled on. One scene per world is “the entire world in one section”, which cannot express a space station and an asteroid field as separate content occupying the same volume, nor make “close the station” different from “walk away from it”. #607 supplied the prerequisite by making entity references survive a save. Consequence: there is always a scene, even before the first save, and entities with no membership are adopted by the active scene when it saves — otherwise anything spawned in the editor would belong to nothing and be written to no file. The reference remap table is keyed by(scene, id), never by id alone: ids are scene-local, so two open scenes both numbering an entity 1 is ordinary. Opening the same file twice is refused, because two copies would share every entity id. Scene transforms and instancing are deliberately deferred — they need a decision on whether the transform bakes at load, as Unreal’s Embedded Level Instances do.
2026-07-26 · A scene is the prefab; instancing and editing are different operations (epic #611)
Decision: prefabs are scenes instanced with their entity ids remapped per instance. No separate format. Built in two phases: runtime instancing first, the linked-with-overrides prefab system after. Why: #609 refuses to open one file twice, which is right for editing and wrong as a limit on instancing — and entity ids were made scene-local in #607 precisely so instances could remap them. Unity’s prefab is a serialised scene file, and Godot says so outright with
PackedScene; both store an instance as a reference to the source plus a list of differences rather than a copy, which is what keeps editing the source propagating to its instances. Consequence: two things must be settled in phase A because they touch already-merged types — whether a scene must have a single root (instancing as a unit with a transform needs one, and our documents are a flat list), and how an outside reference names this instance rather than the prefab, sinceEntityRef::Persistent { scene, id }is ambiguous once a scene is instanced twice. Phase B waits on one decision: whether overrides are per field, as Unity and Godot both do, or whether editing an instance promotes it to its own scene.