Introduction
zut is a hermetic, incremental, content-addressed build system written in Zig 0.16, with Rust and Zig as first-class languages. It is built from scratch — its own Starlark front-end, content-addressed store, sandboxed executor, and incremental engine — in the spirit of Bazel and Buck2.
⚠️ Status: early and experimental. zut is under active development. The design is settled (see the Architecture overview), but APIs, the Starlark dialect, and the CLI will change. It is not yet ready for production builds.
What makes zut zut
- Correct by construction. A build’s output is a pure function of its declared inputs. Undeclared inputs are made physically unavailable to actions by the sandbox, so “works on my machine” drift becomes a build error instead of a mystery.
- Fine-grained incrementality. The unit of caching is the action, not the target or the crate. Change one file and zut rebuilds exactly the actions that depended on it — nothing more.
- Content-addressed caching. Every action’s inputs and outputs are addressed by their hash (BLAKE3). An identical action never runs twice, whether the result is in your local store or a shared remote cache.
- Rust & Zig as peers. A
rust_binarycan depend on azig_libraryand vice versa, linked through the C ABI. Neither language is a bolt-on. - Remote-ready. The execution path is shaped like the Bazel Remote Execution API (REAPI), so remote caching (today) and remote execution (later) slot in without reworking the core.
- Extensible. Build rules like
rust_binaryandzig_libraryare definitions written in Starlark over a small set of primitives (providers + actions) — not hardcoded into the engine.
How this book is organized
- Getting started — build zut, write a
BUILDfile, run your first build, and import crates.io dependencies. - User guide — the CLI, configuration (
.zutrc), logging, remote caching, and the sandbox tiers. - Writing BUILD files — the Starlark dialect, rules and providers, and the built-in Rust/Zig prelude.
- Internals — how the engine fits together, for contributors and the curious.
- Project — how to contribute and where things are headed.
The authoritative design contract is
docs/ARCHITECTURE.md
in the repository; this book is the practical companion to it.
Installation
zut is built from source with the Zig toolchain. There are no other build-time dependencies — the TLS stack, the S3/GCS protocol clients, and the sandbox are all implemented in-tree.
Prerequisites
- Zig 0.16.0 (exactly — zut tracks the 0.16
std.IoAPI). Get it from ziglang.org/download. - Linux for the strong sandbox tiers (user namespaces and/or Landlock). Other platforms build and run, but fall back to a weaker preparation-only sandbox.
- Optional, for building Rust/Zig targets: a host
rustc/cargoand/orzig. zut discovers them at runtime; they are not needed to build zut itself.
Build
$ git clone https://github.com/Sh4d1/zut
$ cd zut
$ zig build # debug build → zig-out/bin/zut
$ zig build -Doptimize=ReleaseFast # optimized build
The binary lands at zig-out/bin/zut. Put it on your PATH, or run it through
the build system:
$ zig build run -- build //:hello
Run the tests
$ zig build test --summary all
This runs the unit-test suite plus the namespace/Landlock sandbox smoketests (which fork real processes, so they live outside the in-process test harness).
Verify it works
$ zig-out/bin/zut
zut 0.0.0-dev — a hermetic, incremental build system (digest: blake3, REAPI #9)
Commands:
build build a target and materialize its outputs into ./zut-out/
test build a target and run its test binary (a cached action)
fetch download crates.io deps into the CAS + generate //crates/BUILD
Run `zut help` for details.
Next: write your first build.
Your first build
A zut workspace is a directory tree with BUILD files. Each BUILD file
declares targets; a target is an instance of a rule (like rust_binary)
with some attributes.
A minimal example
Let’s build a Rust binary that links a Zig static library. Create this layout:
myproject/
BUILD
src/main.rs
greet.zig
greet.zig — a tiny C-ABI function:
export fn zut_add(a: i32, b: i32) callconv(.c) i32 {
return a + b;
}
src/main.rs — calls into it:
extern "C" { fn zut_add(a: i32, b: i32) -> i32; }
fn main() {
println!("2 + 3 = {}", unsafe { zut_add(2, 3) });
}
BUILD — wire them together using the built-in prelude:
load("@builtin//:rust.bzl", "rust_binary")
load("@builtin//:zig.bzl", "zig_library")
zig_library(
name = "greet",
src = "greet.zig",
libname = "greet",
)
rust_binary(
name = "hello",
crate_root = "src/main.rs",
srcs = ["src/main.rs"],
edition = "2021",
out = "hello",
native_deps = [":greet"],
)
Build it
$ zut build //:hello
zut build //:hello
sandbox: namespace (recursive ro host, net+pid isolated)
toolchain: rust[1.xx] zig[0.16.0]
jobs: 8
libgreet.a [zig] built
hello [rustc] built
-> zut-out/hello
done — 2 built, 0 cached
$ ./zut-out/hello
2 + 3 = 5
The label //:hello means “the target named hello in the BUILD file at the
workspace root”. A target in a subdirectory foo/bar is //foo/bar:name.
Incrementality in action
Run the same command again, unchanged:
$ zut build //:hello
...
up to date — 2 action(s), all cached
Nothing rebuilds: zut hashed the action inputs, found the outputs already in its
content-addressed store, and served them. Now touch only the Rust source and
rebuild — only the rustc action re-runs; the Zig library is still a cache hit,
because its inputs didn’t change.
Build state lives under .zut/ in your workspace (the CAS, the action cache,
and a work directory). Delete .zut/ to start cold; it is safe to .gitignore.
Next: import crates.io dependencies.
Fetching crates.io dependencies
zut builds are hermetic: the build phase has no network access. So fetching
third-party crates is a separate, explicit step — zut fetch — that runs once,
downloads everything into the content-addressed store, and generates a BUILD
file the hermetic build then consumes offline.
The workflow
Point zut fetch at a Cargo lockfile:
$ zut fetch Cargo.lock
zut fetch Cargo.lock
fetch serde 1.0.210
fetch serde_derive 1.0.210
...
42 crate(s): 42 fetched, 0 reused → .zut/crates.lock + //crates/BUILD
This does four things:
- Downloads each registry crate’s
.cratearchive over HTTPS (the only networked step — zut ships its own TLS viastd.http.Client). - Checksum-verifies each archive against the lockfile’s SHA-256, then unpacks it into a CAS Tree (a content-addressed directory).
- Records a manifest at
.zut/crates.lockmapping eachname versionto its Tree digest. - Generates
//crates/BUILD—crate_library/crate_proc_macrotargets wired with the right editions, features, and dependency edges.
When cargo and the host triple are available, zut runs cargo metadata for a
faithful model (real editions, resolved features, lib/proc-macro kinds). Without
them it falls back to a lockfile-only model (edition 2021, no features).
Idempotent & incremental
zut fetch is safe to re-run. A crate whose Tree is already in the CAS (per the
prior .zut/crates.lock) is reused, not re-downloaded:
$ zut fetch Cargo.lock
reuse serde 1.0.210
...
42 crate(s): 1 fetched, 41 reused → .zut/crates.lock + //crates/BUILD
Downloads run in parallel across worker threads (--jobs=N to bound them).
Using the fetched crates
Depend on a generated target from your own BUILD:
load("@builtin//:rust.bzl", "rust_binary")
rust_binary(
name = "app",
crate_root = "src/main.rs",
srcs = ["src/main.rs"],
edition = "2021",
out = "app",
deps = ["//crates:serde"],
)
Then zut build //:app runs entirely offline — the sources are grafted from the
CAS Trees recorded earlier.
For the full design — URL construction, the manifest format, the generated model, and the cargo-as-toolchain decision — see
docs/design/crates-io.md.
The command line
zut’s CLI is a small set of commands dispatched from a single table, so zut help is always in sync with what the binary actually does.
$ zut help
Commands
zut build <label> [options]
Evaluate the BUILD file, analyze the target’s transitive graph, run the
resulting actions through the cache, and materialize the target’s declared
outputs into ./zut-out/.
$ zut build //:hello
$ zut build //crates:serde --jobs=4
zut test <label> [options]
Build <label>, then run its (test) binary as a cached action. An unchanged,
previously-passing test is a cache hit — it is not re-run. Accepts the same
options as build.
$ zut test //:mylib_test
//:mylib_test: PASS (cached)
test result: ok. 12 passed; 0 failed
zut fetch [Cargo.lock] [--jobs=N]
Download, checksum-verify, and unpack crates.io dependencies into the CAS, write
.zut/crates.lock, and generate //crates/BUILD. This is the only networked
command. See Fetching crates.io dependencies.
Options
These apply to build and test (and fetch honors --jobs):
| Option | Meaning | Default |
|---|---|---|
--sandbox=prep|namespace|landlock | Hermeticity tier | strongest available |
--jobs=N | Parallel actions | CPU count |
--profile | Per-action wall/cpu/peak-rss table | off |
--remote-cache=<spec> | Back the caches with a shared store | none |
--remote-cache-mode=read|read-write | Remote cache access | read-write |
--log-level=error|warn|info|debug|trace | Diagnostics verbosity | info |
--verbose | Shorthand for --log-level=debug | — |
--quiet | Shorthand for --log-level=off | — |
Every option can also be set in a .zutrc file using the
same name, so the command line stays short.
Labels
A label names a target:
//:name— targetnamein the workspace-rootBUILD.//path/to/pkg:name— targetnameinpath/to/pkg/BUILD.//crates:serde— a generated crates.io target.
Exit status
zut exits non-zero on any failure (missing target, evaluation error, action
failure, test failure). Errors are printed to stderr; machine-readable output is
reserved for stdout.
Configuration
Every build option can be set in three places. They layer, lowest precedence first:
built-in defaults → ~/.zutrc → ./.zutrc → command-line flags
So a project .zutrc overrides your personal ~/.zutrc, and an explicit flag
always wins. Each option is defined exactly once internally — the flag parser
and the .zutrc loader funnel through the same code — so a .zutrc key and its
flag always have the same name and meaning.
The .zutrc file
A flat key = value file, # for comments. Keys match the flag names (without
the leading --); a bare boolean is written key = true.
# ~/.zutrc or ./.zutrc
sandbox = namespace
jobs = 8
remote-cache = s3://my-bucket?region=eu-west-3
remote-cache-mode = read-write
log-level = info
A malformed or unknown value in .zutrc is advisory: zut warns and ignores
it (config never hard-fails a build). An unknown or malformed command-line
flag, by contrast, is an error — typos on the CLI shouldn’t pass silently.
Reference
| Key / flag | Values | Default | Notes |
|---|---|---|---|
sandbox | prep, namespace, landlock | strongest available | Sandboxing |
jobs | positive integer | CPU count | parallel actions |
profile | bool | false | per-action timing table |
remote-cache | a backend spec | none | cmd:/fs:/s3:/gs: |
remote-cache-mode | read, read-write | read-write | |
log-level | error/warn/info/debug/trace/off | info | Logging |
verbose | bool | false | alias for log-level = debug |
quiet | bool | false | alias for log-level = off |
Environment variables
ZUT_LOG— sets the initial log level before any.zutrc/flag is applied (handy for one-off debugging:ZUT_LOG=trace zut build //:x).NO_COLOR— disables colored output (also auto-disabled when stderr is not a terminal).- AWS / GCS credential variables for the remote cache — see Remote caching.
Logging & diagnostics
zut writes to two channels, both on stderr (stdout is reserved for machine-readable output):
- Program output — the build report, help text, the fetch summary. This is the command’s product and is always printed.
- Diagnostics — errors, warnings, and trace messages, gated by a log level.
Levels
From least to most verbose:
| Level | Shows | Use |
|---|---|---|
off (silent, none) | nothing — not even errors | scripting where you only check the exit code |
error | errors | quiet CI |
warn | + warnings | |
info (default) | + informational notes | normal use |
debug | + resolved config, exit detail | troubleshooting your setup |
trace | + fine-grained internal steps | debugging zut itself |
Each level includes everything above it. Diagnostics are tagged (error: ,
warning: , debug: ) and colored when stderr is a color terminal — colors are
suppressed automatically under NO_COLOR, on a pipe, or on redirection.
Setting the level
Three equivalent ways, in increasing precedence:
$ ZUT_LOG=debug zut build //:x # environment, applied first
$ echo 'log-level = debug' >> .zutrc # config file
$ zut build //:x --log-level=debug # flag, wins
$ zut build //:x --verbose # shorthand for --log-level=debug
$ zut build //:x --quiet # shorthand for --log-level=off
Example
$ zut build //:hello --verbose
debug: config: sandbox=namespace jobs=null remote-cache=null (read_write)
zut build //:hello
...
At the default info level you’d see only the build report; --verbose adds the
debug: lines showing the resolved configuration and, on failure, the underlying
error name.
For contributors
The logger is a standalone module (src/log.zig, depends only on std). Inside
the codebase you never call std.debug.print directly — you use:
const log = @import("log");
log.out(" -> {s}\n", .{path}); // always-on program output (verbatim)
log.err("cannot read '{s}'", .{p}); // tagged, gated diagnostics
log.warn(...); log.info(...); log.debug(...); log.trace(...);
log.out is a verbatim drop-in for the old prints (caller supplies the
newline); the leveled functions add the tag and a trailing newline and respect
the global threshold.
Remote caching
A remote cache lets a team (or your CI) share build results. Because zut is content-addressed, sharing is safe by construction: an action’s key is the hash of its inputs, so a result computed on one machine is valid on any other.
Enable it with --remote-cache=<spec> (or remote-cache = <spec> in
.zutrc):
$ zut build //:app --remote-cache=s3://my-bucket?region=eu-west-3
...
remote cache (read_write): 7 hit, 2 miss, 2 uploaded, 0 error
The remote cache is a tier on top of the local store: reads are read-through (miss locally → fetch from remote → populate local), writes are write-through. A misconfigured or unreachable remote degrades gracefully to local-only rather than failing the build, and every blob read back from the remote is digest-verified before use.
Modes
--remote-cache-mode=read-write(default) — read from and upload to the remote.--remote-cache-mode=read— read only (typical for untrusted CI or developer machines that should consume, not publish).
Backends
The scheme in the spec selects the backend:
fs:// — a shared directory
fs:///mnt/nfs/zut-cache
A plain directory (NFS mount, local path). Great for a LAN or for testing.
s3:// — native S3 (and S3-compatible)
s3://my-bucket?region=eu-west-3
s3://my-bucket?region=us-east-1&endpoint=https://minio.local:9000
Speaks the S3 REST protocol directly with hand-rolled AWS SigV4 signing — no AWS
SDK or CLI required. Works against AWS, MinIO, Cloudflare R2, and other
S3-compatible stores via endpoint=.
Credentials follow the standard AWS provider chain:
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_SESSION_TOKEN~/.aws/credentialsand~/.aws/config(honoringAWS_PROFILE)
Region comes from the ?region= query param, else AWS_REGION /
AWS_DEFAULT_REGION.
gs:// — native Google Cloud Storage
gs://my-bucket
gcs://my-bucket
Speaks the GCS XML API with an OAuth2 bearer token from GOOGLE_OAUTH_TOKEN
(e.g. export GOOGLE_OAUTH_TOKEN=$(gcloud auth print-access-token)).
cmd: — bring your own store
cmd:/usr/local/bin/my-cache-helper
zut shells out to your helper with get|put|has <key>, streaming blobs over
stdin/stdout. This is the plug-and-play escape hatch: back the cache with
Redis, a database, an internal artifact service — anything — by writing a small
script. No zut changes required.
On the wire
Blobs are wrapped in a small self-describing envelope (objframe) — a magic
tag, flags, optional metadata, and the payload — with transparent deflate
compression above a size threshold. The envelope lives on zut’s side; backends
only ever move opaque bytes, which keeps the S3/GCS clients general-purpose.
For the tiering semantics, the envelope format, and the SDK boundary, see
docs/design/remote-cache.md.
Sandboxing tiers
Hermeticity is enforced, not trusted. Every action runs inside a sandbox that makes its declared inputs available and undeclared ones unreachable — so a forgotten dependency fails the build instead of silently working on your machine and breaking on someone else’s.
zut picks the strongest tier the host supports by default. Override with
--sandbox=<tier> or sandbox = <tier> in .zutrc.
The tiers (a capability ladder)
namespace — the hermetic-leaning tier (Linux)
Runs the action inside fresh user + mount + pid + net + ipc + uts namespaces:
- user ns — your uid/gid map to root inside, so zut can mount without real privileges (works wherever unprivileged user namespaces are allowed);
- mount ns — the host
/is bind-mounted recursively read-only (so/home,/run,/dev/shmcan’t be written either); the execroot is an overlayfs (materialized inputs as the read-only lower, a writable upper for the action’s outputs);pivot_rootthen makes the contained tree the root; - pid ns — the action runs as PID 1 under a thin reaper that also enforces
timeout_ns; - net ns — no network (builds can’t fetch);
- ipc/uts ns — isolated SysV IPC and hostname.
Result: the action can read the toolchain, its inputs are tamper-proof, and it
can only write to its overlay upper and a private /tmp.
landlock — write-containment (Linux 5.13+)
Uses the Landlock LSM to confine writes to the work directory while allowing reads, without namespaces. A good fit where unprivileged user namespaces are disabled but Landlock is available.
prep — preparation only (portable, non-hermetic)
Materializes the action’s declared inputs into a clean work directory (so the inputs are right) but does not isolate the action from the rest of the filesystem. This is the fallback on non-Linux hosts and the weakest tier — use it only when the stronger tiers are unavailable.
Choosing a tier
$ zut build //:app # strongest available (recommended)
$ zut build //:app --sandbox=namespace # force the namespace tier
$ zut build //:app --sandbox=prep # opt out of isolation (debugging)
The chosen tier is printed in the build header:
sandbox: namespace (recursive ro host, net+pid isolated)
Known limitations
Full input hermeticity still reads the host toolchain (it isn’t a declared, content-addressed input yet); recursive read-only makes that tamper-proof but not yet reproducible across hosts. This — plus seccomp hardening and the Landlock setup-status pipe — is tracked on the roadmap.
The full design, including the threat model and the syscall-level details, is in
docs/ARCHITECTURE.md§7.
BUILD files & Starlark
zut workspaces are described in BUILD files written in Starlark — the same
Python-like configuration dialect used by Bazel and Buck2. zut implements its own
Starlark front-end (lexer, parser, evaluator) in Zig.
Targets, rules, packages
- A package is a directory with a
BUILDfile. - A target is a named instance of a rule declared in that
BUILDfile. - A label addresses a target:
//pkg:name(or//:nameat the root).
load("@builtin//:rust.bzl", "rust_binary")
rust_binary( # the rule
name = "hello", # → label //:hello
crate_root = "src/main.rs",
srcs = ["src/main.rs"],
edition = "2021",
out = "hello",
)
load() and .bzl modules
Reusable definitions (rules, macros, constants) live in .bzl files and are
imported with load():
load("//rules:my_rules.bzl", "my_rule") # a workspace .bzl
load("@builtin//:rust.bzl", "rust_binary", lib = "rust_library") # built-in, with rename
//rules:my_rules.bzl— a.bzlin your workspace.@builtin//:rust.bzl— a module from the built-in prelude that ships inside the zut binary.
load() symbols can be renamed (lib = "rust_library") to avoid clashes.
The Starlark dialect
zut supports the core of Starlark: def functions, if/for, lists, dicts,
strings (with .format(), slicing), comprehensions, struct(...),
and the build-specific builtins below. It is deterministic by design — no clocks,
no randomness, no I/O from Starlark itself; side effects happen only through
declared actions.
Build-specific builtins you’ll use inside rule implementations:
| Builtin | Purpose |
|---|---|
rule(implementation, attrs) | define a rule |
attrs.string(), attrs.label(), … | declare a rule’s attributes |
provider(...) / DefaultInfo(...) | define/return providers |
ctx.actions.declare_file(name) | declare an output file |
ctx.actions.run(executable, arguments, inputs, outputs, env) | register an action |
crate_tree("<name> <ver>", mount) | graft a fetched crates.io source Tree |
Cross-package dependencies are resolved lazily: zut loads a dependency’s BUILD
only when a target actually needs it.
Next: rules & providers.
Rules & providers
Rules in zut are definitions over primitives, not engine built-ins. A rule
says how to turn attributes into actions (commands that produce files) and
what providers (typed results) it hands to the targets that depend on it.
rust_binary and friends are written this way in the prelude —
and you can write your own the same way.
Anatomy of a rule
def _my_tool_impl(ctx):
# 1. Declare outputs.
out = ctx.actions.declare_file(ctx.attr.out)
# 2. Register the action that produces them.
ctx.actions.run(
executable = "/bin/sh",
arguments = ["-c", "my-tool " + ctx.attr.input + " > " + out.path],
inputs = [ctx.attr.input],
outputs = [out],
env = {"PATH": "/usr/bin:/bin"},
)
# 3. Return providers for downstream targets.
return DefaultInfo(files = [out])
my_tool = rule(
implementation = _my_tool_impl,
attrs = {
"input": attrs.string(),
"out": attrs.string(),
},
)
A target then instantiates it:
my_tool(name = "thing", input = "data.in", out = "data.out")
ctx — the rule context
Inside an implementation function, ctx exposes:
ctx.attr.<name>— the target’s attribute values.ctx.actions.declare_file(name)— declare an output; returns a file value with a.path.ctx.actions.run(executable, arguments, inputs, outputs, env)— register an action.inputsmay be files, declared outputs of dependencies, or grafted Trees;outputsare the declared files it produces.ctx.toolchain.rust/ctx.toolchain.zig— the discovered host toolchain (compiler path + environment) for that language.
Actions
An action is the cacheable unit: a command, its input file set, its declared outputs, and its environment. zut hashes all of that into the action key. If the key is already in the cache (locally or remote), the action doesn’t run — its outputs are served. Inputs not declared here are unavailable at run time thanks to the sandbox, which is what makes the cache key sound.
Providers
Providers are the typed values a rule returns for its dependents to consume:
DefaultInfo(files = [...])— the conventional “these are my output files” provider (whatzut buildmaterializes).struct(field = value, ...)— an ad-hoc provider. The prelude’srust_library, for example, returnsstruct(files=..., crate_name=..., rlib=..., rlibs=..., dylibs=...)so a downstreamrust_binarycan wire--externand the transitive rlib closure.provider(...)— define a named provider type for stronger contracts.
A dependent reads a provider’s fields off the dependency value it receives via
its deps-style attributes — that’s how rust_binary discovers the .rlib of
each rust_library it links.
Next: the built-in prelude.
The built-in prelude
zut ships a set of canonical Rust and Zig rules inside the binary. Load them
from the @builtin// namespace — no files to vendor, and a workspace file named
rust.bzl can never shadow the built-in one:
load("@builtin//:rust.bzl", "rust_binary", "rust_library", "rust_test")
load("@builtin//:zig.bzl", "zig_library")
These rules are ordinary Starlark (see the source:
src/rules/rust.bzl,
src/rules/zig.bzl)
— read them as worked examples of writing rules. The
host toolchain comes from ctx.toolchain.rust / ctx.toolchain.zig.
Rust rules
rust_library
Compiles an .rlib and propagates the transitive rlib/dylib closure for
downstream linking.
| Attribute | Meaning |
|---|---|
crate_name | the crate’s import name |
crate_root | entry .rs compiled by rustc |
srcs | all source files (materialized so mod resolves) |
edition | e.g. "2021" |
deps | other rust_library targets (linked via --extern + transitive rlibs) |
macros | rust_proc_macro targets (loaded at compile time) |
native_deps | zig_library (or similar) targets, linked as static libs |
build_script | optional rust_build_script (its cfg directives + OUT_DIR codegen) |
rust_binary
Same dependency model as rust_library, plus an out attribute for the
executable name. Returns DefaultInfo.
rust_test
Like rust_binary but compiled with --test; run as a cached action by
zut test.
Vendored crates.io rules
crate_library and crate_proc_macro are what zut fetch
generates into //crates/BUILD. Their source comes from a checksum-pinned CAS
Tree grafted with src_tree = crate_tree("<name> <ver>", mount), and features
become --cfg feature="x". You normally depend on these (//crates:serde)
rather than writing them by hand.
Zig rules
zig_library
Builds a static library (lib<name>.a) with zig build-lib -fPIC, ready for a
rust_binary to link via native_deps.
| Attribute | Meaning |
|---|---|
src | the Zig root source file |
libname | base name (→ lib<libname>.a, linked as -l static=<libname>) |
A combined example
load("@builtin//:rust.bzl", "rust_binary")
load("@builtin//:zig.bzl", "zig_library")
zig_library(name = "mathz", src = "math.zig", libname = "mathz")
rust_binary(
name = "app",
crate_root = "src/main.rs",
srcs = ["src/main.rs"],
edition = "2021",
out = "app",
native_deps = [":mathz"], # Rust ↔ Zig via the C ABI
deps = ["//crates:serde"], # a fetched crates.io dependency
)
This is the headline feature: Rust and Zig composed as peers in one graph, built hermetically and cached per action.
Architecture overview
This section is a tour of how zut works inside, for contributors and the
curious. The authoritative, detailed design contract is
docs/ARCHITECTURE.md
in the repository — this is the friendly companion.
Design goals
- Correctness by construction — output is a pure function of declared inputs; undeclared inputs are physically unavailable (the sandbox).
- Fine-grained incrementality — the cache unit is the action.
- Speed — parallel by default, content-addressed caching, early cutoff.
- Rust & Zig as peers — linked through the C ABI, neither a bolt-on.
- Remote-ready — the execution path is shaped like the Bazel Remote Execution API (REAPI), so remote caching/execution slot in cleanly.
- Extensible — rules are Starlark definitions over primitives.
The shape of a build
BUILD / .bzl ──► unconfigured graph ──► configured graph ──► actions ──► outputs
(Starlark) (what targets exist) (select() resolved) (cacheable)
Loading, configuration, and analysis are memoized as keys in the incremental engine (DICE); execution is keyed separately by the content-addressed action cache. The next two chapters unpack these:
- The five layers — from
BUILDtext to materialized outputs. - Content addressing & the CAS — the hashing substrate everything stands on.
And the code map:
- Repo & source layout — which directory does what, and which way the dependencies point.
The five layers
A build flows through five layers. Each is (conceptually) a pure function of the previous one, which is what makes the whole pipeline memoizable.
BUILD / .bzl files
│ Starlark evaluation
▼
Unconfigured target graph nodes = (rule, attrs, dep labels) — "what exists"
│ apply configuration (platform, flags, select())
▼
Configured target graph select() resolved, toolchains chosen
│ analysis: run each rule's implementation
▼
Action graph commands + declared inputs/outputs (the cache unit)
│ execution through the action cache + sandbox
▼
Outputs content-addressed files, materialized into zut-out/
-
Loading — the Starlark front-end (
src/starlark/) lexes, parses, and evaluatesBUILD/.bzlfiles.load()pulls in.bzlmodules (including the@builtin//prelude). The result is the unconfigured target graph: which targets exist and how they refer to each other by label. -
Configuration — applies the build configuration (platform, flags,
select()) to produce the configured graph. Toolchains are resolved here. -
Analysis — runs each rule’s
implementationfunction. This is wherectx.actions.run(...)registers actions and rules return providers. The output is the action graph: a DAG of commands, each with a declared input set and declared outputs. -
Execution — the executor runs actions whose results aren’t already cached, in parallel up to
--jobs, each inside the chosen sandbox tier. The result of each action is stored by content. -
Materialization — the requested target’s declared outputs are copied out of the content-addressed store into
./zut-out/.
Two kinds of memoization
zut layers two distinct caches, and the distinction matters:
-
DICE keys memoize loading, configuration, and analysis — the pure, in-process computation graph. These are versioned: when an input changes, DICE invalidates exactly the dependent keys (with early cutoff — if a recomputation produces the same value, dependents don’t recompute).
-
The action cache memoizes execution — keyed by the content hash of an action’s inputs, not by a DICE version. This is what survives across runs and across machines (via the remote cache), because a content hash is machine-independent.
The incremental engine lives in src/dice/; the executor and action-cache
orchestration in src/exec/.
Content addressing & the CAS
Everything in zut is addressed by the hash of its bytes. This is the substrate that makes caching correct and sharing safe.
Digests
A digest is the BLAKE3 hash of some bytes plus their length. zut’s digest
type and hashing live in src/core/digest.zig. The hash function is recorded in
a REAPI-compatible way (the CLI banner shows digest: blake3, REAPI #9), so the
store knows which function produced a given address.
The CAS (content-addressed store)
The CAS (src/core/cas.zig) is a key→bytes store where the key is the
content’s digest. Properties that fall out of that:
- Deduplication — identical bytes are stored once, regardless of how many actions produce or consume them.
- Integrity — a blob read back is verified against the digest it was asked for; corruption or a lying remote is detected.
- Immutability — an address never changes meaning, which is why a cached result computed elsewhere is trustworthy here.
The CAS is tiered: a read misses locally, falls through to the remote cache if configured, verifies the blob, and populates the local tier. Writes go through to the remote (write-through). A remote failure degrades to local-only rather than breaking the build.
Trees (Merkle directories)
A directory is content-addressed as a Tree (src/core/merkle.zig): a node
listing names → child digests (files or sub-trees). The digest of a Tree is
therefore a hash of its entire recursive content. This is how:
- a crate’s unpacked source is grafted into a build as a single immutable input
(
crate_tree(...)), and - an action’s input set is named by one root digest.
The action cache
An action (src/core/action_cache.zig) is keyed by the digest of its
command + input Tree + environment. The action cache maps that key to the action
result (output digests, exit code, captured stdout/stderr). Because the key is
a content hash:
- the same action never runs twice, and
- the cache is valid across machines, which is exactly what the remote cache exploits.
This shape mirrors the Bazel Remote Execution API, so the same store backs both local incrementality and remote caching without a second code path.
See
docs/ARCHITECTURE.md§3 for the precise encodings and the REAPI mapping.
Repo & source layout
Each top-level directory under src/ is a layer with one job, and
dependencies point downward only.
src/
main.zig / root.zig CLI + library root — the only place that wires layers together
│
├── starlark/ the Starlark front-end (lexer → parser → eval → rules)
├── crates/ the crates.io importer (lock, fetch, metadata, gen)
├── exec/ the sandboxed executor + cache orchestration
├── dice/ the incremental engine
├── graph/ target identity (labels) and the build graph
├── core/ the content-addressed substrate: digest, codec, cas,
│ merkle, action_cache
├── cache/ the remote-cache tier on top of core: remote.zig
│ (RemoteStore tiering + backends), objframe.zig (envelope)
├── rules/ the built-in @builtin prelude (rust.bzl, zig.bzl)
├── aws/ STANDALONE AWS SDK — sigv4, credentials, s3 (own module)
├── gcs/ STANDALONE GCS SDK — storage (own module)
└── log.zig STANDALONE leveled logger (own module, std-only)
Why some pieces are standalone modules
aws, gcs, and log are separate build modules, not just directories.
That gives a compiler-enforced boundary: they cannot reach into zut internals,
so they stay reusable on their own.
-
aws/gcsare the seeds of general-purpose Zig SDKs (aws.s3.Client,gcs.storage.Client) that have nothing to do with build systems. zut’s remote cache is a thin adapter over them; framing/compression (objframe) stays on zut’s side so the SDKs only move opaque bytes. Each can later be extracted to its own package. -
logdepends on nothing butstd, so every layer imports it as@import("log")— no../log.zigreaching across directories — and there is exactly one module instance, hence one process-global log threshold.
The dependency arrows
The front-end and importer sit on top; core/ is the content-addressed
substrate everything stands on; cache/ builds the remote tier on core/ using
the SDKs as backends without leaking them. Nothing imports main.
The full rationale and the migration history are in
docs/design/layout.md.
Contributing
Contributions are welcome. This page is the quick version; the canonical, more
detailed guide is
CONTRIBUTING.md in
the repository.
Build & test
$ zig build # build the CLI → zig-out/bin/zut
$ zig build test --summary all # unit tests + sandbox smoketests
You need Zig 0.16.0 exactly. See Installation.
Conventions
- No raw
std.debug.print. Use the logger:@import("log")thenlog.outfor program output,log.err/warn/info/debug/tracefor diagnostics. - Tests live with their code. When a source file grows large, its tests move
to a sibling
<name>_test.zig, pulled in from the original viatest { _ = @import("<name>_test.zig"); }. White-box tests drive code through its public API. - Layers point downward. Respect the source layout;
aws/gcs/logare standalone modules that must not reach into zut. - Match the surrounding code — comment density, naming, and idiom. zut’s source is heavily commented with the why; keep that up.
- Every change keeps the suite green and adds tests for new behavior.
Where to start
- Good entry points are tracked on the roadmap (Phase 2/3 work).
- Read
docs/ARCHITECTURE.mdfirst — it’s the design contract and explains why things are shaped the way they are. - Small, reviewable commits with a clear message are preferred; pure moves (e.g. splitting a file) should carry no logic change.
Roadmap
zut is built in phases, each ending in a demoable, tested capability. This is
a summary; the full plan with rationale lives in
docs/ROADMAP.md.
| Phase | Theme | Status |
|---|---|---|
| 0 | Vertical slice — rust_binary ← zig_library over the C ABI, no interpreter | ✅ done |
| 1 | Starlark front-end — real BUILD/.bzl evaluation, rules as definitions | ✅ done |
| 2 | Rust & Zig depth — proc-macros, build scripts, crates.io importer | in progress |
| 3 | Configuration & query — select(), platforms, configuration transitions | planned |
| 4 | Remote cache & execution — cash in the REAPI shape | remote cache ✅, execution planned |
| 5 | Daemon & polish — long-lived server, structured events, great errors | planned |
Cross-cutting, every phase: hermeticity is enforced not trusted; the execution path stays REAPI-shaped; correctness invariants are tested as they’re built.
If you’d like to help with any of this, see Contributing.