Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

zut is a hermetic, incremental, content-addressed build system written in Zig 0.16, with Rust and Zig as first-class languages. It is built from scratch — its own Starlark front-end, content-addressed store, sandboxed executor, and incremental engine — in the spirit of Bazel and Buck2.

⚠️ Status: early and experimental. zut is under active development. The design is settled (see the Architecture overview), but APIs, the Starlark dialect, and the CLI will change. It is not yet ready for production builds.

What makes zut zut

  • Correct by construction. A build’s output is a pure function of its declared inputs. Undeclared inputs are made physically unavailable to actions by the sandbox, so “works on my machine” drift becomes a build error instead of a mystery.
  • Fine-grained incrementality. The unit of caching is the action, not the target or the crate. Change one file and zut rebuilds exactly the actions that depended on it — nothing more.
  • Content-addressed caching. Every action’s inputs and outputs are addressed by their hash (BLAKE3). An identical action never runs twice, whether the result is in your local store or a shared remote cache.
  • Rust & Zig as peers. A rust_binary can depend on a zig_library and vice versa, linked through the C ABI. Neither language is a bolt-on.
  • Remote-ready. The execution path is shaped like the Bazel Remote Execution API (REAPI), so remote caching (today) and remote execution (later) slot in without reworking the core.
  • Extensible. Build rules like rust_binary and zig_library are definitions written in Starlark over a small set of primitives (providers + actions) — not hardcoded into the engine.

How this book is organized

  • Getting started — build zut, write a BUILD file, run your first build, and import crates.io dependencies.
  • User guide — the CLI, configuration (.zutrc), logging, remote caching, and the sandbox tiers.
  • Writing BUILD files — the Starlark dialect, rules and providers, and the built-in Rust/Zig prelude.
  • Internals — how the engine fits together, for contributors and the curious.
  • Project — how to contribute and where things are headed.

The authoritative design contract is docs/ARCHITECTURE.md in the repository; this book is the practical companion to it.

Installation

zut is built from source with the Zig toolchain. There are no other build-time dependencies — the TLS stack, the S3/GCS protocol clients, and the sandbox are all implemented in-tree.

Prerequisites

  • Zig 0.16.0 (exactly — zut tracks the 0.16 std.Io API). Get it from ziglang.org/download.
  • Linux for the strong sandbox tiers (user namespaces and/or Landlock). Other platforms build and run, but fall back to a weaker preparation-only sandbox.
  • Optional, for building Rust/Zig targets: a host rustc/cargo and/or zig. zut discovers them at runtime; they are not needed to build zut itself.

Build

$ git clone https://github.com/Sh4d1/zut
$ cd zut
$ zig build               # debug build → zig-out/bin/zut
$ zig build -Doptimize=ReleaseFast   # optimized build

The binary lands at zig-out/bin/zut. Put it on your PATH, or run it through the build system:

$ zig build run -- build //:hello

Run the tests

$ zig build test --summary all

This runs the unit-test suite plus the namespace/Landlock sandbox smoketests (which fork real processes, so they live outside the in-process test harness).

Verify it works

$ zig-out/bin/zut
zut 0.0.0-dev — a hermetic, incremental build system (digest: blake3, REAPI #9)

Commands:
  build   build a target and materialize its outputs into ./zut-out/
  test    build a target and run its test binary (a cached action)
  fetch   download crates.io deps into the CAS + generate //crates/BUILD

Run `zut help` for details.

Next: write your first build.

Your first build

A zut workspace is a directory tree with BUILD files. Each BUILD file declares targets; a target is an instance of a rule (like rust_binary) with some attributes.

A minimal example

Let’s build a Rust binary that links a Zig static library. Create this layout:

myproject/
  BUILD
  src/main.rs
  greet.zig

greet.zig — a tiny C-ABI function:

export fn zut_add(a: i32, b: i32) callconv(.c) i32 {
    return a + b;
}

src/main.rs — calls into it:

extern "C" { fn zut_add(a: i32, b: i32) -> i32; }

fn main() {
    println!("2 + 3 = {}", unsafe { zut_add(2, 3) });
}

BUILD — wire them together using the built-in prelude:

load("@builtin//:rust.bzl", "rust_binary")
load("@builtin//:zig.bzl", "zig_library")

zig_library(
    name = "greet",
    src = "greet.zig",
    libname = "greet",
)

rust_binary(
    name = "hello",
    crate_root = "src/main.rs",
    srcs = ["src/main.rs"],
    edition = "2021",
    out = "hello",
    native_deps = [":greet"],
)

Build it

$ zut build //:hello
zut build //:hello
  sandbox: namespace (recursive ro host, net+pid isolated)
  toolchain: rust[1.xx] zig[0.16.0]
  jobs: 8
  libgreet.a                   [zig]  built
  hello                        [rustc]  built
  -> zut-out/hello
  done — 2 built, 0 cached

$ ./zut-out/hello
2 + 3 = 5

The label //:hello means “the target named hello in the BUILD file at the workspace root”. A target in a subdirectory foo/bar is //foo/bar:name.

Incrementality in action

Run the same command again, unchanged:

$ zut build //:hello
  ...
  up to date — 2 action(s), all cached

Nothing rebuilds: zut hashed the action inputs, found the outputs already in its content-addressed store, and served them. Now touch only the Rust source and rebuild — only the rustc action re-runs; the Zig library is still a cache hit, because its inputs didn’t change.

Build state lives under .zut/ in your workspace (the CAS, the action cache, and a work directory). Delete .zut/ to start cold; it is safe to .gitignore.

Next: import crates.io dependencies.

Fetching crates.io dependencies

zut builds are hermetic: the build phase has no network access. So fetching third-party crates is a separate, explicit step — zut fetch — that runs once, downloads everything into the content-addressed store, and generates a BUILD file the hermetic build then consumes offline.

The workflow

Point zut fetch at a Cargo lockfile:

$ zut fetch Cargo.lock
zut fetch Cargo.lock
  fetch  serde 1.0.210
  fetch  serde_derive 1.0.210
  ...
  42 crate(s): 42 fetched, 0 reused → .zut/crates.lock + //crates/BUILD

This does four things:

  1. Downloads each registry crate’s .crate archive over HTTPS (the only networked step — zut ships its own TLS via std.http.Client).
  2. Checksum-verifies each archive against the lockfile’s SHA-256, then unpacks it into a CAS Tree (a content-addressed directory).
  3. Records a manifest at .zut/crates.lock mapping each name version to its Tree digest.
  4. Generates //crates/BUILD — crate_library / crate_proc_macro targets wired with the right editions, features, and dependency edges.

When cargo and the host triple are available, zut runs cargo metadata for a faithful model (real editions, resolved features, lib/proc-macro kinds). Without them it falls back to a lockfile-only model (edition 2021, no features).

Idempotent & incremental

zut fetch is safe to re-run. A crate whose Tree is already in the CAS (per the prior .zut/crates.lock) is reused, not re-downloaded:

$ zut fetch Cargo.lock
  reuse  serde 1.0.210
  ...
  42 crate(s): 1 fetched, 41 reused → .zut/crates.lock + //crates/BUILD

Downloads run in parallel across worker threads (--jobs=N to bound them).

Using the fetched crates

Depend on a generated target from your own BUILD:

load("@builtin//:rust.bzl", "rust_binary")

rust_binary(
    name = "app",
    crate_root = "src/main.rs",
    srcs = ["src/main.rs"],
    edition = "2021",
    out = "app",
    deps = ["//crates:serde"],
)

Then zut build //:app runs entirely offline — the sources are grafted from the CAS Trees recorded earlier.

For the full design — URL construction, the manifest format, the generated model, and the cargo-as-toolchain decision — see docs/design/crates-io.md.

The command line

zut’s CLI is a small set of commands dispatched from a single table, so zut help is always in sync with what the binary actually does.

$ zut help

Commands

zut build <label> [options]

Evaluate the BUILD file, analyze the target’s transitive graph, run the resulting actions through the cache, and materialize the target’s declared outputs into ./zut-out/.

$ zut build //:hello
$ zut build //crates:serde --jobs=4

zut test <label> [options]

Build <label>, then run its (test) binary as a cached action. An unchanged, previously-passing test is a cache hit — it is not re-run. Accepts the same options as build.

$ zut test //:mylib_test
  //:mylib_test: PASS (cached)
    test result: ok. 12 passed; 0 failed

zut fetch [Cargo.lock] [--jobs=N]

Download, checksum-verify, and unpack crates.io dependencies into the CAS, write .zut/crates.lock, and generate //crates/BUILD. This is the only networked command. See Fetching crates.io dependencies.

Options

These apply to build and test (and fetch honors --jobs):

OptionMeaningDefault
--sandbox=prep|namespace|landlockHermeticity tierstrongest available
--jobs=NParallel actionsCPU count
--profilePer-action wall/cpu/peak-rss tableoff
--remote-cache=<spec>Back the caches with a shared storenone
--remote-cache-mode=read|read-writeRemote cache accessread-write
--log-level=error|warn|info|debug|traceDiagnostics verbosityinfo
--verboseShorthand for --log-level=debug—
--quietShorthand for --log-level=off—

Every option can also be set in a .zutrc file using the same name, so the command line stays short.

Labels

A label names a target:

  • //:name — target name in the workspace-root BUILD.
  • //path/to/pkg:name — target name in path/to/pkg/BUILD.
  • //crates:serde — a generated crates.io target.

Exit status

zut exits non-zero on any failure (missing target, evaluation error, action failure, test failure). Errors are printed to stderr; machine-readable output is reserved for stdout.

Configuration

Every build option can be set in three places. They layer, lowest precedence first:

built-in defaults  →  ~/.zutrc  →  ./.zutrc  →  command-line flags

So a project .zutrc overrides your personal ~/.zutrc, and an explicit flag always wins. Each option is defined exactly once internally — the flag parser and the .zutrc loader funnel through the same code — so a .zutrc key and its flag always have the same name and meaning.

The .zutrc file

A flat key = value file, # for comments. Keys match the flag names (without the leading --); a bare boolean is written key = true.

# ~/.zutrc or ./.zutrc
sandbox      = namespace
jobs         = 8
remote-cache = s3://my-bucket?region=eu-west-3
remote-cache-mode = read-write
log-level    = info

A malformed or unknown value in .zutrc is advisory: zut warns and ignores it (config never hard-fails a build). An unknown or malformed command-line flag, by contrast, is an error — typos on the CLI shouldn’t pass silently.

Reference

Key / flagValuesDefaultNotes
sandboxprep, namespace, landlockstrongest availableSandboxing
jobspositive integerCPU countparallel actions
profileboolfalseper-action timing table
remote-cachea backend specnonecmd:/fs:/s3:/gs:
remote-cache-moderead, read-writeread-write
log-levelerror/warn/info/debug/trace/offinfoLogging
verboseboolfalsealias for log-level = debug
quietboolfalsealias for log-level = off

Environment variables

  • ZUT_LOG — sets the initial log level before any .zutrc/flag is applied (handy for one-off debugging: ZUT_LOG=trace zut build //:x).
  • NO_COLOR — disables colored output (also auto-disabled when stderr is not a terminal).
  • AWS / GCS credential variables for the remote cache — see Remote caching.

Logging & diagnostics

zut writes to two channels, both on stderr (stdout is reserved for machine-readable output):

  • Program output — the build report, help text, the fetch summary. This is the command’s product and is always printed.
  • Diagnostics — errors, warnings, and trace messages, gated by a log level.

Levels

From least to most verbose:

LevelShowsUse
off (silent, none)nothing — not even errorsscripting where you only check the exit code
errorerrorsquiet CI
warn+ warnings
info (default)+ informational notesnormal use
debug+ resolved config, exit detailtroubleshooting your setup
trace+ fine-grained internal stepsdebugging zut itself

Each level includes everything above it. Diagnostics are tagged (error: , warning: , debug: ) and colored when stderr is a color terminal — colors are suppressed automatically under NO_COLOR, on a pipe, or on redirection.

Setting the level

Three equivalent ways, in increasing precedence:

$ ZUT_LOG=debug zut build //:x          # environment, applied first
$ echo 'log-level = debug' >> .zutrc    # config file
$ zut build //:x --log-level=debug      # flag, wins
$ zut build //:x --verbose              # shorthand for --log-level=debug
$ zut build //:x --quiet                # shorthand for --log-level=off

Example

$ zut build //:hello --verbose
debug: config: sandbox=namespace jobs=null remote-cache=null (read_write)
zut build //:hello
  ...

At the default info level you’d see only the build report; --verbose adds the debug: lines showing the resolved configuration and, on failure, the underlying error name.

For contributors

The logger is a standalone module (src/log.zig, depends only on std). Inside the codebase you never call std.debug.print directly — you use:

const log = @import("log");

log.out("  -> {s}\n", .{path});   // always-on program output (verbatim)
log.err("cannot read '{s}'", .{p}); // tagged, gated diagnostics
log.warn(...); log.info(...); log.debug(...); log.trace(...);

log.out is a verbatim drop-in for the old prints (caller supplies the newline); the leveled functions add the tag and a trailing newline and respect the global threshold.

Remote caching

A remote cache lets a team (or your CI) share build results. Because zut is content-addressed, sharing is safe by construction: an action’s key is the hash of its inputs, so a result computed on one machine is valid on any other.

Enable it with --remote-cache=<spec> (or remote-cache = <spec> in .zutrc):

$ zut build //:app --remote-cache=s3://my-bucket?region=eu-west-3
  ...
  remote cache (read_write): 7 hit, 2 miss, 2 uploaded, 0 error

The remote cache is a tier on top of the local store: reads are read-through (miss locally → fetch from remote → populate local), writes are write-through. A misconfigured or unreachable remote degrades gracefully to local-only rather than failing the build, and every blob read back from the remote is digest-verified before use.

Modes

  • --remote-cache-mode=read-write (default) — read from and upload to the remote.
  • --remote-cache-mode=read — read only (typical for untrusted CI or developer machines that should consume, not publish).

Backends

The scheme in the spec selects the backend:

fs:// — a shared directory

fs:///mnt/nfs/zut-cache

A plain directory (NFS mount, local path). Great for a LAN or for testing.

s3:// — native S3 (and S3-compatible)

s3://my-bucket?region=eu-west-3
s3://my-bucket?region=us-east-1&endpoint=https://minio.local:9000

Speaks the S3 REST protocol directly with hand-rolled AWS SigV4 signing — no AWS SDK or CLI required. Works against AWS, MinIO, Cloudflare R2, and other S3-compatible stores via endpoint=.

Credentials follow the standard AWS provider chain:

  1. AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_SESSION_TOKEN
  2. ~/.aws/credentials and ~/.aws/config (honoring AWS_PROFILE)

Region comes from the ?region= query param, else AWS_REGION / AWS_DEFAULT_REGION.

gs:// — native Google Cloud Storage

gs://my-bucket
gcs://my-bucket

Speaks the GCS XML API with an OAuth2 bearer token from GOOGLE_OAUTH_TOKEN (e.g. export GOOGLE_OAUTH_TOKEN=$(gcloud auth print-access-token)).

cmd: — bring your own store

cmd:/usr/local/bin/my-cache-helper

zut shells out to your helper with get|put|has <key>, streaming blobs over stdin/stdout. This is the plug-and-play escape hatch: back the cache with Redis, a database, an internal artifact service — anything — by writing a small script. No zut changes required.

On the wire

Blobs are wrapped in a small self-describing envelope (objframe) — a magic tag, flags, optional metadata, and the payload — with transparent deflate compression above a size threshold. The envelope lives on zut’s side; backends only ever move opaque bytes, which keeps the S3/GCS clients general-purpose.

For the tiering semantics, the envelope format, and the SDK boundary, see docs/design/remote-cache.md.

Sandboxing tiers

Hermeticity is enforced, not trusted. Every action runs inside a sandbox that makes its declared inputs available and undeclared ones unreachable — so a forgotten dependency fails the build instead of silently working on your machine and breaking on someone else’s.

zut picks the strongest tier the host supports by default. Override with --sandbox=<tier> or sandbox = <tier> in .zutrc.

The tiers (a capability ladder)

namespace — the hermetic-leaning tier (Linux)

Runs the action inside fresh user + mount + pid + net + ipc + uts namespaces:

  • user ns — your uid/gid map to root inside, so zut can mount without real privileges (works wherever unprivileged user namespaces are allowed);
  • mount ns — the host / is bind-mounted recursively read-only (so /home, /run, /dev/shm can’t be written either); the execroot is an overlayfs (materialized inputs as the read-only lower, a writable upper for the action’s outputs); pivot_root then makes the contained tree the root;
  • pid ns — the action runs as PID 1 under a thin reaper that also enforces timeout_ns;
  • net ns — no network (builds can’t fetch);
  • ipc/uts ns — isolated SysV IPC and hostname.

Result: the action can read the toolchain, its inputs are tamper-proof, and it can only write to its overlay upper and a private /tmp.

landlock — write-containment (Linux 5.13+)

Uses the Landlock LSM to confine writes to the work directory while allowing reads, without namespaces. A good fit where unprivileged user namespaces are disabled but Landlock is available.

prep — preparation only (portable, non-hermetic)

Materializes the action’s declared inputs into a clean work directory (so the inputs are right) but does not isolate the action from the rest of the filesystem. This is the fallback on non-Linux hosts and the weakest tier — use it only when the stronger tiers are unavailable.

Choosing a tier

$ zut build //:app                      # strongest available (recommended)
$ zut build //:app --sandbox=namespace  # force the namespace tier
$ zut build //:app --sandbox=prep       # opt out of isolation (debugging)

The chosen tier is printed in the build header:

  sandbox: namespace (recursive ro host, net+pid isolated)

Known limitations

Full input hermeticity still reads the host toolchain (it isn’t a declared, content-addressed input yet); recursive read-only makes that tamper-proof but not yet reproducible across hosts. This — plus seccomp hardening and the Landlock setup-status pipe — is tracked on the roadmap.

The full design, including the threat model and the syscall-level details, is in docs/ARCHITECTURE.md §7.

BUILD files & Starlark

zut workspaces are described in BUILD files written in Starlark — the same Python-like configuration dialect used by Bazel and Buck2. zut implements its own Starlark front-end (lexer, parser, evaluator) in Zig.

Targets, rules, packages

  • A package is a directory with a BUILD file.
  • A target is a named instance of a rule declared in that BUILD file.
  • A label addresses a target: //pkg:name (or //:name at the root).
load("@builtin//:rust.bzl", "rust_binary")

rust_binary(           # the rule
    name = "hello",    # → label //:hello
    crate_root = "src/main.rs",
    srcs = ["src/main.rs"],
    edition = "2021",
    out = "hello",
)

load() and .bzl modules

Reusable definitions (rules, macros, constants) live in .bzl files and are imported with load():

load("//rules:my_rules.bzl", "my_rule")        # a workspace .bzl
load("@builtin//:rust.bzl", "rust_binary", lib = "rust_library")  # built-in, with rename
  • //rules:my_rules.bzl — a .bzl in your workspace.
  • @builtin//:rust.bzl — a module from the built-in prelude that ships inside the zut binary.

load() symbols can be renamed (lib = "rust_library") to avoid clashes.

The Starlark dialect

zut supports the core of Starlark: def functions, if/for, lists, dicts, strings (with .format(), slicing), comprehensions, struct(...), and the build-specific builtins below. It is deterministic by design — no clocks, no randomness, no I/O from Starlark itself; side effects happen only through declared actions.

Build-specific builtins you’ll use inside rule implementations:

BuiltinPurpose
rule(implementation, attrs)define a rule
attrs.string(), attrs.label(), …declare a rule’s attributes
provider(...) / DefaultInfo(...)define/return providers
ctx.actions.declare_file(name)declare an output file
ctx.actions.run(executable, arguments, inputs, outputs, env)register an action
crate_tree("<name> <ver>", mount)graft a fetched crates.io source Tree

Cross-package dependencies are resolved lazily: zut loads a dependency’s BUILD only when a target actually needs it.

Next: rules & providers.

Rules & providers

Rules in zut are definitions over primitives, not engine built-ins. A rule says how to turn attributes into actions (commands that produce files) and what providers (typed results) it hands to the targets that depend on it. rust_binary and friends are written this way in the prelude — and you can write your own the same way.

Anatomy of a rule

def _my_tool_impl(ctx):
    # 1. Declare outputs.
    out = ctx.actions.declare_file(ctx.attr.out)

    # 2. Register the action that produces them.
    ctx.actions.run(
        executable = "/bin/sh",
        arguments = ["-c", "my-tool " + ctx.attr.input + " > " + out.path],
        inputs = [ctx.attr.input],
        outputs = [out],
        env = {"PATH": "/usr/bin:/bin"},
    )

    # 3. Return providers for downstream targets.
    return DefaultInfo(files = [out])

my_tool = rule(
    implementation = _my_tool_impl,
    attrs = {
        "input": attrs.string(),
        "out": attrs.string(),
    },
)

A target then instantiates it:

my_tool(name = "thing", input = "data.in", out = "data.out")

ctx — the rule context

Inside an implementation function, ctx exposes:

  • ctx.attr.<name> — the target’s attribute values.
  • ctx.actions.declare_file(name) — declare an output; returns a file value with a .path.
  • ctx.actions.run(executable, arguments, inputs, outputs, env) — register an action. inputs may be files, declared outputs of dependencies, or grafted Trees; outputs are the declared files it produces.
  • ctx.toolchain.rust / ctx.toolchain.zig — the discovered host toolchain (compiler path + environment) for that language.

Actions

An action is the cacheable unit: a command, its input file set, its declared outputs, and its environment. zut hashes all of that into the action key. If the key is already in the cache (locally or remote), the action doesn’t run — its outputs are served. Inputs not declared here are unavailable at run time thanks to the sandbox, which is what makes the cache key sound.

Providers

Providers are the typed values a rule returns for its dependents to consume:

  • DefaultInfo(files = [...]) — the conventional “these are my output files” provider (what zut build materializes).
  • struct(field = value, ...) — an ad-hoc provider. The prelude’s rust_library, for example, returns struct(files=..., crate_name=..., rlib=..., rlibs=..., dylibs=...) so a downstream rust_binary can wire --extern and the transitive rlib closure.
  • provider(...) — define a named provider type for stronger contracts.

A dependent reads a provider’s fields off the dependency value it receives via its deps-style attributes — that’s how rust_binary discovers the .rlib of each rust_library it links.

Next: the built-in prelude.

The built-in prelude

zut ships a set of canonical Rust and Zig rules inside the binary. Load them from the @builtin// namespace — no files to vendor, and a workspace file named rust.bzl can never shadow the built-in one:

load("@builtin//:rust.bzl", "rust_binary", "rust_library", "rust_test")
load("@builtin//:zig.bzl", "zig_library")

These rules are ordinary Starlark (see the source: src/rules/rust.bzl, src/rules/zig.bzl) — read them as worked examples of writing rules. The host toolchain comes from ctx.toolchain.rust / ctx.toolchain.zig.

Rust rules

rust_library

Compiles an .rlib and propagates the transitive rlib/dylib closure for downstream linking.

AttributeMeaning
crate_namethe crate’s import name
crate_rootentry .rs compiled by rustc
srcsall source files (materialized so mod resolves)
editione.g. "2021"
depsother rust_library targets (linked via --extern + transitive rlibs)
macrosrust_proc_macro targets (loaded at compile time)
native_depszig_library (or similar) targets, linked as static libs
build_scriptoptional rust_build_script (its cfg directives + OUT_DIR codegen)

rust_binary

Same dependency model as rust_library, plus an out attribute for the executable name. Returns DefaultInfo.

rust_test

Like rust_binary but compiled with --test; run as a cached action by zut test.

Vendored crates.io rules

crate_library and crate_proc_macro are what zut fetch generates into //crates/BUILD. Their source comes from a checksum-pinned CAS Tree grafted with src_tree = crate_tree("<name> <ver>", mount), and features become --cfg feature="x". You normally depend on these (//crates:serde) rather than writing them by hand.

Zig rules

zig_library

Builds a static library (lib<name>.a) with zig build-lib -fPIC, ready for a rust_binary to link via native_deps.

AttributeMeaning
srcthe Zig root source file
libnamebase name (→ lib<libname>.a, linked as -l static=<libname>)

A combined example

load("@builtin//:rust.bzl", "rust_binary")
load("@builtin//:zig.bzl", "zig_library")

zig_library(name = "mathz", src = "math.zig", libname = "mathz")

rust_binary(
    name = "app",
    crate_root = "src/main.rs",
    srcs = ["src/main.rs"],
    edition = "2021",
    out = "app",
    native_deps = [":mathz"],     # Rust ↔ Zig via the C ABI
    deps = ["//crates:serde"],    # a fetched crates.io dependency
)

This is the headline feature: Rust and Zig composed as peers in one graph, built hermetically and cached per action.

Architecture overview

This section is a tour of how zut works inside, for contributors and the curious. The authoritative, detailed design contract is docs/ARCHITECTURE.md in the repository — this is the friendly companion.

Design goals

  • Correctness by construction — output is a pure function of declared inputs; undeclared inputs are physically unavailable (the sandbox).
  • Fine-grained incrementality — the cache unit is the action.
  • Speed — parallel by default, content-addressed caching, early cutoff.
  • Rust & Zig as peers — linked through the C ABI, neither a bolt-on.
  • Remote-ready — the execution path is shaped like the Bazel Remote Execution API (REAPI), so remote caching/execution slot in cleanly.
  • Extensible — rules are Starlark definitions over primitives.

The shape of a build

BUILD / .bzl  ──►  unconfigured graph  ──►  configured graph  ──►  actions  ──►  outputs
   (Starlark)        (what targets exist)     (select() resolved)   (cacheable)

Loading, configuration, and analysis are memoized as keys in the incremental engine (DICE); execution is keyed separately by the content-addressed action cache. The next two chapters unpack these:

And the code map:

The five layers

A build flows through five layers. Each is (conceptually) a pure function of the previous one, which is what makes the whole pipeline memoizable.

  BUILD / .bzl files
        │   Starlark evaluation
        ▼
  Unconfigured target graph     nodes = (rule, attrs, dep labels) — "what exists"
        │   apply configuration (platform, flags, select())
        ▼
  Configured target graph       select() resolved, toolchains chosen
        │   analysis: run each rule's implementation
        ▼
  Action graph                  commands + declared inputs/outputs (the cache unit)
        │   execution through the action cache + sandbox
        ▼
  Outputs                       content-addressed files, materialized into zut-out/
  1. Loading — the Starlark front-end (src/starlark/) lexes, parses, and evaluates BUILD/.bzl files. load() pulls in .bzl modules (including the @builtin// prelude). The result is the unconfigured target graph: which targets exist and how they refer to each other by label.

  2. Configuration — applies the build configuration (platform, flags, select()) to produce the configured graph. Toolchains are resolved here.

  3. Analysis — runs each rule’s implementation function. This is where ctx.actions.run(...) registers actions and rules return providers. The output is the action graph: a DAG of commands, each with a declared input set and declared outputs.

  4. Execution — the executor runs actions whose results aren’t already cached, in parallel up to --jobs, each inside the chosen sandbox tier. The result of each action is stored by content.

  5. Materialization — the requested target’s declared outputs are copied out of the content-addressed store into ./zut-out/.

Two kinds of memoization

zut layers two distinct caches, and the distinction matters:

  • DICE keys memoize loading, configuration, and analysis — the pure, in-process computation graph. These are versioned: when an input changes, DICE invalidates exactly the dependent keys (with early cutoff — if a recomputation produces the same value, dependents don’t recompute).

  • The action cache memoizes execution — keyed by the content hash of an action’s inputs, not by a DICE version. This is what survives across runs and across machines (via the remote cache), because a content hash is machine-independent.

The incremental engine lives in src/dice/; the executor and action-cache orchestration in src/exec/.

Content addressing & the CAS

Everything in zut is addressed by the hash of its bytes. This is the substrate that makes caching correct and sharing safe.

Digests

A digest is the BLAKE3 hash of some bytes plus their length. zut’s digest type and hashing live in src/core/digest.zig. The hash function is recorded in a REAPI-compatible way (the CLI banner shows digest: blake3, REAPI #9), so the store knows which function produced a given address.

The CAS (content-addressed store)

The CAS (src/core/cas.zig) is a key→bytes store where the key is the content’s digest. Properties that fall out of that:

  • Deduplication — identical bytes are stored once, regardless of how many actions produce or consume them.
  • Integrity — a blob read back is verified against the digest it was asked for; corruption or a lying remote is detected.
  • Immutability — an address never changes meaning, which is why a cached result computed elsewhere is trustworthy here.

The CAS is tiered: a read misses locally, falls through to the remote cache if configured, verifies the blob, and populates the local tier. Writes go through to the remote (write-through). A remote failure degrades to local-only rather than breaking the build.

Trees (Merkle directories)

A directory is content-addressed as a Tree (src/core/merkle.zig): a node listing names → child digests (files or sub-trees). The digest of a Tree is therefore a hash of its entire recursive content. This is how:

  • a crate’s unpacked source is grafted into a build as a single immutable input (crate_tree(...)), and
  • an action’s input set is named by one root digest.

The action cache

An action (src/core/action_cache.zig) is keyed by the digest of its command + input Tree + environment. The action cache maps that key to the action result (output digests, exit code, captured stdout/stderr). Because the key is a content hash:

  • the same action never runs twice, and
  • the cache is valid across machines, which is exactly what the remote cache exploits.

This shape mirrors the Bazel Remote Execution API, so the same store backs both local incrementality and remote caching without a second code path.

See docs/ARCHITECTURE.md §3 for the precise encodings and the REAPI mapping.

Repo & source layout

Each top-level directory under src/ is a layer with one job, and dependencies point downward only.

src/
  main.zig / root.zig    CLI + library root — the only place that wires layers together
    │
    ├── starlark/        the Starlark front-end (lexer → parser → eval → rules)
    ├── crates/          the crates.io importer (lock, fetch, metadata, gen)
    ├── exec/            the sandboxed executor + cache orchestration
    ├── dice/            the incremental engine
    ├── graph/           target identity (labels) and the build graph
    ├── core/            the content-addressed substrate: digest, codec, cas,
    │                      merkle, action_cache
    ├── cache/           the remote-cache tier on top of core: remote.zig
    │                      (RemoteStore tiering + backends), objframe.zig (envelope)
    ├── rules/           the built-in @builtin prelude (rust.bzl, zig.bzl)
    ├── aws/             STANDALONE AWS SDK — sigv4, credentials, s3 (own module)
    ├── gcs/             STANDALONE GCS SDK — storage (own module)
    └── log.zig          STANDALONE leveled logger (own module, std-only)

Why some pieces are standalone modules

aws, gcs, and log are separate build modules, not just directories. That gives a compiler-enforced boundary: they cannot reach into zut internals, so they stay reusable on their own.

  • aws / gcs are the seeds of general-purpose Zig SDKs (aws.s3.Client, gcs.storage.Client) that have nothing to do with build systems. zut’s remote cache is a thin adapter over them; framing/compression (objframe) stays on zut’s side so the SDKs only move opaque bytes. Each can later be extracted to its own package.

  • log depends on nothing but std, so every layer imports it as @import("log") — no ../log.zig reaching across directories — and there is exactly one module instance, hence one process-global log threshold.

The dependency arrows

The front-end and importer sit on top; core/ is the content-addressed substrate everything stands on; cache/ builds the remote tier on core/ using the SDKs as backends without leaking them. Nothing imports main.

The full rationale and the migration history are in docs/design/layout.md.

Contributing

Contributions are welcome. This page is the quick version; the canonical, more detailed guide is CONTRIBUTING.md in the repository.

Build & test

$ zig build                    # build the CLI → zig-out/bin/zut
$ zig build test --summary all # unit tests + sandbox smoketests

You need Zig 0.16.0 exactly. See Installation.

Conventions

  • No raw std.debug.print. Use the logger: @import("log") then log.out for program output, log.err/warn/info/ debug/trace for diagnostics.
  • Tests live with their code. When a source file grows large, its tests move to a sibling <name>_test.zig, pulled in from the original via test { _ = @import("<name>_test.zig"); }. White-box tests drive code through its public API.
  • Layers point downward. Respect the source layout; aws/gcs/log are standalone modules that must not reach into zut.
  • Match the surrounding code — comment density, naming, and idiom. zut’s source is heavily commented with the why; keep that up.
  • Every change keeps the suite green and adds tests for new behavior.

Where to start

  • Good entry points are tracked on the roadmap (Phase 2/3 work).
  • Read docs/ARCHITECTURE.md first — it’s the design contract and explains why things are shaped the way they are.
  • Small, reviewable commits with a clear message are preferred; pure moves (e.g. splitting a file) should carry no logic change.

Roadmap

zut is built in phases, each ending in a demoable, tested capability. This is a summary; the full plan with rationale lives in docs/ROADMAP.md.

PhaseThemeStatus
0Vertical slice — rust_binary ← zig_library over the C ABI, no interpreter✅ done
1Starlark front-end — real BUILD/.bzl evaluation, rules as definitions✅ done
2Rust & Zig depth — proc-macros, build scripts, crates.io importerin progress
3Configuration & query — select(), platforms, configuration transitionsplanned
4Remote cache & execution — cash in the REAPI shaperemote cache ✅, execution planned
5Daemon & polish — long-lived server, structured events, great errorsplanned

Cross-cutting, every phase: hermeticity is enforced not trusted; the execution path stays REAPI-shaped; correctness invariants are tested as they’re built.

If you’d like to help with any of this, see Contributing.