Builds and CI
The rules
- Every repository builds with Bazel, at a version pinned in
.bazelversion.bazel build //...andbazel test //...work the same way in every repo. - Tools are pinned by the repository, never assumed from the machine. Programs a build runs come from nixpkgs through
rules_nixpkgs. A program nixpkgs doesn’t have comes from anhttp_archivewith asha256. Upgrading a tool is then a one-line commit, and Bazel rebuilds exactly what depends on it. - CI uses one shared base image. It contains Nix, Bazelisk, the AWS CLI and
bazel-cache-sync, and nothing project-specific. No repository maintains its own CI image. - Everyone reads the shared cache; only trusted machines write to it. Trusted means a protected-branch CI pipeline in the
svrbcgroup, or an SVRBC administrator’s workstation (see below). - Secrets live in GitLab CI/CD variables, never in a repository. The cache itself needs none.
How the cache decides what to rebuild
Bazel models a build as actions. An action is one command with declared inputs and outputs, such as “engrave psalm-82 with MuseScore into an SVG.” The cache has two halves:
cas/(content-addressable store): files, each stored under the SHA-256 of its own bytes. A blob can’t be stored under the wrong name, and identical files are stored once.ac/(action cache): results, each stored under the hash of the action itself: its command line, its environment, the hash of every input file and of every tool. A result says which outputs the action produced, by theircas/hashes.
If an action’s ac/ entry exists, Bazel downloads its outputs instead of running it. Timestamps play no part, so a fresh clone on another machine gets the same hits.
Bazel also stops early. If an action re-runs but produces byte-identical output, the hashes downstream don’t change, so nothing downstream re-runs.
Caching is only correct if every input is declared. Bazel runs actions in a sandbox that hides undeclared files, so a missing input fails the build instead of producing a wrong cache hit.
Why only trusted machines write: a cas/ blob checks itself, but an ac/ entry is a claim (“this action produced that file”) that nobody can check without re-running the action. A bad writer could point a legitimate action at a poisoned file, and every reader would use it.
The shared cache
One cache serves every SVRBC project. Keys hash the exact command, tools and inputs, so projects can’t collide, and identical work is stored once. There is no cache server:
Reading. Every repository’s
.bazelrcpoints Bazel at the cache over plain HTTPS, and tells it never to upload:shcommon --remote_cache=https://bazel-cache.svrbc.org common --remote_upload_local_results=false common --disk_cache=~/.cache/bazel-diskCloudFront serves the private bucket read-only. A CloudFront function passes only
ac/<sha256>andcas/<sha256>paths, so the bucket can’t be listed, and compression is off because Bazel checks every blob against its hash. Reads are public. That’s fine for public repositories, whose outputs anyone could build from the source; finding a private repository’s outputs would take its source too.Writing. Bazel can’t write to S3 (it can’t sign requests), and doesn’t need to. A trusted machine builds into its disk cache as usual, then runs
bazel-cache-sync, which uploads the new entries: everycas/blob before anyac/entry, so a result never appears before its outputs.Retention. S3 lifecycle rules expire
ac/entries after 60 days andcas/blobs after 67, so a result rarely outlives its outputs. When one does, Bazel notices the missing blob, re-runs that one action and carries on (tested with Bazel 9.2 in both download modes).Misses are quiet. CloudFront may read the bucket’s key list, so a missing key is a
404(a miss), not a403(an error). Listing itself is still blocked by the path filter.
How it was built, as AWS CLI commands, is in infra/bazel-cache/.
Trusted workstations
A workstation is trusted when it belongs to an SVRBC administrator, has full-disk encryption, and signs in to AWS through the access portal. There’s no shared secret to hand out: write access is the person’s own AWS role, and revoking it is removing their access.
To upload what this machine has built (see Getting started for the one-time setup):
aws sso login --profile svrbc
AWS_PROFILE=svrbc bazel-cache-syncUpload only what you built from committed, reviewed code; an upload is trusted by every other machine.
CI
Every repository includes the shared template rather than writing its own cache and runner setup:
include:
- project: svrbc/developers.svrbc.org
file: ci/bazel.yml
build:
extends: .svrbc-bazel
script:
- bazel build //...
- bazel test //....svrbc-bazel runs the job in the shared base image. On a protected branch or tag, its after_script asks GitLab for an OIDC token, exchanges it for short-lived AWS credentials (the svrbc-bazel-cache-writer role trusts only master, main and tags of svrbc projects), and runs bazel-cache-sync. Everywhere else the job only reads. A failed upload is reported but never fails the job: the cache only ever saves time.
GitLab.com’s shared runners start every job from a clean state, so each job downloads its toolchain from the Nix binary cache. If that proves slow, the fix is a generated warm image (built from the same pins with nix2container) or a Nix binary cache near the runners. A hand-written Dockerfile is never the fix.
Why this setup
The goal: never redo a build step whose inputs haven’t changed, whether on a workstation, in CI, or on another machine, while keeping local edit-and-rebuild fast.
- Why not Nix on its own? Nix caches whole derivations and decides whether to rebuild by inputs. It has no early stop on unchanged output (that needs content-addressed derivations, which are experimental), and every build pays for evaluation and copying the source into the store. That’s fine for CI and slow for an edit-and-see loop. We still use Nix, for what it does best: pinned, reproducible tools.
- Why not a Docker image per project? Images drift from the build’s idea of its inputs. Upgrade Chromium in an image, forget to tell the build, and the cache serves stale artifacts. They also add a pipeline stage and tag juggling to every repo.
- Why no cache server? Bazel’s remote cache speaks gRPC, or HTTP
GETandPUT. S3 needs signed requests, so Bazel can’t write to it. But it can read from it, and writes don’t have to go through Bazel: a trusted machine uploads its disk cache afterwards. That removes the server (bazel-remoteon EC2, with its patching, TLS and password file) and leaves a bucket, a CDN and an upload script. The cost: new results appear when a job finishes, not as each action completes. - Why not hand-rolled input hashes? doreancon.org’s
build-stamps.jsonalready skips unchanged hymns. But it only hashes the inputs someone remembered to list (not tool versions), and it can’t share results between machines.
Pilot results
doreancon.org’s hymn engraving, measured 2026-10-01 on the bazel-pilot branch. There are 18 hymns. MuseScore 4.7.4, Xvfb and Node 24 come from one pinned nixpkgs commit through rules_nixpkgs_core 0.14.0, with Bazel 9.2.0.
| Case | Result |
|---|---|
| Engrave all 18 hymns, warm toolchain | 12.6 s |
| Edit one hymn | 1.3–1.4 s, one action (1.8 s before the fonts were pinned) |
| Revert that edit | 0.09 s, from the disk cache |
| Change a manifest field the engraver doesn’t read | 0.23 s: the per-hymn cuts re-run, all engravings skipped |
| Fresh checkout, shared cache | 36/36 cached; 12.7 s, all of it Bazel start-up |
| Cold container (base image, nothing cached) | 60 s, about 55 s of it toolchain download |
| Cold container, remote cache warmed by trusted CI | 36/36 remote hits, 0.46 s critical path, 42 s total |
| Same hymn engraved directly, warm MuseScore, no Bazel | 0.8 s |
Fresh output base, reading bazel-cache.svrbc.org | 36/36 remote hits, no warnings |
examples/hello on GitLab.com, second master pipeline | 1/1 remote hit, uploaded by the first through OIDC |
What we learned:
- A single local edit is not faster. It’s about 0.5 s slower than running the engraver directly (Bazel’s sandbox and start-up). The gains are elsewhere: no work is repeated across machines, nothing re-runs when an output comes out unchanged, and the tool versions are part of every cache key.
- Nix’s fontconfig reads the host’s fonts. On a non-NixOS machine it falls back to
/etc/fontsand scans/usr/share/fonts, so text could set differently per machine with nothing in the cache key to show it. Pin afonts.confthat lists only Nix fonts and a prebuilt cache, and pass it asFONTCONFIG_FILE(doreancon.org’snix/fonts.nix). That fixed hermeticity and also removed most of the per-action cost: a freshHOMEno longer rescans fonts. - MuseScore’s output isn’t byte-stable, so nothing downstream of a fresh engraving can be skipped. The skip happens before MuseScore, at the per-hymn manifest cut.
xvfb-run -araces when actions run in parallel: most X servers die on a shared display number. StartXvfb -displayfddirectly instead.- GTK starts gvfs daemons that leave a FUSE mount in
$HOME. SetGIO_USE_VFS=local. - Bazel falls back to
processwrapper-sandboxon Ubuntu workstations (AppArmor restricts unprivileged user namespaces) and in unprivileged containers. It works, but it can’t hide absolute paths outside the sandbox.
Open questions
The doreancon.org pilot exists to answer these before other repositories convert:
- [x] Does
rules_nixpkgswork under Bzlmod with Bazel 9? Yes (0.14.0, 9.2.0). - [x] MuseScore from nixpkgs or the AppImage? nixpkgs: same version, and every hymn’s crop matches the Docker-built SVG exactly.
- [x] Does a build work in an unprivileged container? Yes, with
processwrapper-sandbox, in the base image under local Docker. - [x] The same on GitLab.com’s shared runners? Yes, and better: they allow Bazel’s full
linux-sandbox, and the non-rootbuilderuser can write to the checkout (verified 2026-10-02 withexamples/hello). - [ ] Per-job toolchain download on GitLab.com (about 55 s locally).
- [x] Where does the cache server run? Nowhere: reads go through CloudFront and trusted machines upload with
bazel-cache-sync. - [x] Does Bazel survive a result whose output blob has expired? Yes: it re-runs that action (Bazel 9.2,
toplevelandalldownload modes). - [x] Does CI write through GitLab OIDC? Yes: one
masterpipeline ofexamples/hellouploaded, and the next got1 remote cache hit(2026-10-02). - [x] Is a single local edit fast enough? 1.3–1.4 s against 0.8 s direct, after pinning the fonts. Judged acceptable for the pilot.