Skip to content

Builds and CI ​

The rules ​

  1. Every repository builds with Bazel, at a version pinned in .bazelversion. bazel build //... and bazel test //... work the same way in every repo.
  2. Tools are pinned by the repository, never assumed from the machine. Programs a build runs come from nixpkgs through rules_nixpkgs. A program nixpkgs doesn’t have comes from an http_archive with a sha256. Upgrading a tool is then a one-line commit, and Bazel rebuilds exactly what depends on it.
  3. CI uses one shared base image. It contains Nix, Bazelisk, the AWS CLI and bazel-cache-sync, and nothing project-specific. No repository maintains its own CI image.
  4. Everyone reads the shared cache; only trusted machines write to it. Trusted means a protected-branch CI pipeline in the svrbc group, or an SVRBC administrator’s workstation (see below).
  5. Secrets live in GitLab CI/CD variables, never in a repository. The cache itself needs none.

How the cache decides what to rebuild ​

Bazel models a build as actions. An action is one command with declared inputs and outputs, such as “engrave psalm-82 with MuseScore into an SVG.” The cache has two halves:

  • cas/ (content-addressable store): files, each stored under the SHA-256 of its own bytes. A blob can’t be stored under the wrong name, and identical files are stored once.
  • ac/ (action cache): results, each stored under the hash of the action itself: its command line, its environment, the hash of every input file and of every tool. A result says which outputs the action produced, by their cas/ hashes.
Looking up an action in the cacheAn action's command, inputs and tools are hashed into an ac/ key. The ac/ entry lists the action's outputs by their cas/ hashes, and each cas/ entry holds the bytes. If there is no ac/ entry, Bazel runs the action itself.An actioncommand + environment+ every input’s hash+ every tool’s hashac/d6fa21f4…“exit 0; psalm-82.svgis cas/9a3eb477…”a claim: only trusted writerscas/9a3eb477…the SVG’s byteschecks itself: name = hashhashfetchno entry: run it, then storemiss

If an action’s ac/ entry exists, Bazel downloads its outputs instead of running it. Timestamps play no part, so a fresh clone on another machine gets the same hits.

Bazel also stops early. If an action re-runs but produces byte-identical output, the hashes downstream don’t change, so nothing downstream re-runs.

Caching is only correct if every input is declared. Bazel runs actions in a sandbox that hides undeclared files, so a missing input fails the build instead of producing a wrong cache hit.

Why only trusted machines write: a cas/ blob checks itself, but an ac/ entry is a claim (“this action produced that file”) that nobody can check without re-running the action. A bad writer could point a legitimate action at a poisoned file, and every reader would use it.

The shared cache ​

One cache serves every SVRBC project. Keys hash the exact command, tools and inputs, so projects can’t collide, and identical work is stored once. There is no cache server:

Who reads and who writes the shared cacheEvery machine reads through bazel-cache.svrbc.org, a read-only CloudFront distribution in front of the private S3 bucket svrbc-bazel-cache. Protected CI pipelines write with bazel-cache-sync using GitLab OIDC, and trusted workstations write with bazel-cache-sync using AWS SSO.Any machineworkstations, MR pipelinesProtected CImaster and tags, svrbc groupTrusted workstationan SVRBC administrator’sbazel-cache.svrbc.orgCloudFrontGET only; ac/ and cas/ keysanything else is a 404S3 bucketsvrbc-bazel-cacheprivate, us-west-2ac/ expires at 60 dayscas/ at 67 daysreadsbazel-cache-sync, GitLab OIDCbazel-cache-sync, AWS SSO
  • Reading. Every repository’s .bazelrc points Bazel at the cache over plain HTTPS, and tells it never to upload:

    sh
    common --remote_cache=https://bazel-cache.svrbc.org
    common --remote_upload_local_results=false
    common --disk_cache=~/.cache/bazel-disk

    CloudFront serves the private bucket read-only. A CloudFront function passes only ac/<sha256> and cas/<sha256> paths, so the bucket can’t be listed, and compression is off because Bazel checks every blob against its hash. Reads are public. That’s fine for public repositories, whose outputs anyone could build from the source; finding a private repository’s outputs would take its source too.

  • Writing. Bazel can’t write to S3 (it can’t sign requests), and doesn’t need to. A trusted machine builds into its disk cache as usual, then runs bazel-cache-sync, which uploads the new entries: every cas/ blob before any ac/ entry, so a result never appears before its outputs.

  • Retention. S3 lifecycle rules expire ac/ entries after 60 days and cas/ blobs after 67, so a result rarely outlives its outputs. When one does, Bazel notices the missing blob, re-runs that one action and carries on (tested with Bazel 9.2 in both download modes).

  • Misses are quiet. CloudFront may read the bucket’s key list, so a missing key is a 404 (a miss), not a 403 (an error). Listing itself is still blocked by the path filter.

How it was built, as AWS CLI commands, is in infra/bazel-cache/.

Trusted workstations ​

A workstation is trusted when it belongs to an SVRBC administrator, has full-disk encryption, and signs in to AWS through the access portal. There’s no shared secret to hand out: write access is the person’s own AWS role, and revoking it is removing their access.

To upload what this machine has built (see Getting started for the one-time setup):

sh
aws sso login --profile svrbc
AWS_PROFILE=svrbc bazel-cache-sync

Upload only what you built from committed, reviewed code; an upload is trusted by every other machine.

CI ​

Every repository includes the shared template rather than writing its own cache and runner setup:

yaml
include:
  - project: svrbc/developers.svrbc.org
    file: ci/bazel.yml

build:
  extends: .svrbc-bazel
  script:
    - bazel build //...
    - bazel test //...

.svrbc-bazel runs the job in the shared base image. On a protected branch or tag, its after_script asks GitLab for an OIDC token, exchanges it for short-lived AWS credentials (the svrbc-bazel-cache-writer role trusts only master, main and tags of svrbc projects), and runs bazel-cache-sync. Everywhere else the job only reads. A failed upload is reported but never fails the job: the cache only ever saves time.

GitLab.com’s shared runners start every job from a clean state, so each job downloads its toolchain from the Nix binary cache. If that proves slow, the fix is a generated warm image (built from the same pins with nix2container) or a Nix binary cache near the runners. A hand-written Dockerfile is never the fix.

Why this setup ​

The goal: never redo a build step whose inputs haven’t changed, whether on a workstation, in CI, or on another machine, while keeping local edit-and-rebuild fast.

  • Why not Nix on its own? Nix caches whole derivations and decides whether to rebuild by inputs. It has no early stop on unchanged output (that needs content-addressed derivations, which are experimental), and every build pays for evaluation and copying the source into the store. That’s fine for CI and slow for an edit-and-see loop. We still use Nix, for what it does best: pinned, reproducible tools.
  • Why not a Docker image per project? Images drift from the build’s idea of its inputs. Upgrade Chromium in an image, forget to tell the build, and the cache serves stale artifacts. They also add a pipeline stage and tag juggling to every repo.
  • Why no cache server? Bazel’s remote cache speaks gRPC, or HTTP GET and PUT. S3 needs signed requests, so Bazel can’t write to it. But it can read from it, and writes don’t have to go through Bazel: a trusted machine uploads its disk cache afterwards. That removes the server (bazel-remote on EC2, with its patching, TLS and password file) and leaves a bucket, a CDN and an upload script. The cost: new results appear when a job finishes, not as each action completes.
  • Why not hand-rolled input hashes? doreancon.org’s build-stamps.json already skips unchanged hymns. But it only hashes the inputs someone remembered to list (not tool versions), and it can’t share results between machines.

Pilot results ​

doreancon.org’s hymn engraving, measured 2026-10-01 on the bazel-pilot branch. There are 18 hymns. MuseScore 4.7.4, Xvfb and Node 24 come from one pinned nixpkgs commit through rules_nixpkgs_core 0.14.0, with Bazel 9.2.0.

CaseResult
Engrave all 18 hymns, warm toolchain12.6 s
Edit one hymn1.3–1.4 s, one action (1.8 s before the fonts were pinned)
Revert that edit0.09 s, from the disk cache
Change a manifest field the engraver doesn’t read0.23 s: the per-hymn cuts re-run, all engravings skipped
Fresh checkout, shared cache36/36 cached; 12.7 s, all of it Bazel start-up
Cold container (base image, nothing cached)60 s, about 55 s of it toolchain download
Cold container, remote cache warmed by trusted CI36/36 remote hits, 0.46 s critical path, 42 s total
Same hymn engraved directly, warm MuseScore, no Bazel0.8 s
Fresh output base, reading bazel-cache.svrbc.org36/36 remote hits, no warnings
examples/hello on GitLab.com, second master pipeline1/1 remote hit, uploaded by the first through OIDC

What we learned:

  • A single local edit is not faster. It’s about 0.5 s slower than running the engraver directly (Bazel’s sandbox and start-up). The gains are elsewhere: no work is repeated across machines, nothing re-runs when an output comes out unchanged, and the tool versions are part of every cache key.
  • Nix’s fontconfig reads the host’s fonts. On a non-NixOS machine it falls back to /etc/fonts and scans /usr/share/fonts, so text could set differently per machine with nothing in the cache key to show it. Pin a fonts.conf that lists only Nix fonts and a prebuilt cache, and pass it as FONTCONFIG_FILE (doreancon.org’s nix/fonts.nix). That fixed hermeticity and also removed most of the per-action cost: a fresh HOME no longer rescans fonts.
  • MuseScore’s output isn’t byte-stable, so nothing downstream of a fresh engraving can be skipped. The skip happens before MuseScore, at the per-hymn manifest cut.
  • xvfb-run -a races when actions run in parallel: most X servers die on a shared display number. Start Xvfb -displayfd directly instead.
  • GTK starts gvfs daemons that leave a FUSE mount in $HOME. Set GIO_USE_VFS=local.
  • Bazel falls back to processwrapper-sandbox on Ubuntu workstations (AppArmor restricts unprivileged user namespaces) and in unprivileged containers. It works, but it can’t hide absolute paths outside the sandbox.

Open questions ​

The doreancon.org pilot exists to answer these before other repositories convert:

  • [x] Does rules_nixpkgs work under Bzlmod with Bazel 9? Yes (0.14.0, 9.2.0).
  • [x] MuseScore from nixpkgs or the AppImage? nixpkgs: same version, and every hymn’s crop matches the Docker-built SVG exactly.
  • [x] Does a build work in an unprivileged container? Yes, with processwrapper-sandbox, in the base image under local Docker.
  • [x] The same on GitLab.com’s shared runners? Yes, and better: they allow Bazel’s full linux-sandbox, and the non-root builder user can write to the checkout (verified 2026-10-02 with examples/hello).
  • [ ] Per-job toolchain download on GitLab.com (about 55 s locally).
  • [x] Where does the cache server run? Nowhere: reads go through CloudFront and trusted machines upload with bazel-cache-sync.
  • [x] Does Bazel survive a result whose output blob has expired? Yes: it re-runs that action (Bazel 9.2, toplevel and all download modes).
  • [x] Does CI write through GitLab OIDC? Yes: one master pipeline of examples/hello uploaded, and the next got 1 remote cache hit (2026-10-02).
  • [x] Is a single local edit fast enough? 1.3–1.4 s against 0.8 s direct, after pinning the fonts. Judged acceptable for the pilot.