Install with Docker Compose
Everything an operator needs to install, verify, upgrade, troubleshoot, and remove a Repave deployment — without access to the application source.
This is the single entry point. Two companions:
docker-compose-deployment.md— deep reference: every.envvariable, backup/restore, nginx, UAT Docker accessairgapped-install.md— offline specifics
1. Host requirements
- Linux host or VM, x86-64 — Debian/Ubuntu or the RHEL family (RHEL 10 and its rebuilds, CentOS Stream, Fedora)
- Docker Engine + Docker Compose v2 (
docker compose) - Persistent disk for database, runtime artifacts, and project workspaces
- 8 GB RAM minimum; 20 GB free disk minimum, 50 GB+ recommended
- Network access to the release registry (online installs only)
No other host tooling is required. The installer carries its own jq,
openssl, curl, and Docker CLI, so host versions cannot affect the result.
If Docker is not installed yet, run the bootstrap script — shipped in the
release bundle's scripts/ directory (download the bundle from the release
page or GCS; it needs no Docker to unpack):
tar -xzf legacy-modernization-<version>-online.tgz
sudo legacy-modernization-<version>-online/scripts/bootstrap-docker.sh --data-root /data/docker
--data-root places Docker's storage on a larger volume, and must be set
before the daemon first starts — relocating a populated data-root later is a
stop-the-daemon operation. Skip the flag to use the default. Run
--check-only first to see what it would do.
The script installs Docker with apt on Debian and Ubuntu and with dnf on
the RHEL family, choosing Docker's repository from /etc/os-release. A host in
neither family is refused by --check-only, before anything on the host is
modified.
Recommended filesystem layout
Two directories grow without bound and should be separate mounts, so neither can fill the root volume:
| Mount | Holds | Size |
|---|---|---|
/var/lib/docker, or the --data-root you choose | Image layers, container writable layers, volumes | 150 GB+ |
/opt/app-rewrite | Database, runtime artifacts, project workspaces | 100 GB+ |
Image layers dominate: the platform images, every UAT stack an agent builds, and the nested daemons' own storage all land under the data-root.
RHEL-family hosts
RHEL 10 and its rebuilds (AlmaLinux, Rocky, Oracle Linux), CentOS Stream, and Fedora all install through the same script. Two things differ from Debian.
Podman is removed, but never while it is in use. RHEL ships Podman, which
conflicts with the docker-ce packages. If Podman holds any container, image,
or per-user rootless storage, the script stops and tells you what it found
rather than deleting it. Migrate or remove those workloads, then re-run. On a
host where Podman is installed but empty, the script removes it and continues.
Docker CE publishes per-major-version repositories. The script pins the
repository to the major version /etc/os-release reports, so a subscription
pinned to a minor release (VERSION_ID="10.2") still resolves to Docker's
10 directory.
SELinux
Install with SELinux enforcing. Do not set it permissive, and do not change the Docker daemon's own SELinux setting.
Docker CE ships with container labelling off — docker info lists seccomp
and cgroupns under security options, and no selinux. SELinux still confines
the daemon; it does not additionally label the containers the daemon starts.
That is the configuration this platform is tested and supported on, and on it
the full stack runs under Enforcing with no policy changes and no denials:
| What | Under Enforcing, daemon default |
|---|---|
| Install-root bind mounts read and written as uid 10001 | works |
TLS certificate directory mounted read-only from /etc/pki | works |
/var/run/docker.sock mounted into a container | works |
| Privileged nested daemons, and image builds inside them | works |
Turning the daemon's labelling on breaks the deployment. With
"selinux-enabled": true in /etc/docker/daemon.json, measured on an el10
host: the application can no longer write its own data directory, and nginx
can no longer read its certificate. Both failures are silent in the audit log —
the denials are suppressed by dontaudit rules, so ausearch -m AVC returns
nothing and the only symptom is a permission error inside a container. If you
need to confirm this on your own host, semodule -DB makes the suppressed
denials visible and semodule -B restores the default.
Making that configuration work would mean relabelling each mounted host path
into the container policy's own type — including the certificate directory
under /etc/pki, which belongs to the operating system and should not be
relabelled to suit one application. The supported configuration avoids the
question entirely.
The one labelling change the install does make is its own: a --data-root
outside /var/lib/docker is registered as an equivalent of it
(semanage fcontext -a -e /var/lib/docker <path>) and relabelled, before the
daemon first starts. /var/lib/docker carries a type the container policy
expects; a directory you nominate instead would otherwise inherit its parent's.
firewalld
Leave firewalld running. Docker creates its own docker zone and places the
bridge in it, while the host's external interface stays in whichever zone you
have assigned. Published ports and container-to-container traffic both work
with firewalld active; no rule needs to be added by hand for the stack itself.
Open 80 and 443 to your users in the usual way.
System-wide crypto policy
A host-wide crypto policy — including the AD-SUPPORT-LEGACY subpolicy — does
not affect the containers. Each image ships its own OpenSSL and reads the
policy files of the image, not the host. Under FIPS mode the same holds for the
containers, but the host's own outbound TLS is restricted, which can affect
registry pulls; raise it with us before enabling FIPS on a deployment host.
Application allowlisting
Endpoint allowlisting products (Carbon Black App Control, fapolicyd) block
execution of binaries they have not seen. Agents build and run code by design,
so their output is new every time. Exclude these paths from execution control,
or the platform's own containers will not start and agent jobs will fail with
errors that do not mention the allowlisting product:
| Path | Why |
|---|---|
The Docker data-root (/var/lib/docker, or the --data-root you chose) | Every container's filesystem, including the platform's own |
/var/lib/containerd | Container runtime state and unpacked image content |
/opt/app-rewrite/workspaces | Project working copies; agents compile and run code here |
/usr/bin/docker*, /usr/libexec/docker | The Docker CLI and its plugins |
Airgapped RHEL hosts
The script needs Docker's repository to be reachable. On a host with no
internet egress, either mirror download.docker.com into your own repository
manager (Satellite, Nexus, or a local dnf repo) and point
/etc/yum.repos.d/docker-ce.repo at it, or install the docker-ce,
docker-ce-cli, containerd.io, docker-buildx-plugin and
docker-compose-plugin RPMs by hand. container-selinux, which docker-ce
requires, comes from your own distribution repositories rather than from
Docker. Once docker compose version works, the rest of
Airgapped hosts applies unchanged.
2. Install
Authenticate once
sudo mkdir -p /opt/app-rewrite/installer-docker-config
sudo docker --config /opt/app-rewrite/installer-docker-config login -u _json_key --password-stdin \
asia-southeast1-docker.pkg.dev < /path/to/artifact-reader.json
Use the per-client read-only service-account key, not gcloud auth login
— human credentials expire and can only be refreshed interactively, which has
stranded a deploy mid-upgrade before.
Log in with
--configpointed at a directory under the install root, not your own~/.docker. The installer runs in its own container — a filesystem separate from the host's — so a plaindocker loginwrites credentials the installer can never see, and every pull then failsUnauthenticated requesteven though the login itself reported success. Logging in with--config /opt/app-rewrite/installer-docker-configand mounting that same directory into the installer (below) puts both sides on the same credentials.preflightchecks for this mount and fails clearly if it is missing, rather than lettinginstalldiscover it mid-pull.This is a separate directory from
docker-config/, whichdeploy-client-release.shand theonboard-client-registry-authrunbook use for a different mechanism: agcloud-managedcredHelpersentry that fetches a fresh token on every pull. The installer image carries nogcloud(dropping it is most of how it stays under 200 MB), so it cannot use a credential helper at all — a static, non-expiringdocker loginis the only option available to it. Writing that static credential into the same directory as an existingcredHelpersentry is a real, previously-hit failure mode: a staticauthsentry alongsidecredHelpersmakes Docker intermittently use the stale entry instead of invoking the helper. Keeping the two directories distinct means a host that uses both mechanisms — a credential helper for its own scripts, and this installer for upgrades — cannot reproduce that incident, because the installer never writes intodocker-config/at all.
Run the installer
export INSTALLER=asia-southeast1-docker.pkg.dev/repave-prod/legacy-modernization/installer:<version>
sudo mkdir -p /opt/app-rewrite
sudo docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
-v /opt/app-rewrite/installer-docker-config:/root/.docker \
"$INSTALLER" install
The installer creates the directory layout, generates secrets, pins image tags, sets data ownership, pulls images, runs migrations, starts the stack, and verifies the result.
It is idempotent — if it fails partway, fix the cause and run the same
command again. It resumes rather than requiring cleanup, and never regenerates
existing secrets or discards your .env edits.
Commands
All use the same image and the socket + install-root mounts. The credentials
mount is needed by install and upgrade — the commands that actually pull
images — and preflight checks that it is present and usable:
| Command | Purpose | Needs credentials mount |
|---|---|---|
preflight | Check host readiness. Read-only — safe at any time, including against a live deployment. | yes (checks it's there) |
install | Install, or resume a failed install. | yes |
upgrade | Upgrade an existing deployment to this image's version. | yes |
verify | Confirm the deployment actually works. | no |
doctor | Print diagnostics for a misbehaving deployment. | no |
uninstall | Tear down. Preserves data unless --delete-data. | no |
Options
| Option | Meaning |
|---|---|
--install-root <path> | Deploy elsewhere than /opt/app-rewrite. Change both sides of the -v to match. |
--airgap | Never contact a registry; images must be preloaded. |
--data-root <path> | Report-only check that Docker's storage is where you expect. |
--health-url <url> | Non-default health endpoint. |
--yes | Skip confirmation prompts (required for non-interactive runs). |
--stop-running-containers | Stop the open code-server (VS Code) session containers a deploy has to drain, instead of waiting for them. An interactive run offers this at a prompt; this flag is how an unattended one opts in. --yes does not imply it. Job containers (agents, BDD/unit-test runs) need no flag: they are always stopped, and their jobs resume on their own. |
--delete-data | uninstall only — also remove data, workspaces, and volumes. |
Why the mount uses the same path on both sides. The app launches sibling containers through the Docker socket, and the daemon resolves their mount paths on the host, not inside the app container — so
.envmust record host paths. Mounting at a matching path means every path is valid on both sides, with nothing to translate. The installer refuses to run if the install root is not a real mount point, because Docker silently creates a missing-vsource and a forgotten mount would otherwise write a whole deployment into a throwaway container filesystem and report success.To deploy elsewhere, change both sides:
-v /srv/app-rewrite:/srv/app-rewrite ... install --install-root /srv/app-rewrite
Airgapped hosts
docker load -i installer-<version>.tar
docker load -i app-images-<version>.tar
sudo docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
"$INSTALLER" install --airgap
--airgap lists anything missing up front rather than failing partway through.
That covers every image the release needs offline, not only the ones Compose
starts: the dind image Repave IDE sessions run beside the agent runner is
checked too, because it is consumed later when a session starts and this host
has no registry to fall back on. Sessions run the agent runner image itself, so
there is no separate session image to load.
Five third-party images (postgres, valkey, the docker-host sidecar,
cli-proxy-api, nginx) are not shipped in the bundle at all — you are
expected to preload each yourself, the same way you would any image
this installer doesn't ship. --airgap only checks each is present and
reports by name any that are missing; it never pulls on your behalf.
postgres and valkey are required to run the stack, so a fully offline
install needs a manual docker pull (if this host can reach their
registries) or docker load from an archive obtained separately;
nginx/cli-proxy-api only matter once you enable their Compose profile, but
the same requirement applies then.
See airgapped-install.md for bundle contents.
After installing
Open the URL, import the signed customer license, then create the first account. The first account becomes the installation's administrator, who manages the license, users and organization-wide API keys from the Admin portal. Set up email there too, so people can reset a forgotten password and receive project invitations.
If the import fails, the message tells you which side the problem is on:
| Message | Meaning |
|---|---|
| "This license is invalid. Check the license file and try again." | The file is not valid JSON, is missing fields, or was signed by a key this release does not trust. Check it survived the transfer intact. |
| "The license signature is invalid. Import a valid Repave license." | The signature does not match the payload — the file was altered after signing. |
| "The license could not be saved because of a server error…" | The license file is fine. Something behind it failed; the cause is in docker compose logs app. |
| "Could not reach the server…" | The request never arrived — the app is down, or a proxy answered instead of it. |
The license is never partially applied: if the import reports an error, no license was recorded, so it is always safe to fix the cause and import the same file again.
Every project needs an API key: its own, set in the project's Settings, or the organization's, set once in the Admin portal and used by every project without one of its own. Either way the key is registered with the CLIProxyAPI gateway under that project automatically.
3. Verify
verify never pulls, so it doesn't need the credentials mount:
sudo docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
"$INSTALLER" verify
Checks real behaviour, not just that containers are up:
- Health endpoint returns
okand the expected version — a stale container answering happily is exactly what this catches - Postgres accepting connections
repaveCLI usable inside the app container- Agent-runner image present locally (absent, agent jobs fail at run time, long after install looks successful)
- CLIProxyAPI gateway listening
data/andworkspaces/fully owned by the app image's uid- Where nginx is configured: health through the proxy, since a recreated app container gets a new IP that nginx caches while the direct check looks clean
4. Upgrade
export INSTALLER=<registry>/installer:<new-version>
sudo docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
-v /opt/app-rewrite/installer-docker-config:/root/.docker \
"$INSTALLER" upgrade
First time using the installer on this host — e.g. the deployment was
installed by an earlier release's manual procedure or the host-side scripts?
Do §2's "Authenticate once" first: without installer-docker-config/
populated and mounted, preflight stops the upgrade at the credentials check.
What happens, in order: preflight → backup (database dump, .env,
compose file, bind-mounted directories) → refresh deployment assets → pull
images → align data ownership with the new image → run migrations →
recreate containers → verify.
Two things worth knowing:
-
Ownership is self-migrating. The target uid is read from the app image rather than hardcoded, so the same step migrates forward on upgrade and reverses cleanly if you roll back to an older image. No manual
chown, no version gate. -
docker-compose.override.ymlis never touched. Client-specific additions (TLS, custom domains) belong there and survive upgrades. The release ownsdocker-compose.ymland overwrites it. -
.envfollowscurrent/. Every run aligns.env's owner with the owner of the deploy directory, so on a host set up byinstall-repave-user.shit belongs to the operator even though the installer itself runs as root. That matters because.envis0600and Compose reads it for interpolation: a root-owned.envmakes every laterdocker compose ps|logs|restartfail withpermission deniedfor an operator who is in thedockergroup but not a sudoer. A deployment already in that state — installed before this fix — is repaired by the nextinstallorupgrade. The mode itself is not preserved:installre-asserts0600, so widening.envto0640for a group is not a durable change; chowncurrent/to the operator instead. -
Running containers are drained first. Agent jobs, BDD/unit-test runs and open VS Code sessions write to the data the upgrade is about to re-own, so they have to be gone before it proceeds.
- Agent jobs are stopped without asking, and carry on by themselves. The
upgrade stops the app first, then the job containers. Each interrupted job
goes back to the queue and, when the new app starts, continues where it
stopped in the same agent session. Its agent log says so: "Interrupted by a
platform restart … It resumes automatically in the same session (automatic
resume 1 of 3)." A job interrupted a fourth time in a row is left FAILED for
someone to resume by hand. A host reboot or a
docker stopof a job container is handled the same way. - Open VS Code sessions need your say. Stopping one closes that Repave IDE and
loses unsaved editor state, so the upgrade lists them — while the app is still
serving — and asks. Answer
nto wait instead (5 minutes, then the upgrade aborts and rolls back); a session never ends on its own, so that only ever times out. Unattended runs need--stop-running-containersto stop them.
- Agent jobs are stopped without asking, and carry on by themselves. The
upgrade stops the app first, then the job containers. Each interrupted job
goes back to the queue and, when the new app starts, continues where it
stopped in the same agent session. Its agent log says so: "Interrupted by a
platform restart … It resumes automatically in the same session (automatic
resume 1 of 3)." A job interrupted a fourth time in a row is left FAILED for
someone to resume by hand. A host reboot or a
Backups land in backups/pre-upgrade/<timestamp>/. Keep the most recent until
you have confirmed the upgrade in the browser, not just via the health check.
5. Troubleshooting
Start with (doctor reports, never pulls, so it doesn't need the credentials
mount either):
sudo docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
"$INSTALLER" doctor
doctor reports versions, disk, compose state, image presence, health, data
ownership, gateway credential registration, and recent errors. It never prints
.env values, so its output is safe to paste into a support thread.
Symptom map
Agent jobs fail with unknown provider for model proj-<id>/<model>, or 502
from the gateway
Restarting or recreating cli-proxy-api wipes every registered per-project
credential. The app re-registers them by itself when it starts, and again before
dispatching a job the gateway has forgotten, so this should heal without help —
check docker compose logs app | grep GatewayCredentials first. If the reconcile
itself is failing (unreachable gateway, mismatched GATEWAY_MANAGEMENT_SECRET),
re-save each affected project's Settings to re-register. doctor compares
registered credentials against projects that should have them.
upgrade fails with N managed container(s) still running after 300s
Something the app launched is still holding the data — almost always an open Repave
IDE session, which the upgrade will not close without consent. If you declined,
or the run was unattended, re-run with --stop-running-containers; it closes
that Repave IDE and loses unsaved editor state, so check with whoever is using it
first. A job container listed here ignored its stop; its job resumes once the
container is gone.
Uploads or agent jobs fail with EACCES
Data ownership does not match the app image's uid. Re-running upgrade fixes
this automatically. To inspect without changing anything:
sudo scripts/patch-repave-ownership.sh; add APPLY=1 to fix in place.
The app tries to build app-rewrite-agent-runner:latest
AGENT_RUNNER_IMAGE is unset or the image was never pulled. verify catches
this; re-run install/upgrade to repin and pull.
Runner logs repeatedly show POST ... fetch failed
Set DOCKER_AGENT_NETWORK to the Compose network and DOCKER_HOST_URL to
http://app:3000, then recreate the app container.
Everything healthy on localhost:3000, but real traffic returns 502
nginx is holding the previous app container's IP. Restart nginx. verify
detects this when nginx is a defined service.
Large uploads fail at nginx
Set client_max_body_size 11g in the nginx config and reload it.
It must sit above the application's own upload cap (MAX_UPLOAD_BYTES, 10 GB), not equal to it: nginx measures the whole request body while the application caps the uploaded file alone, so a 10 GB file arrives as slightly more than 10 GB. Set them equal and nginx always refuses first, with a generic HTML error instead of the application's message naming the file. Set it lower and that lower number silently becomes your real limit, whatever the application tells users.
nginx buffers request bodies to disk. They go to the nginx_body_temp volume, so the space is explicit and relocatable, but a named volume still sits under Docker's data-root by default — the disk whose exhaustion wedges the daemon. Size that disk for client_max_body_size × the concurrent uploads you expect, or point the volume at another disk; see docker-compose-deployment.md.
Upgrading an existing deployment: the live vhost config under
nginx/conf.d/is per-deployment and not tracked in git, so a release upgrade will not change it. An install that predates this setting still enforces its oldclient_max_body_size— commonly2g— while the application now tells users the limit is 10 GB. Edit the file and reload nginx (docker compose exec nginx nginx -s reload) as part of the upgrade, or the lower value stays your real limit. Addclient_body_temp_path /var/cache/nginx/client_temp;at the same time: without it nginx uses its compiled-in default, which happens to match the mounted volume for thenginx:1.27-alpineimage this ships with, but not for every build —nginx-unprivilegedand several distro packages default to/var/lib/nginx/..., where nothing is mounted and buffered bodies land back on Docker's data-root.
HTTPS does not work
Confirm the Docker nginx service owns ports 80/443, certificates are mounted
read-only, and any host nginx is stopped.
Image pulls fail with Unauthenticated request
Using the installer: almost always the installer-docker-config mount is missing, or
the login used to populate it wasn't done with --config pointed at that
directory — docker login on the bare host writes to the host's own
~/.docker/config.json, a different filesystem from the installer container,
so a successful-looking login there does nothing for the installer's own
pulls. preflight catches this before install starts; re-run it if you
skipped straight to install. See §2's "Authenticate once" for the exact
commands.
Using the host-side scripts directly (deploy-client-release.sh and
similar): usually a stale static auths entry conflicting with the gcloud
credential helper in docker-config/config.json — not an IAM problem. See
docker-compose-deployment.md.
A pull fails with no space left on device
Docker's data-root filled. preflight checks this in advance. See
bootstrap-docker.sh --data-root and the relocation notes in the deployment
reference.
exec format error when running the installer
Wrong architecture for this host. Use the release image built for x86-64.
6. Backup and restore
Upgrades back up automatically. For an on-demand backup, and for the full
restore procedure, see
docker-compose-deployment.md — the
database dump plus .env, compose file, and bind-mounted directories are what
a restore needs.
7. Uninstall
Like verify and doctor, uninstall never pulls an image, so it does not
need the credentials mount:
# Stop and remove containers; ALL DATA PRESERVED
sudo docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
"$INSTALLER" uninstall
# Also delete data, workspaces, and volumes — irreversible
sudo docker run --rm -it \
-v /var/run/docker.sock:/var/run/docker.sock \
-v /opt/app-rewrite:/opt/app-rewrite \
"$INSTALLER" uninstall --delete-data
Without --delete-data, data is left in place and a later install adopts it.
With it, the installer lists exactly what will be destroyed and requires
confirmation; backups/ is never removed.
8. Manual installation
The installer automates the procedure documented in
docker-compose-deployment.md, which remains
supported and is the reference for anything the installer does not cover.
The manual path in brief — unpack the release bundle, cp .env.example .env,
set NEXTAUTH_SECRET/INTERNAL_API_KEY (openssl rand -base64 32), set image
tags from metadata/images.json, set APP_DATA_HOST_DIR/
APP_WORKSPACE_HOST_DIR to host paths, create those directories owned by the
image's repave uid (10001), then:
docker compose pull
docker compose up -d postgres docker-host cli-proxy-api claude-mem-server claude-mem-worker
docker compose run --rm migrate
docker compose up -d --no-deps app
curl -fsS http://localhost:3000/api/health
On a fresh install the migration service bootstraps the database from the
packaged schema and records the packaged migrations as applied, leaving future
upgrades on normal Prisma migration history. Do not substitute prisma db push.
Older pilot databases created by schema push must be baselined once — see the
PRISMA_BASELINE_EXISTING_SCHEMA flow in the deployment reference.
Historical note: earlier releases created a dedicated host
repaveuser at uid 1000. That is no longer needed — the images use uid 10001, which no cloud image's login user occupies, so container-writable directories are simply owned by that numeric uid andls -lshowing a bare number is expected.