Skip to main content

UIS CLI Reference

The ./uis command manages the UIS provision-host container and all services within it. Commands are organized into host-level (managing the container) and service-level (managing Kubernetes services inside the container).

Container Management

These commands run on the host machine and manage the UIS container.

CommandDescription
./uis startStart the UIS provision-host container
./uis stopStop the container
./uis restartRestart the container
./uis containerShow container status
./uis shellOpen interactive bash shell in the container
./uis exec <command>Execute a command inside the container
./uis logs [--tail N]Show container logs (default: last 50 lines)
./uis buildBuild the container image locally as uis-provision-host:local

Platform Management

UIS targets multiple Kubernetes platforms (Rancher Desktop, Azure AKS, …). The uis platform subcommands surface them all under a single command interface. See Platforms overview for the full mechanic.

CommandDescription
./uis platform list [--offline|--deep]List all platforms and their state. --offline skips reachability probe; --deep adds per-platform extras (e.g. cluster version, cost).
./uis platform use [<name>] [--offline]Switch the active platform — kubectl context + cluster-config.sh flip together. No arg → interactive picker over reachable platforms. --offline allows switching to an unreachable platform (e.g. to clean up stale state).
./uis platform init <provider>Interactive setup wizard for a cloud platform. Writes .uis.secrets/cloud-accounts/<provider>-default.env.
./uis platform up <provider>Provision the cluster end-to-end. Chains bootstrap + tofu apply + post-apply configuration. Auto-flips active platform to the new cluster on success.
./uis platform status <provider>Show cluster state, external IP, and rough cost estimate. Does not target the active platform — reports on the named one.
./uis platform down <provider>Tear down the cluster. Requires typing the cluster name to confirm (irreversible). Auto-resets active platform back to rancher-desktop on success.

platform list — canonical output

$ ./uis platform list
Active: rancher-desktop

PLATFORM STATUS
rancher-desktop ✓ running (active) local k3s
azure-aks · configured, not running (run './uis platform up azure-aks' to start it)

Four possible state values per row: ✓ running, · configured, not running, · not initialized, ✗ unreachable. See Platforms overview for what each means.

platform use — canonical output

$ ./uis platform use rancher-desktop
✓ Switched: azure-aks → rancher-desktop
$ ./uis platform use      # no arg → interactive picker
PLATFORM STATUS
[1] rancher-desktop ✓ running (currently active) local k3s
azure-aks · configured, not running (run './uis platform up azure-aks' to start it)

Pick a platform [1-1]:

Only running platforms get selectable numbers. Switching to a not initialized or unreachable platform doesn't have a meaningful outcome.

./uis deploy, ./uis undeploy, ./uis list, ./uis status, ./uis configure, ./uis expose, ./uis stack install, and ./uis test all all print a one-line banner identifying the active platform before running:

$ ./uis deploy nginx
ℹ Platform: azure-aks (reachable)
(deploy output follows…)

If no platform is active or the active platform is unreachable, the banner aborts the command with a recovery hint. See Platforms overview for all four banner cases.

Network Management

UIS supports two networking providers — Cloudflare (production-grade tunnels with WAF + your own domain) and Tailscale (per-service Funnel for dev sharing on any network). The uis network subcommands manage both under a single command surface. See Networking for the comparison + walkthrough.

CommandDescription
./uis network listList both providers and their state.
./uis network init <provider>Interactive setup wizard. <provider> is cloudflare or tailscale.
./uis network up <provider> [flags]Deploy the provider into the cluster. Tailscale supports --with-cluster-funnel for an opt-in catch-all device.
./uis network down <provider>Tear down the provider's cluster footprint.
./uis network status <provider>Show provider state, tunnel/route state, pod health.
./uis network verify <provider>Run the provider's diagnostics.
./uis network expose tailscale <service> [--yes]Expose a service via a per-service Tailscale Funnel device. Namespace auto-detected. Tailscale-specific — Cloudflare uses cluster-wide HostRegexp routing instead.
./uis network unexpose tailscale <service>Undo per-service Funnel exposure.

Service Management

Discovery

CommandDescription
./uis listList all services with deployment status
./uis list --category <id>Filter by category (e.g., DATABASES, OBSERVABILITY)
./uis list --allShow all services including disabled
./uis statusShow deployed services health and cluster context
./uis categoriesList all service categories

Deploy and Undeploy

CommandDescription
./uis deployDeploy all enabled/autostart services
./uis deploy <service-id>Deploy a specific service (auto-enables it)
./uis undeploy <service-id>Remove a service from the cluster

Autostart Configuration

Services can be marked for automatic deployment.

"Autostart" does not mean "starts at boot"

Nothing in UIS runs when the machine boots. The enabled list only decides what ./uis deploy deploys when you run it with no arguments — it is a default argument list, not a startup mechanism.

What actually happens after a host restart:

  • Deployed services come back on their own. Deployments, Services, IngressRoutes, Secrets and database roles are cluster state; once the cluster is running again the kubelet restarts the pods. Nothing needs redeploying, and re-running ./uis deploy is not the fix.
  • The cluster itself is not started by UIS. Rancher Desktop is installed at OS level, so whether it comes up with the machine is a host setting outside UIS's control.

So if services are unreachable after a reboot, check whether the cluster is running before redeploying anything.

CommandDescription
./uis enable <service-id>Add service to autostart (deploys on next ./uis deploy)
./uis disable <service-id>Remove from autostart (does not undeploy)
./uis list-enabledShow all services in autostart configuration
./uis syncAuto-enable all currently deployed services

Verification

CommandDescription
./uis verify <service-id>Run service-specific verification checks

Stack Management

Stacks are pre-configured groups of related services deployed together.

CommandDescription
./uis stacksList all available stacks
./uis stack info <stack-id>Show stack details (components, dependencies)
./uis stack install <stack-id>Install all services in a stack in order
./uis stack install <stack-id> --skip-optionalSkip optional services
./uis stack remove <stack-id>Remove all services in a stack

Available stacks: observability, ai-local, analytics

Template Management

A template installs an application that spans several services — a database with migrations, per-app instances, exposures — under one name and one app_name. It is the uis half of Rules for Deploying Applications: templates provision, ArgoCD deploys workloads.

CommandDescription
./uis template listList available UIS templates from the registry
./uis template info <id>Show one template's details
./uis template install <id> [--dry-run] [--param k=v]...Deploy and configure every service the template declares
./uis template remove <id> [--app <name>] [--purge] [--yes]Remove an installed application. Data is kept unless --purge. --app picks one tenant when a template has several

--dry-run

Prints every deploy/configure the install would run, in order, with params resolved, and runs nothing. ⚠️ It is a dry run of the install, not of the fetch — the definition artifact is pulled, because that is how the plan is known. Nothing else is written.

Several tenants of one template

--param app_name= installs a second, independent tenant — a live one and a test one, say. The record holds one entry per tenant, keyed on app_name, so they do not collide.

🔴 Before 1.6.35 the record was keyed on the template id, so a second install silently replaced the first's record: the first tenant stayed deployed, healthy and serving traffic, and could no longer be removed by the tool that installed it.

Consequences worth knowing:

  • remove <id> is ambiguous once a template has two tenants, and refuses, listing the --app values it knows: ./uis template remove atlas --app atlas-t
  • requires: <id> is ambiguous the same way and refuses for the same reason — two tenants export different URLs and nothing can say which you meant
  • a record with no app_name (written before 1.6.29) cannot be keyed, so recording one is refused rather than allowed to collide

--param app_name= and what remove remembers

The install records the effective app_name in .uis.extend/applications.yaml, and remove derives every per-app name from that — never from the application id. They are the same string only when --param app_name= was not used.

🔴 A record written before 1.6.29 has no app_name. remove falls back to the id, says so loudly, and refuses --yes: the plan cannot be verified against what was installed, and the failure mode is undeploying a different tenant's instance. Read the plan and confirm interactively, or remove by hand.

⚠️ app_name becomes a SQL identifier and a Kubernetes name. Letters, digits, underscore and hyphen only; anything else is refused before any SQL runs. A hyphen is fine — the database keeps it, the Postgres role converts it to an underscore — so --param app_name=my-app yields database my-app and role my_app.

What remove does and does not do

It removes what the install added, not what the application produced:

removedkept
the code-location entries, then dagster is redeployeddatabases, roles and secrets
per-app instances of multi-instance servicessingle-instance shared services
the application record

⚠️ Single-instance services are never undeployed. postgresql is shared; an application does not own it, and removing an application must not take the platform's database with it.

--purge additionally drops the per-app Postgres roles and secrets. An application's database is the one thing reinstalling cannot reconstruct, which is why it is opt-in — the same line undeploy and configure --purge already draw.

Removal refuses while another installed application requires it, naming the dependant.

The declaration

A template ships a template-info.yaml:

install_type: stack
params:
app_name: myapp # substituted anywhere as {{ params.app_name }}

provides:
services:
- service: postgresql
config:
database: "{{ params.app_name }}"
init: migrations/ # a file OR a directory
namespace: dagster # where to write the secret
secret_name_prefix: "{{ params.app_name }}-database"
- service: postgrest
config:
schemas: api_v1
url_prefix: api-myapp
stacks:
- observability # expanded to its services, deploy-only

config: keys

Every key maps to a uis configure flag. An unrecognised key is rejected, so a typo such as url-prefix fails the install rather than being silently ignored.

KeyPassed asNotes
database--databaseDeclared once, by whichever service owns it; every other configurable service in the same install is passed the same value rather than deriving its own
init--init-file -a file, or a directory — see below
schemas--schemasPostgREST
url_prefix--url-prefixPostgREST
namespace--namespace⚠️ requires secret_name_prefix
secret_name_prefix--secret-name-prefix⚠️ requires namespace
code_location(not a flag)A mapping, not a scalar — see below. Dagster only

⚠️ The resulting Secret is named <secret_name_prefix>-db and its key is always DATABASE_URL — see Dagster's tenant contract for why that matters when another service consumes it.

code_location: — contributing to Dagster

A provides entry for dagster may carry a code_location mapping. The entry is written into .uis.extend/dagster-code-locations.yaml and Dagster is redeployed, so an application that orchestrates with Dagster installs in one step:

provides:
services:
- service: dagster
config:
code_location:
name: "{{ params.app_name }}-data"
image: ghcr.io/terchris/atlas-data
tag: v20260909-abc1234
module: atlas_data.definitions
why: "the ETL that fills api_v1"
env_secrets: "{{ params.app_name }}-database-db"
FieldRequiredNotes
nameyesThe code-location name. {{ params.* }} is substituted before it is recorded
imageyes
tagyes⚠️ Must be immutable — the same rule the Dagster validator applies
moduleyesThe Python module holding Definitions
whynoFree text, kept in the file for whoever reads it next
env_secretsnoExtra Secrets whose keys become environment variables. A scalar or a list; both are accepted. See below — you usually do not need it

Re-installing at the same pin rewrites the same entry and changes nothing — Helm rolls only when the image field changes. A new pin rolls the tag.

🔴 UIS wires the Secret it created for you — do not restate it. If the same install ran configure postgresql --namespace <ns> --secret-name-prefix <p>, the Secret <p>-db is added to the code location's env_secrets automatically. Naming it again in the definition is harmless (it is de-duped) but wrong in principle: the name is UIS's own construction, so a definition that repeats it is a second place that must agree — and one that hard-codes it breaks under --param app_name.

Use env_secrets only for Secrets this install did not create.

⚠️ Without that wiring, a clean install comes up unable to reach its own databaseEXIT=0, schema present, API answering, and the pipeline dead on the first run. It took a machine that had never seen the application to expose it, because leftover state supplied the Secret on every cluster that had one.

⚠️ This is a code-location writer, not a generic "contribute to another service's extend file" mechanism. Two other files would qualify (prometheus-targets.yaml, monitors.yaml) and the generic form was declined: one consumer does not tell you the shape of three.

commands: — letting an operator ask whether your output is still true

An optional top-level block in the artifact. UIS runs what it declares and interprets nothing:

commands:
check:
description: "Does the output reflect the input?"
run: /opt/atlas/atlas-status.sh
in: code-location

Surfaced as uis template check <id>, which execs run in the code location's pod and passes the application's own exit code straight through.

Every other check in this platform asks whether a component is healthy

An application served a deleted company over its public API for 7.5 hours while every operator-visible signal was green:

feed job SUCCESS every 30 min · exit_code 0 · backlog 0
watermark advancing · GET /entity 200 · 5 instigators RUNNING

meanwhile the transform had failed 16 consecutive times
119 changes unapplied — 22 of them deletions
the register 8.4 hours stale

It was found because a human asked a question. uis verify dagster would have passed throughout — it proves the daemon can fire schedules, not that the data is right.

note
Why check and not status

status already means is it up in four places, and verify means does this component work in about six. This command asks neither. Giving it either word would give the same name to the claim that was true during those 7.5 hours and the claim that was false.

The command must ship IN THE IMAGE, with what it needs

Measured against a real artifact before this shipped: the application's script lived in its repository, not its image, and psql was absent from the image too. Run as declared it printed a header of blank values and exited 0.

A blank reading as "nothing to report" when the truth is "I could not look" is the failure this command exists to end — so UIS pre-flights that the command is present and executable, and treats exit 127 as could not be asked rather than as a failed check.

The exit-code contract

Your check's exit code is the verdict. UIS maps it and never re-judges it:

exitmeaning
0the output reflects the input
1it does not — a definite claim
2could not look — the check could not reach what it needed
127could not look — the command or a dependency is missing
anything elsetreated as a problem, and reported as outside the contract
An application must be able to say "I could not look"

Until 1.6.81 every non-zero code collapsed into UNHEALTHY, so a check that had lost its database connection made a definite claim that the data was wrong. Measured: an application exiting 2 for cannot was reported as UNHEALTHY — reported a problem (exit 2) while its data was fine throughout.

The platform builds four states to keep cannot look apart from unhealthy, and it collapsed at the one hop nobody guarded — the application's own exit code.

Why there is still a catch-all, and why it points the other way

A catch-all must fail toward alarm, never toward reassurance.

UIS FAILED TO EVAL exists because the could not be asked state used to have an "anything else" bucket — and a defect in UIS looked like an honest "cannot tell". That catch-all failed toward reassurance, so it was removed.

An undefined exit code goes the other way: it lands in the alarming state, and the output says UIS is interpreting rather than relaying a meaning the contract defines.

UIS relays the answer; it does not certify it

On exit 0 the output says "reported success. UIS relayed this; it did not verify it."

The application's own point, applied to the platform: "exits 0" is not "the output reflects the input", and a criterion that accepts the former will be satisfied by a stub. UIS cannot judge whether the numbers are right — that is the application's to own. It can refuse to claim it did.

An application that declares nothing says so

uis template check on an application with no commands.check reports that it cannot tell you whether its output reflects its input, and exits 2 — distinct from 1, so a script can tell "checked, and it is wrong" from "there is nothing here that can check". A missing pod exits 2 as well, saying NOTHING WAS CHECKED.

Silence would make this a thing one application has and nobody else does.

Four states, and conflating any two is the defect:

1declared, healthy
2declared, unhealthy
3declared nothing
4could not be asked

"3 of 4 healthy" silently drops what it could not ask. A stopped application must still produce a line, in state 4, with a reason — not vanish from the list.

operational: — what installing this will actually do

An optional top-level block in the artifact, rendered by uis template info. UIS reads nothing from it and validates nothing in it: the application owns the content, the platform only displays it.

Nest a fact beside the sentence it completes

install.takes says the data load afterwards is the long part; install.first_load says how long is long — row counts, table counts, disk. Same question, same reader, so they render one under the other.

⚠️ Nesting a key under a known container does not make it render. The children of install and first_data are listed individually in the table below; a new child is as invisible as a new top-level key until it appears there. An application moved a key under install believing that was enough, and their reasoning was right while the mechanism disagreed with it.

tip
unscheduled and manual_only are different claims
  • unscheduledcannot run. No private data, no credential, nothing to do.
  • manual_onlymust run, once, by hand, and then never again.

They render as different kinds of sentence, not as two lists with different adjectives:

  Run ONCE by hand — nothing will ever trigger it: brreg_bootstrap
Never runs, and nothing to launch: redcross-branches, frr

⚠️ The first wording was "Run once by hand, never on a schedule" against "No schedule at all". Both are schedule-negative, so the reader had to spot the difference in the qualifiers — and "no schedule at all" reads as "you will have to run it yourself", which is the other field's meaning exactly.

Collapsing them loses the one instruction an operator cannot skip. An application needed the distinction and invented manual_only rather than overload the key that existed; it is now rendered on both surfaces, and at install it appears next to the first-data job list, which is the moment it means something.

note
Why it is not redundant with first_data.how

It looks redundant, and on uis template info it is — that field explains in prose why the job is the odd one out. first_data.how does not render at the end of an install.

An application deleted manual_only as a duplicate, then checked which surface each field reaches and put it back: the deletion left the operator about to run the chain seeing the job in an ordered list, absent from unscheduled, with nothing saying it is a one-time load. The fact belongs where the person about to act is looking; the reason belongs where someone investigating is looking.

:::

A key UIS does not know about has no designed layout

UIS renders a known set of operational.* keys with a designed layout. A key outside that set is not an error and is not lost:

  • uis template info shows it anyway, verbatim, under "also declared". UIS reads nothing from this block and validates nothing in it — a platform that only displays content has no business deciding which of it is displayable.
  • The install names it rather than presenting it, because the install is read once and its job is to be short enough to be read.

⚠️ Before 1.6.68 an unknown key was shown nowhere at all. An application moved an upgrade remedy into operational.troubleshooting precisely so operators would see it, and it went from a place they would not look to a place they could not. A second application then had two invented keys silent at once and had asked about only one of them.

If a key deserves a designed layout, ask — adding one is a few lines.

It answers the questions an operator has before installing, and which provides: cannot:

operational:
automation: "Ships stopped. No data is fetched until an operator enables the schedules."
timezone: Europe/Oslo
install:
deploys: [postgresql, postgrest, dagster]
takes: a few minutes
note: "the API is live and serves zero endpoints until the first pipeline run"
first_data:
why: "enabling the schedules does not backfill"
how: "launch these jobs, in this order"
jobs: [annual_sources_refresh, klass_refresh, transform_and_publish]
takes: "~11 minutes"
cadence:
- { cron: "0 2 * * 0", what: "~37 annual public-sector sources" }
external_services: [SSB, FHI]
unscheduled: [parked-source]

Which surface each field reaches

Write each field for the reader who will actually see it. template info is read before installing, by someone deciding. The install summary is read after, by someone who has just watched it finish and is deciding what to do next. They are not the same person and often not the same sentence.

fieldtemplate infoinstall summary
automationyesyes
unscheduledyesyes
install.noteyesyes
first_data.jobsyesyes
first_data.takesyesyes
troubleshootingyesyes (a pointer)
manual_onlyyesyes
install.deploysyesno
install.takesyesno
install.first_loadyesno
first_data.whyyesno
first_data.howyesno
cadenceyesno
external_servicesyesno
timezoneyesno
A field marked "no" cannot be leaned on by one marked "yes"

first_data.why"enabling the schedules does not backfill" — is info-only. So a sentence in automation, which the installer does print, must not assume that warning was read. Say it again if the reader needs it at that moment.

An application discovered this the hard way: its automation sentence was correct, and until 1.6.65 the installer did not render it at all. It had been written blind, for a surface it never reached.

warning
troubleshooting is the one field the install does not print in full

Its reader is someone whose install has already gone wrong — at 02:00, with an error in front of them. Printing remedies at the end of a successful install would be noise, and noise here is expensive: it trains people to skip the block that also carries the automation warning.

So the install prints a pointer"this application ships its own remedies: ./uis template info <id>" — because the end of a successful install is the one moment the operator is certainly reading, and nobody discovers a command at 02:00 that they have never seen. info holds the content.

This table is not documentation of intent — it is checked

test-operational-fields-reach-install.sh parses this table and compares it against both renderers in template.sh. A field added to one and not the other, or a row here that stops matching the code, fails the suite.

That is deliberate: an application author who reads this table is making editorial decisions on the strength of it, and a stale table would quietly make those decisions wrong.

Rendered in two places, deliberately. template info prints the whole block before an install; the install summary prints the short form after one, immediately below Endpoints:.

⚠️ Both, because a user who runs install without info would otherwise never see it — and template list offers info and install as two equal options with nothing marking the first as a prerequisite. Even a reader who did run info met those job names several minutes and several hundred lines earlier; the end of the output is the part that gets read.

🔴 automation is the single most important line. Does installing this start anything? An operator deciding whether to install needs that before the service list, and nothing else in the definition says it.

⚠️ And it must be true on its own, including about sensors. An asset driven by an automation condition has no schedule to switch on — it runs from a sensor, and default_automation_condition_sensor ships stopped too. "Enables the schedules" is the wrong instruction for such an asset: someone who enabled every schedule would still not be running it. Say schedules and sensors, and list the asset in unscheduled.

⚠️ first_data exists because enabling schedules does not backfill. A cron is a next fire, not a catch-up, so a Thursday install can sit empty until Sunday. An application whose data arrives on a schedule should say how to load it now.

⚠️ It lives in the artifact, not the catalogue entry, so it is version-locked to the code it describes — the same reasoning that keeps params: and provides: out of the registry. info therefore pulls the definition at its digest to render this; the pull is cached, so repeated calls cost nothing, and a fetch failure degrades to the registry half rather than failing the command.

requires: and exports: — one application reading another

An application may export values a dependant needs, and declare what it needs:

# in atlas
exports:
api-url: "http://api-{{ params.app_name }}.localhost"

# in atlas-frontend
requires:
- application: atlas
provides: api-url
params:
api_base: "{{ requires.atlas.api-url }}"

⚠️ requires: refuses, it never auto-installs. A missing dependency names the application and the command that installs it, and stops. Installing one application must not silently install another: the second one's exposure decisions are a person's to make.

The install records what it installed in .uis.extend/applications.yaml — the id, artifact, tag, pin (digest), the services and code locations it created, its requires, and its resolved exports. That record is what makes requires a refusal rather than a guess: without it, "is atlas installed?" could only be answered by probing the cluster for symptoms, and a probe that infers presence eventually infers it wrongly.

⚠️ A fixture or stack template installed with no artifact pin writes no record, and says so. A dependant's requires: will not see it.

The source, the allowlist and the pin

An application's install definition is its own OCI artifact, published beside its image and pulled with oras on the provision host — no pod, no cluster, no docker:

source:
artifact: ghcr.io/terchris/atlas-data/uis
tag: v20260909-abc1234
digest: sha256:<64 hex>
CheckBehaviour
allowlistghcr.io/helpers-no/* and ghcr.io/terchris/* by default. Anything else is refused, naming the value and the allowlist. Extend it in .uis.extend/template-allowlist.conf — one glob per line
pinThe digest is pulled; the tag is only shown. A missing or malformed digest is refused — a tag alone is not a pin, because tags are mutable at a registry
immutable taglatest, main, master, head and an empty tag are refused even with a valid digest

The allowlist is a security boundary, not a convenience: the definition is fed to configure --init-file, which applies SQL as the database owner. A merged typo in the catalogue must not be able to point a platform at a stranger's SQL.

Private artifacts use oras login ghcr.io with GITHUB_USERNAME / GITHUB_ACCESS_TOKEN from the master secrets. Public artifacts are pulled anonymously — the platform token is never touched. When the credentials are missing, the error names the secret, not the URL.

⚠️ oras ships in uis-provision-host 1.6.16 and later. Installing an application on an older provision host fails with oras: not found on a machine whose ./uis version may look current — run ./uis pull.

./uis pull — and the version it may not be able to give you

pull fetches ghcr.io/helpers-no/uis-provision-host:latest, a moving tag published by the container build. The update notice compares the installed version against version.txt on main — which main gains the moment a release commit merges, minutes before the image exists, and permanently if that build fails.

So pull can succeed and leave you where you were:

Update available: 1.6.49 -> 1.6.50   (run: ./uis pull)
$ ./uis pull
Image updated successfully
Now running version: 1.6.49

🔴 Since 1.6.51 it says so, and exits 3. After pulling, pull reads back what actually arrived and compares it with main. When they differ it distinguishes three cases, because they need different actions:

what it foundwhat it tells you
<repo>:<version> is not in the registrythe build is still running or it failed — nothing is wrong with your machine, and it names the Actions page
<repo>:<version> is in the registry but :latest is oldera tagging fault in the release; take it by version: UIS_IMAGE=<repo>:<version> ./uis pull
the registry could not be reached"do not know", stated as such — never reported as "not built yet"

🔴 pull, stop and restart refuse while a UIS command is running inside the container (1.6.52). All three stop it, and template install is minutes long — interrupting it leaves a half-built application: a database with a partial schema, or a code location written to .uis.extend that Dagster never loaded. No UIS command repairs that. The refusal names the process it found and says how long to wait; UIS_FORCE=1 ./uis pull overrides it.

⚠️ Exit 3 means the pull worked and did not deliver the advertised version. A script that treats any non-zero as failure will now notice; one that only checks for zero was previously told success. Exit 0 still means you are on the version main advertises.

⚠️ A pulled image is not a running image

docker pull fetches an image. It does not touch a container that is already running, and ./uis start on a running host is a deliberate no-op. So this sequence upgrades nothing and reports success:

$ docker pull ghcr.io/helpers-no/uis-provision-host:latest
$ ./uis start

The host keeps executing the old code, ./uis version is still correct, and every surface agrees the upgrade worked. imac hit this upgrading to 1.6.112 and only avoided a false test result by removing the container first out of habit.

./uis pull and ./uis restart are not affected. Both stop the container first, and starting it again recreates it from the image on disk. The trap is only a pull performed outside the launcher.

🔴 Since 1.6.114 ./uis start says so. When the running container was created from a different image than the one now on disk, it prints both image ids and names the command that applies it:

[UIS] The running container was created from a DIFFERENT image than the one on disk.
[UIS] running from: 3f9a…
[UIS] on disk now: c71b…
[UIS] A pulled image does not replace a running container, so this host is
[UIS] still executing the OLD code. 'uis start' will not change that.
[UIS] Apply it with: ./uis restart (or ./uis pull to fetch and apply)

It is a warning, not a refusal — running a container deliberately pinned to an older image is a legitimate thing to be doing, and start is not the command that should second-guess it.

The registry entry — the seam with the catalogue

The source block above lives in the catalogue, not in the artifact. The artifact carries the definition; the registry carries the pointer to it.

🔴 The shape first, because a list of field names is not a shape. An earlier version of this section gave only the table below, and a reader authored visibility inside source from it — which UIS reads as null and defaults to public, so it works, ships, and keeps working until the first private artifact fails by never asking for credentials (dev-templates, urb-agents#479). Their words for it: "correct by accident is the failure mode neither of us can see." So:

{
"id": "atlas",
"templateKind": "application",
"visibility": "public",
"category": "APPLICATION",
"version": "v20260909-853c696",
"name": "Atlas Data",
"description": "…",
"abstract": "…",
"tags": ["data", "dagster"],

"source": {
"artifact": "ghcr.io/terchris/atlas-data/uis",
"tag": "v20260909-853c696",
"digest": "sha256:def7b9d2…3d3a6c54"
}
}

⚠️ visibility is a sibling of source, not a member of it. Nesting it is accepted silently and defaults to public.

These are the only fields uis template install reads:

fieldrequiredwhat UIS does with it
templateKindyesMust be application. Read as .templateKind // .kind. ⚠️ The registry already has both templateKind and install_type — reuse one, do not add a third discriminator
source.artifactyesThe OCI artifact, by convention <image>/uis. Checked against the allowlist
source.tagyesShown to a human, never pulled by
source.digestyes🔴 What is actually pulled — and it must be AUTHORED, not resolved by the catalogue build. See below
visibilitynopublic (default) or private; decides whether the pull needs oras login. Top-level, beside source
categoryyesMust name a category whose context is uis. An application entry also stays listed on its templateKind alone, so a miscategorised one is visible rather than silently absent
version, name, description, abstract, tagsfor displayWhat uis template info prints

🔴 source.digest must be authored, not resolved at catalogue-build time. A build that re-resolves the tag on every run tracks the tag, so a tag re-pointed at a different artifact is blessed by the next unrelated build — and UIS cannot tell, because it reads whatever the latest published registry says and pins no version of it. Pulling a digest gives integrity (you get what the digest names); only a committed, reviewable digest gives provenance (a human approved this one). dev-templates established this on urb-agents#479 against an earlier claim of mine that UIS could catch it at install time — it cannot, and does not try.

⚠️ The artifact must agree with the entry about its own id. A definition whose id: conflicts with the entry it was fetched for is refused, naming both: installing it would record the application under a name its own definition never claimed.

Everything else in an entry is display metadata for the website and uis template info. In particular params: and provides: come from the artifact's template-info.yaml — the single source of truth — even if the catalogue inlines a resolved copy for the site.

A commented, valid worked example lives at provision-host/uis/tests/fixtures/catalogue/registry-entry.example.json.

Testing an application install with no catalogue at all

REGISTRY_URL_PRIMARY accepts a file:// URL, so a one-entry registry on disk is enough to install a real published artifact before the catalogue carries it:

./uis template install atlas \
REGISTRY_URL_PRIMARY=file:///mnt/urbalurbadisk/my-registry.json

⚠️ Set it on the ./uis command line — docker exec does not inherit the caller's environment, and the launcher forwards these by name.

🔴 Put the file somewhere that survives ./uis pull. /mnt/urbalurbadisk/ is inside the container, which is recreated on every pull, so a registry written there disappears and the symptom is "the entry is not published" rather than "the file is gone". .uis.extend/ is mounted and survives.

Editing that file and re-running takes effect immediately. The registry cache is keyed by the URL, and a file:// source is never cached — read every time, because reading a local file is free and caching it is what makes editing it confusing. A remote registry is still cached for an hour; REGISTRY_CACHE_TTL=0 forces a refetch.

⚠️ A file:// registry that cannot be read refuses; it does not fall back to the catalogue. You asked for that file, so silently resolving a different entry from the published registry would be worse than failing.

init: — a file or an ordered directory

  • a file — applied as-is
  • a directory — every *.sql in it, concatenated in LC_ALL=C sort order

Order is part of the contract: migrations are numbered (001_…, 050_…) because DDL is order-dependent. The count and the ordered file list are printed before anything is applied, so a partial apply is recoverable from the log. Non-.sql files are ignored, and an empty directory fails rather than installing nothing.

{{ params.* }} is substituted into the concatenated content, so a parameter may appear in any file.

Where your application is, after it installs

The completion summary ends with the application's endpoints, taken from its exports::

Endpoints:
api-url http://api-atlas.localhost

🔴 Do not guess this URL from --url-prefix. The route matches on hostnameHostRegexp('api-atlas\..+') — so http://api-atlas.localhost/ answers and http://localhost/api-atlas/ returns a bare Traefik 404. The prefix is the subdomain, not a path.

⚠️ Only what the definition declares in exports: appears here. An application that exports nothing prints no endpoints, and there is currently no other command that will tell you — status, list and verify all omit it.

What a fresh install actually gives you

🔴 An application whose data arrives from a pipeline serves an empty API on day one, and that is correct. template install guarantees that the schema exists, the grants are right, and the API answers — not that there is anything in it. If the application's own orchestrator owns the migrations and the ingest, the first rows appear on its first run, not at install.

This is worth stating because an empty-but-correct API is indistinguishable from a broken install to someone seeing it for the first time, and the instinct is to go looking for the failure. It is also why data freshness belongs in a monitor rather than in uis verify: the platform cannot know when an application's first pipeline run is due.

An application's catalogue entry should say so in its own words — atlas does.

init: on a database that already exists

Installing onto a database that is already there — a re-install, or an application already running on the cluster — takes a different path, and it is worth knowing what it does:

the database and its rolekept, never recreated
the passwordpreserved. It is read back from the Secret UIS wrote. --rotate mints a new one; nothing else does
init:re-applied, and the result reports init_applied
a failing init:refuses and does not drop the database — that data predates the command. ⚠️ It is not left untouched: statements before the failure are already committed, so the schema can be left part-applied. See PLAN-cli-init-file-partial-apply

A role that outlived its database

Removing an application without --purge drops nothing, and dropping the database by hand afterwards leaves the per-app role behind. The next install therefore meets a missing database and an existing role — the create path, with the role already there.

the rolekept, never dropped — it may own objects this command knows nothing about
its password🔴 RESET to the one this install publishes, and the command says so
rollbackif a later step fails, a role that predates the command is never dropped; one this command created is

⚠️ The reset invalidates any other consumer of that role until its pods restart — the same hazard --rotate carries. It is still the right trade: the alternative is publishing a credential that is wrong for everyone.

⚠️ Before 1.6.49 this branch was silent and produced a broken install. CREATE USER failed, the failure was discarded because a guard checked only that the role existed, and the install reported status: ok while writing a password that had never been set on the role. The application came up unable to authenticate (ops, urb-agents#595).

🔴 --rotate will break a running workload until its pods restart and re-read the Secret. Environment-variable consumers — a Dagster code location, for instance — read the credential once at pod start.

⚠️ This used to happen on every re-install, unasked. The rotation existed because "UIS does not store per-app passwords" — but it does, in the Secret it wrote, so it now reads it back instead. A re-install that reports EXIT=0 and leaves the application unable to authenticate is not a trade worth making for a credential nobody asked to change. The JSON reports rotated either way.

⚠️ Re-applying init: is safe by contract, not by luck. An init: must satisfy the schema after one application equals the schema after two. That is a stronger requirement than "each statement is idempotent" — a set of individually-idempotent files can still converge on the wrong schema, which is exactly what happened to the first application's migrations and was only found by measuring pass 1 against pass 2 (urb-agents#362).

⚠️ Zero-pad your numbers. Sorting is lexicographic, so 9_, 10_, 100_ apply in reverse — and out-of-order DDL can succeed while leaving the wrong schema, which is the one failure this ordering contract exists to prevent. UIS warns when numeric and lexicographic order disagree and shows what numeric order would have been, but it applies the lexicographic order regardless: it cannot know which you meant. 001_, 002_, … 010_ is immune.

Testing a template before publishing it

REGISTRY_URL_PRIMARY, REGISTRY_URL_FALLBACK and TEMPLATE_REPO honour an environment override, so a template can be exercised before it reaches the registry:

TEMPLATE_REPO=/path/to/local/dev-templates ./uis template install my-fixture

⚠️ These are forwarded into the container by name. docker exec does not inherit the caller's environment, so a variable the launcher does not forward is silently ignored from the host while working inside the container — set them on the ./uis command line as above and they will arrive.

There is deliberately no fixture template in the registry. It was considered and declined on 2026-09-09: the registry is what uis template list shows a user, so anything in it is something someone may install, and a fixture whose purpose is to exercise edge cases is not that. The local override above covers the testing need it would have served, and an example belongs in documentation — where it can be read without being installable.

Multi-instance services, and the order of operations

./uis deploy receives --app <app_name> automatically for any service whose multiInstance is true in services.json (today: postgrest). configure always receives --app, because a single-instance service can still hold per-app resources — configure postgresql --app creates a per-app database in the shared instance.

⚠️ Which runs first is per-service, and the install prints it. It follows from what multi-instance means:

orderwhy
single-instancedeploy, then configurethe service is shared and already running; configure postgresql execs into the running pod
multi-instanceconfigure, then deploy --appdeploy --app creates the per-app instance and consumes what configure produced. configure cannot want the instance running, because it does not exist yet

A multi-instance service declared with no config: is rejected: a per-app instance would have nothing to consume.

A template does not yet cover every surface

A web frontend has no service to declare: webapp — a multi-instance Deployment + Service + IngressRoute — is Phase 5 of PLAN-templates-002 and is not built. An application whose frontend is a container image still deploys that half through ArgoCD, which is the documented path for workloads anyway; only the route is missing from provides:.


Secrets Management

CommandDescription
./uis secrets initCreate .uis.secrets/ directory with templates
./uis secrets statusShow which secrets are configured vs missing
./uis secrets editOpen secrets config in editor
./uis secrets generateGenerate Kubernetes secrets from templates
./uis secrets applyApply generated secrets to the cluster
./uis secrets validateValidate secrets config and check required values

Testing

CommandDescription
./uis test-allDeploy and undeploy all services (full integration test)
./uis test-all --dry-runShow test plan without executing
./uis test-all --cleanUndeploy everything first, then run tests
./uis test-all --only <svc> [svc...]Test only specified services and their dependencies

Service-Specific Commands

Tailscale and Cloudflare

Tailscale and Cloudflare are managed through the unified uis network family. The legacy uis tailscale <verb> and uis cloudflare <verb> invocations print a redirect stub and exit non-zero.

ArgoCD

CommandDescription
./uis argocd register <name> <repo-url>Register a GitHub repo as ArgoCD application. Name is used as namespace, repo-url must be full HTTPS URL
./uis argocd remove <name>Remove an ArgoCD application and its namespace
./uis argocd listList registered ArgoCD applications with health and sync status
./uis argocd verifyRun ArgoCD health checks

Host Configuration

Manage configurations for different deployment targets.

CommandDescription
./uis host addList available host templates
./uis host add <template-id>Add a host configuration from template
./uis host listList configured hosts with status

Other Commands

CommandDescription
./uis initFirst-time setup wizard (cluster type, domain, project name)
./uis setupInteractive TUI menu for browsing and deploying services
./uis tools listList optional tools with installation status. See Tools.
./uis tools install <tool-id>Install an optional tool (aws-cli, azure-cli, etc.). See Tools.
./uis docs generate [dir]Generate JSON data files for website documentation
./uis versionShow UIS version
./uis helpShow help

Environment Variables

VariablePurposeDefault
UIS_IMAGEOverride container imageghcr.io/helpers-no/uis-provision-host:latest
UIS_KUBECONFIG_DIROverride kubeconfig directory$HOME/.kube