Kanzo UI
AI

The gateway

The server half of @kanzo-tech/llm — LiteLLM as the platform's one door to models, reached by alias through an application's own server, with a team and a key per tenant.

@kanzo-tech/llm names a model by alias behind a gateway. services/ai/, in this repository and under the same release tag, is that gateway: LiteLLM, which holds the providers' keys, maps each alias to an upstream and speaks OpenAI's chat-completions protocol whatever the upstream speaks. Each capability ships both halves.

Two aliases

An application names an alias, never a provider:

AliasFor
chatconversations and agents — Chat, a ToolLoopAgent
completeshort completions asked on every pause in typing — Assist's ghost text and suggestions

Which upstream answers each one is the gateway's profile, so the same gateway("chat") runs on a laptop's GPU in development and on a provider in production. Changing models is an edit to the profile, not a release of every application that calls it.

The key stays on a server

browser ──▶ the application's server (BFF) ──Bearer <tenant key>──▶ gateway ──▶ upstream

The browser never holds a key. It calls its own application's route — the baseURL it hands createGateway — and that server checks the session, checks the alias, caps the tokens, names the caller, injects its tenant's key and streams the answer back. Why the key is never in the page has the route.

Run it

From services/ai/:

docker compose up -d --wait   # the gateway on http://localhost:4000, both aliases answering

The development profile, litellm.dev.yaml, serves both aliases from open models through Docker Model Runner, on the host's GPU — Docker Desktop 4.40 or later, with Model Runner turned on. Compose declares the models and points the gateway at them. The first up pulls them, about 6.5 GB; after that the loop is offline and costs nothing. The admin console is /ui, admin / sk-dev-master-key.

An application's own compose includes services/ai/compose.yml the way it includes the identity service, from a checkout pinned to the release of its packages. The containers are ai-gateway, ai-postgres and ai-cache, so nothing of the application's collides with them.

A team per tenant

Spend, limits and logs belong to a tenant, not to the application. The module at services/ai/modules/team gives one tenant a LiteLLM team, carrying its budget and the aliases it may call, and one service-account key in it — the key that tenant's server presents. An application instantiates it from its own repository, once per tenant:

infra/ai.tf
module "acme" {
  source     = "git::https://github.com/Kanzo-Tech/ui.git//services/ai/modules/team?ref=v0.30.0"
  team_id    = "board-acme"
  alias      = "board — acme"
  max_budget = 50 # USD per 30 days; null is unlimited
}

module.acme.key is sensitive, and the application's server is the only thing that reads it. The module drives LiteLLM's management API through the litellm provider, configured with the gateway's URL and its master key.

Deploying

services/ai/modules/gateway runs the gateway, its Postgres and its cache as Swarm services on a network the deployment owns; applications on that network reach it at its url output. The deployment passes what is its own:

  • profile — its LiteLLM config: which upstream answers each alias. A production profile is the deployment's, not this repository's; the development one shows the shape. The config is named by its hash, so changing it rolls the gateway and no application changes.
  • upstream_keys — env name to key, for every upstream the profile names.
  • admin_hostname and admin_allow — optionally, an IP-allowlisted route to the console and to the management API that modules/team drives. Without them there is no route at all.

The module's master_key output is that management key, and it is never an application's.

The image is pinned by digest, in the compose file and in the module alike. Two LiteLLM releases, 1.82.7 and 1.82.8, were published compromised in March 2026: move the pin deliberately, to a release that has been out for a few days, never by tag.

Caching

The cache is default_off. A request opts in where the same input is often asked twice — complete, the same field value on the same pause — and a conversation turn never does, or Retry would answer the same thing again.

On this page