Production deploy: qkt-prod¶
An example runbook for a managed Dokploy deployment. Replace <prod-host>,
<deploy-user>, and <deploy-root> with values from your private operations
inventory. For a fresh Docker Compose host, start with
Deploy with Docker.
Stack overview¶
| Piece | Where |
|---|---|
| Host | <prod-host> (SSH as a least-privileged <deploy-user>) |
| Orchestrator | Dokploy on Docker Swarm |
| qkt image | ghcr.io/elitekaycy/qkt:<tag> (built by .github/workflows/release.yml on tag push) |
| MT5 gateway image | elitekaycy/mt5-gateway-api@sha256:... (pinned by digest in compose.yml) |
| Compose source | Private deployment repo cloned at <deploy-root> |
| qkt state | ./state/ bind mount → /var/lib/qkt inside container |
| qkt logs | ./logs/ bind mount → /var/lib/qkt/logs inside container |
| Dokploy state | dokploy-postgres container (project env vars, deploy history) |
QKT_IMAGE_TAG is the single source of truth for what's deployed. It is set
in Dokploy's per-project env (UI) and written into <deploy-root>/.env
on each deploy. The compose file requires it to be set (${QKT_IMAGE_TAG:?...});
there is no fallback — a missing env fails the deploy loud.
"What's deployed right now?"¶
Three ways, in order of preference:
# 1. Ask the running binary. This is the version-drift fix's whole point:
ssh <deploy-user>@<prod-host> 'docker exec qkt qkt --version'
# qkt <version> (<git-sha>) built <timestamp>
# 2. Read the running image tag:
ssh <deploy-user>@<prod-host> "docker inspect qkt --format '{{.Config.Image}}'"
# ghcr.io/elitekaycy/qkt:v<version>
# 3. Read the Dokploy-managed env (canonical, but requires SSH):
ssh <deploy-user>@<prod-host> 'grep QKT_IMAGE_TAG <deploy-root>/.env'
(1) is the operator default. The version string maps directly to a git SHA, so
git show <git-sha> tells you exactly what is running.
Releasing a new version¶
Tagging triggers the image build; bumping QKT_IMAGE_TAG in Dokploy triggers
the redeploy. The qkt-prod repo no longer needs a "bump" commit per release —
the image tag lives in Dokploy's env, not in compose.yml.
-
Tag the release on the
qktrepo:git checkout main && git pull # Update VERSION file to the new version, commit, push, then: git tag v<version> && git push origin v<version>The
release.ymlworkflow builds and pushesghcr.io/elitekaycy/qkt:v<version>. Wait for the workflow to finish before continuing. -
Bump
QKT_IMAGE_TAGin Dokploy:- Dokploy UI → project → Environment → set
QKT_IMAGE_TAG=v<version>. - Click "Deploy" (or use the API).
- Dokploy UI → project → Environment → set
-
Verify:
Rolling back¶
Same path, in reverse. Set QKT_IMAGE_TAG back to the prior version and redeploy:
The image is already in GHCR (releases are immutable), so the redeploy is just a container restart against the existing image. ETA: ~30s.
If you don't know the prior version, check the deploy history in the Dokploy UI, or the image label of any saved container:
Applying a strategy change¶
Use deploy only for a new daemon name. For an edit to an already-running
strategy or portfolio, resync the same name so unrelated strategies keep
running and a failed replacement leaves the previous session registered:
ssh <deploy-user>@<prod-host> \
'docker exec qkt qkt resync /strategies/<strategy>.qkt --as <strategy> --dry-run'
ssh <deploy-user>@<prod-host> \
'docker exec qkt qkt resync /strategies/<strategy>.qkt --as <strategy>'
ssh <deploy-user>@<prod-host> 'docker exec qkt qkt status <strategy>'
Production gates are evaluated on resync the same way they are on deploy. Use a waiver only with an explicit reason:
docker exec qkt qkt resync /strategies/<strategy>.qkt --as <strategy> \
--waive all --reason "emergency broker-side fix"
Tailing logs¶
# qkt daemon stdout (per-strategy events, control plane errors):
ssh <deploy-user>@<prod-host> 'docker logs -f --tail 200 qkt'
# Per-strategy file logs (logback-rotated, mounted at ./logs/):
ssh <deploy-user>@<prod-host> 'tail -F <deploy-root>/logs/<strategy>.log'
# MT5 gateway:
ssh <deploy-user>@<prod-host> 'docker logs -f --tail 200 qkt-mt5'
The container json-file driver is capped at 10 MB × 5 files per service, so
docker-level logs are bounded. The per-strategy logback appenders rotate
independently (50 MB / 7-day) and live on the bind mount, so they survive
container rebuilds.
State recovery¶
The ./state/ bind mount under the compose directory holds the daemon's
control-port file, strategy state snapshots, and trade history. The
State backup runbook covers cadence and restore.
If you nuke the container but keep ./state/, the daemon restarts in the same
state. If you lose ./state/ you lose the trade-history ring buffer and any
pending in-flight orders — see the backup runbook for the restore path.
MT5 broker login¶
The bundled MT5 gateway (qkt-mt5 container) runs MetaTrader 5 in a VNC
desktop. If MT5 logs out (broker session timeout, credential change), the
daemon's qkt brokers list will show gateway: down.
# SSH tunnel the VNC port locally:
ssh -L 3020:127.0.0.1:3020 <deploy-user>@<prod-host>
# Then open http://127.0.0.1:3020 in a local browser, log in to MT5.
VNC password lives in MT5_VNC_PASSWORD in Dokploy env.
Required: daily reconciliation¶
qkt reconcile <strategy> compares the engine's book against the broker's
(per-symbol net positions + equity both sides) and exits non-zero on any delta.
Run it daily from cron and page on failure — slow silent state drift is exactly
the bug class a human only notices by accident:
0 6 * * * /usr/local/bin/qkt reconcile mystrategy --json || curl -fsS "https://api.telegram.org/bot$TOKEN/sendMessage" -d chat_id=$CHAT --data-urlencode text="qkt reconcile DELTA — check immediately"
Every order event is also journaled append-only under state/journal/<strategy>/
(one JSONL per UTC day, fsynced) — the reconstructible audit trail for "what
exactly happened, in order".
Margin floor and stop-out levels¶
risk.margin_floor_pct (default 200) blocks NEW entries while the venue margin
level sits below it; risk-reducing orders always pass. Set it against the broker's
REAL stop-out level, not a guess — MT5 brokers force-close the largest losers when
margin level hits their stop-out, the broker's choice of position, during the exact
volatility spike that caused it:
| Broker profile | Margin call | Stop-out |
|---|---|---|
| Exness (standard) | 60% | 0% (!) — Exness liquidates at 0, but gaps can blow through |
| ICMarkets | 100% | 50% |
| FTMO | n/a (drawdown rules instead) | account breach rules |
| Pepperstone | 90% | 50% |
Verify the figures for YOUR account type in the broker terminal before trusting them; they vary by account tier and jurisdiction. A 200% floor keeps roughly 2x coverage — the practitioner norm. Weekend/news exposure reduction is tracked in
398 and not yet automated: reduce manually before weekends until it ships.¶
Required: external deadman watchdog¶
The daemon cannot report its own death — a host reboot or OOM on a Friday evening
goes unnoticed until someone looks, with open leveraged positions protected only
by venue-side stops. Run scripts/deadman-watchdog.sh from a SECOND machine via
cron (every minute), pointed at the daemon's /health through an SSH tunnel or
private network:
* * * * * QKT_HEALTH_URL=http://127.0.0.1:8200/health TELEGRAM_BOT_TOKEN=... TELEGRAM_CHAT_ID=... /opt/qkt/deadman-watchdog.sh
It pages once per outage (daemon down, or any running strategy silent past 15 minutes — the wedged-session signature) and once on recovery. A hosted dead-man service (healthchecks.io) works as an alternative; the point is that the checker does not share the trading host's fate.
Trust boundary and host clock¶
Two assumptions the deployment relies on — written down so a future change can't silently break them:
- Control plane and gateway are loopback-only and unauthenticated. The daemon's
control HTTP (deploy/halt/kill/flatten) binds
127.0.0.1, and the mt5-gateway HTTP has no auth of its own. That is safe exactly as long as nothing rebinds them to a routable interface or port-forwards them. Anyone who can reach those ports can submit orders. If remote control is ever needed, put an authenticated reverse proxy in front — do not just change the bind address. - The host clock must be NTP-disciplined (
chronyon the prod host).SCHEDULEfiring and the daily-PnL UTC-midnight rollover use host time; candle boundaries use broker tick time and are immune. A drifting clock fires schedules at the wrong time and rolls daily loss budgets early or late.
Disaster: full host loss¶
- Provision a new host (any Docker Swarm / Compose-capable Linux).
- Install Dokploy from the official install script.
- Restore Dokploy's
dokploy-postgresfrom backup → all project env vars (includingQKT_IMAGE_TAG) come back. - Reconnect the
qkt-prodGitHub repo to Dokploy → it re-clones the compose source. - Restore the
state/bind mount fromstate-backup. - Trigger deploy → container starts on the same image tag as before.
- Log back in to MT5 via VNC (broker session is stored in
mt5_confignamed volume — if that's lost, full re-login is required).
The container itself is stateless; everything that matters is in ./state/ and
Dokploy's postgres.
See also¶
- Deploy with Docker — generic compose deployment, not Dokploy-specific.
- State backup — backup/restore cadence for
./state/. - Monitoring — what to scrape, what to page on.
- Troubleshooting — symptom → cause → fix.