Platform · scoped estate

Command Centre

Metrics, events, logs and traces for the selected tenant / client / environment — dig into evidence when something breaks. Read-only; no autonomous remediation.

Health

85

estate index

P1 open

2

escalated

Agents

4

2 high-risk

Read-only Agent OS
  • No shell execution
  • No cluster-admin
  • No secret reads
  • No database writes
  • No firewall changes
  • No autonomous remediation

Observability signals

Metrics · events · log pipelines · traces — unified fleet view
  • Nodes

    298

    infra

  • Clusters

    13

    infra

  • Agents

    30

    APM

  • Error events

    2

    events

  • P2

    1

    events

  • SLA risks

    3

    traces

  • Log drains

    4

    logs

  • Approvals

    3

    gates

Live telemetry

Fleet metrics & monitors

Streaming timeseries widgets and threshold monitors — Datadog-style dashboard layout, vendor-neutral simulated feed (1.5s ticks).

LIVEupdated now

CPU

41.9%

fleet avg

Latency p95

239ms

model gateway

Error rate

0.39%

production

Throughput

834rps

agent invoke

Agents busy

45%

active workers

CPU utilization

Agent hosts · last ~60s

Request latency

Gateway p95 · ms

Live event stream

Rolling ingest from collectors & agents

  • Agent sidecar scrape completed

    04:55:08

  • Trace sample ingested · investigation span

    04:55:07

  • Approval SLA tick · pending queue

    04:55:14

  • Approval SLA tick · pending queue

    04:55:17

Error rate

Failed invokes / total

Throughput

Requests per second

Active monitors

Threshold checks on live series — monitors-as-code pattern (no vendor lock-in)

  • Fleet CPU anomaly

    ok

    avg(last_5m):cpu.utilization{scope:agents}

    41.9% · thr > 85%

  • Gateway latency p95

    ok

    avg(last_5m):gateway.latency.p95

    239ms · thr > 400ms

  • Error rate spike

    ok

    sum(last_5m):errors.rate{env:production}

    0.39% · thr > 2.5%

  • Log pipeline lag

    ok

    avg(last_5m):pipeline.lag.p95

    1.6s · thr > 3s

Telemetry pipeline

Collectors forward structured container logs from agent sidecars — ops pattern analogous to Fluent Bit DaemonSets tailing /var/log/containers/*.log.

  • Log collectors (DaemonSet)

    4/4 nodes

    healthy
  • Structured JSON parse

    cri-o · containerd

    healthy
  • Export endpoint

    EU residency

    healthy
  • Pipeline lag p95

    1.4s · live

    healthy
Open evidence / log viewer

Application latency

Model gateway request performance — live p95 overlay on seeded providers

OpenAI

203 ms live · US / EU routing

healthy

Anthropic

234 ms live · US / EU routing

healthy

Google Gemini

265 ms live · US

degraded

Azure OpenAI

297 ms live · EU (Sweden Central)

healthy

AWS Bedrock

328 ms live · EU (Frankfurt)

healthy

Ollama (on-prem)

359 ms live · On-premise

healthy

vLLM Cluster

390 ms live · On-premise (sovereign)

healthy
Open Model Gateway

Infrastructure health heatmap

Composite health by customer and environment — see everything in one place

Customerproductionstagingdevdr
FS Core Banking Platform94898468
Nordic Payments Rail99998479
Card Issuing Services99938883
Grid Telemetry Fabric88837873
SCADA Edge Estate85807570
Clinical Data Platform93888378
Imaging AI Workloads99988277
National Registry Services96918670

Error & incident timeline

Historical event volume (seeded) — live series above for last-minute fleet health

Token spend and retry waste

USD per day across all tenants