Investigate · workspace

P1rca-readySLA at risk

Why is fs-prod-cs-tool2 NotReady?

Worker node fs-prod-cs-tool2 flipped to NotReady with pods stuck in ContainerCreating.

Opened Sunday, August 2, 2026 at 6:41:00 AM · 2026-08-02T06:41:00Z

Organisation · client scope

Tenant / org
Nordic Federated Bank
tn-nordic · eu-north-1 · EU
Client / customer
FS Core Banking Platform
cu-fsprod · Financial Services · platinum · SLA 99.99%
Environment
production
Client owner
Ingrid Halvorsen

Elapsed

1150h 57m 55s

live clock

Reference

inc-4821

recurrence 3x

Tenant scope

Nordic Federated Bank

FS Core Banking Platform · production

Investigation workspace

Evidence-backed timeline — every step is read-only, bounded and audited. Select a custom date to review previous history, logs, and graphs.

Read-only Agent OS

History window

Aug 2, 06:26 AM → Aug 2, 07:11 AM

Filters timeline, logs, graphs, and workload monitoring

Affected resources

Hostname, IP, application, and cluster identity for SRE / platform triage

3 resources

fs-prod-cs-tool2

worker
Application
calico-node
Hostname
fs-prod-cs-tool2
IP address
10.42.6.21
Cluster
fs-prod-k8s
Namespace
kube-system
Pod
calico-node-7xk2m
Node
fs-prod-cs-tool2
FQDN / endpoint
fs-prod-cs-tool2.fsprod.corp.internal
Region
eu-north-1
Role
worker

registry.corp.internal

registry
Application
corp container registry
Hostname
registry.corp.internal
IP address
198.51.100.44
Cluster
fs-prod-k8s
FQDN / endpoint
registry.corp.internal
Region
eu-north-1
Role
registry

fs-prod-cs-tool1

worker
Application
peer worker nodes
Hostname
fs-prod-cs-tool1
IP address
10.42.6.11
Cluster
fs-prod-k8s
Region
eu-north-1
Role
worker

Status

rca-ready

Application

calico-node / container networking

fs-prod-cs-tool2

Worker node

fs-prod-cs-tool2

IP 10.42.6.21

Cluster

fs-prod-k8s

ns kube-system

Tenant

Nordic Federated Bank

FS Core Banking Platform

Lead agent

Kubernetes Agent 01

Environment: production

RCA confidence

88%

FQDN / endpoint

fs-prod-cs-tool2.fsprod.corp.internal

eu-north-1

Investigation timeline

Severity P1 · 14 of 14 steps in window · expand for formation, load graph, and logs

  1. Detailed formation

    Session opened 2026-08-02 06:41:12 UTC under operator role Platform SRE. Passport ag-kubernetes-01 signature chain verified against trust store rev-2026-07. Tenant boundary tn-nordic and customer cu-fsprod bound to the investigation; production write verbs remain stripped. Guardrail POL-001 (tenant isolation) and POL-004 (no secret read) evaluated → allow-read.

    Logs · filtered by time

    [2026-08-02T06:41:12.012Z] guardrails: begin session incident=inc-4821 agent=ag-kubernetes-01
    [2026-08-02T06:41:12.048Z] passport: signature OK kid=nordic-k8s-01 exp=2026-12-03T00:00:00Z
    [2026-08-02T06:41:12.061Z] scope: tenant=tn-nordic customer=cu-fsprod env=production
    [2026-08-02T06:41:12.088Z] policy: POL-001 allow · POL-004 allow-read · POL-009 deny-write
    [2026-08-02T06:41:12.101Z] guardrails: session VERIFIED · append-only evidence channel open
    • passport signature valid · 2026-08-02T06:41:12.048Z
    • tenant boundary check passed · tn-nordic / cu-fsprod
    • read-only mode confirmed · write verbs=0
  2. Detailed formation

    Read-only journalctl -u kubelet --since '2026-08-02 06:35:00' --until '2026-08-02 06:43:00'. Artefact ev-2. No permission elevation; secrets redacted by POL-004 sanitiser.

    Logs · filtered by time

    2026-08-02T06:38:11.204Z kubelet[1184]: E0814 06:38:11.204112   1184 kuberuntime_manager.go:901] createPodSandbox for pod "calico-node-7xk2m_kube-system" failed: rpc error: code = Unknown desc = failed to pull image "registry.corp.internal/cni/calico-node:v3.27.2": failed to pull and unpack image: read tcp 10.42.6.21:52344->198.51.100.44:443: read: connection reset by peer
    2026-08-02T06:38:11.204Z kubelet[1184]: E0814 06:38:11.204401   1184 pod_workers.go:1300] Error syncing pod 8f2a… (calico-node-7xk2m), skipping: failed to "CreatePodSandbox" for "calico-node-7xk2m_kube-system" with CreatePodSandboxError: "CreatePodSandboxError"
    2026-08-02T06:38:46.118Z kubelet[1184]: W0814 06:38:46.118002   1184 image_pull.go:112] Back-off pulling image "registry.corp.internal/cni/calico-node:v3.27.2"
    2026-08-02T06:39:22.551Z kubelet[1184]: E0814 06:39:22.551440   1184 remote_image.go:238] PullImage "registry.corp.internal/cni/calico-node:v3.27.2" from image service failed: rpc error: code = Unknown desc = connection reset by peer
    2026-08-02T06:41:04.902Z kubelet[1184]: E0814 06:41:04.902771   1184 kubelet_node_status.go:694] Error updating node status, will retry: timed out waiting for lastHeartbeat (last=2026-08-02T06:37:26Z)
  3. Detailed formation

    Artefact ev-3. Layer sha256:6b2f… is 38.4MB; every attempt aborts near 5MB after a successful TLS 1.3 handshake — pattern consistent with mid-stream RST on an inspection appliance, not auth or DNS failure.

    Load graph · Layer transfer progress (MB) and RST count · filtered

    Logs · filtered by time

    2026-08-02T06:38:09.441Z containerd: pulling registry.corp.internal/cni/calico-node:v3.27.2@sha256:a91e…
    2026-08-02T06:38:09.512Z containerd: resolving host=registry.corp.internal → 198.51.100.44
    2026-08-02T06:38:09.630Z containerd: TLS handshake OK (TLSv1.3, cipher=TLS_AES_256_GCM_SHA384, 118ms)
    2026-08-02T06:38:09.701Z containerd: fetch layer sha256:6b2f… size=38.4MB started
    2026-08-02T06:38:14.228Z containerd: transfer aborted at 5.2MB/38.4MB — connection reset by peer (errno=104)
    2026-08-02T06:39:01.884Z containerd: retry 2 · aborted at 4.9MB/38.4MB — connection reset by peer
    2026-08-02T06:39:57.103Z containerd: retry 3 · aborted at 5.1MB/38.4MB — connection reset by peer
    2026-08-02T06:40:03.440Z containerd: TLS handshake OK, stream terminated by remote before Content-Length complete
    2026-08-02T06:40:03.441Z containerd: note: pull of registry.corp.internal/base/pause:3.9 via internal mirror SUCCEEDED (12.1MB in 1.4s)
  4. Detailed formation

    PromQL window 2026-08-02 06:30–06:44 UTC. Host utilisation flat; only pull_image error counter climbs. Confirms resource-exhaustion hypothesis is unsupported before formal rejection at s10.

    Load graph · Host load % and PullImage errors · filtered

    Logs · filtered by time

    # capturedAt: 2026-08-02T06:44:10.220Z  node=fs-prod-cs-tool2
    node_cpu_utilisation{node="fs-prod-cs-tool2"}                     0.21
    node_memory_utilisation{node="fs-prod-cs-tool2"}                  0.48
    node_filesystem_used_ratio{node="fs-prod-cs-tool2"}               0.39
    container_runtime_operations_errors_total{operation="pull_image"} 9
    container_runtime_operations_errors_total{operation="create_container"} 0
    rate(container_runtime_operations_errors_total{operation="pull_image"}[10m]) 0.015/s

Final root cause

Confidence 88% · risk low · no production write required

Registry egress traffic from fs-prod-cs-tool2 is being reset mid-transfer, most likely by SSL inspection on the outbound path introduced in change CHG-20482. Container image layers for the CNI plugin cannot complete, so the container runtime network never becomes ready and the node reports NotReady.

Recommendation

Validate outbound TCP 443 connectivity from the node subnet to the registry egress range and confirm SSL-inspection exclusions cover registry.corp.internal and the upstream mirror. Re-run the image pull after the exclusion is verified.

Open full RCA

Opened Sunday, August 2, 2026 at 6:41:00 AM · inc-4821