Pingdom customer impact
External signal1 active item(s) in this window.
Generated 2026-07-19 22:19 for 2026-07-11 07:00 to 2026-07-18 07:00 from Pingdom checks, Slack #_alerts_prod, and AWS SNS alerts.
Bottom line: Pingdom observed recent customer-facing glitches (email unconfirmed) and application-level critical paths are present.
1 active item(s) in this window.
5 critical, 9 non-critical active item(s).
No AWS alarm emails were captured in this window.
No active issue listed in this category.
| Pingdom Check | Status | Events | Downtime | Last Seen | Likely Services | Correlated Evidence |
|---|---|---|---|---|---|---|
| Adservio Ro | Recovered recently | 1 | 0m | 2026-07-17 18:05 | unclassified | Pingdom-only evidence so far |
Pingdom rows show externally visible signal first. The correlated evidence column helps tie the failing check back to services, Slack alert families, or AWS alarms when those links exist.
This view attributes alerts to the workload or resource named in the alert text. Grafana, Loki, and Tempo are treated as observability components and are excluded when a more specific impacted target is also present.
| Impacted Service / Resource | Highest Severity | Count | Last Seen | Status | Top Alert Types | Discussion Signal | Latest Thread Note |
|---|---|---|---|---|---|---|---|
| grafana | Critical | 17 | 2026-07-17 23:05 | Seen today | PlatformLatencyP95Critical1s (1)PagerDutyTest (1)Watchdog (6)KubeHpaMaxedOut (2)PlatformLatencyP95Warning400ms (1) | None | No thread note |
| accommodations-api Grouped 4 variantsVariant mentions 14Active variants 4 | Critical | 14 | 2026-07-16 01:04 | Recent (72h) | TraefikServiceHighErrorRate (10)KubePodContainerRestartingFrequently (2)KubePodCrashLooping (1)KubeDeploymentReplicasMismatch (1) | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
| etcd | Critical | 7 | 2026-07-15 18:50 | Recent (72h) | etcdInsufficientMembers (1)TargetDown (3)etcdMembersDown (3) | Release / migration issueGeneral investigation | tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket |
| web-80 | Critical | 2 | 2026-07-16 01:00 | Recent (72h) | TraefikServiceHighErrorRate (1)TraefikServiceHighLatency (1) | None | No thread note |
| core-grafana-80 | Critical | 1 | 2026-07-16 12:21 | Recent (72h) | TraefikServiceHighErrorRate (1) | General investigation | This should go away caused due codex agent running queries |
| admission-api | Warning | 19 | 2026-07-17 22:31 | Seen today | TraefikServiceHighLatency (18)RESOLVED - TraefikServiceHighLatency (1) | None | No thread note |
| core-grafana | Warning | 2 | 2026-07-17 16:23 | Seen today | KubePodContainerRestartingFrequently (2) | None | No thread note |
| library-api Grouped 5 variantsVariant mentions 8Active variants 5 | Warning | 5 | 2026-07-16 01:12 | Recent (72h) | KubePodCrashLooping (2)KubePodContainerRestartingFrequently (2)KubeDeploymentReplicasMismatch (1) | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
| rooms-api Grouped 2 variantsVariant mentions 3Active variants 2 | Warning | 4 | 2026-07-16 01:04 | Recent (72h) | KubePodContainerRestartingFrequently (2)KubePodCrashLooping (1)KubeDeploymentReplicasMismatch (1) | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
| send-codes Grouped 2 variantsVariant mentions 3Active variants 2 | Warning | 3 | 2026-07-16 05:25 | Recent (72h) | KubeJobFailed (2)KubePodContainerRestartingFrequently (1) | None | No thread note |
| metrics-server | Warning | 3 | 2026-07-15 18:49 | Recent (72h) | KubeAggregatedAPIDown (3) | None | No thread note |
| colecteaza-sms-note-abs | Warning | 1 | 2026-07-16 01:04 | Recent (72h) | KubePodContainerRestartingFrequently (1) | None | No thread note |
| download-album | Warning | 1 | 2026-07-16 01:04 | Recent (72h) | KubePodContainerRestartingFrequently (1) | None | No thread note |
| rezumat | Warning | 1 | 2026-07-16 01:04 | Recent (72h) | KubePodContainerRestartingFrequently (1) | None | No thread note |
| Alert | Severity | Count | Last Seen | Status | Threads | Top Impacted Services | Discussion Signal | Latest Thread Note |
|---|---|---|---|---|---|---|---|---|
| TraefikServiceHighErrorRate | Critical | 12 | 2026-07-16 12:21 | Recent (72h) | 1 | accommodations-api (10)web-80 (1)core-grafana-80 (1) | General investigation | This should go away caused due codex agent running queries |
| PlatformLatencyP95Critical1s | Critical | 1 | 2026-07-16 01:05 | Recent (72h) | 0 | grafana (1) | None | |
| PagerDutyTest | Critical | 1 | 2026-07-15 10:41 | Recent (72h) | 0 | grafana (1) | None | |
| etcdInsufficientMembers | Critical | 1 | 2026-07-15 10:40 | Recent (72h) | 1 | etcd (1) | Release / migration issue | Looks good i did a helm deployed on tuiasi that rolled out alert manger pod and these alerts are triggered all good on tuiasi cluster . | false positive alerts. |
| TraefikServiceHighLatency | Warning | 19 | 2026-07-17 22:26 | Seen today | 0 | admission-api (18)web-80 (1) | None | |
| Watchdog | Warning | 6 | 2026-07-17 23:05 | Seen today | 0 | grafana (6) | None | |
| KubePodContainerRestartingFrequently | Warning | 4 | 2026-07-17 16:23 | Seen today | 0 | accommodations-api (2)library-api (2)rooms-api (2)core-grafana (2)colecteaza-sms-note-abs (1) | None | |
| RESOLVED - TraefikServiceHighLatency | Warning | 1 | 2026-07-17 22:31 | Seen today | 0 | admission-api (1) | None | |
| TargetDown | Warning | 3 | 2026-07-15 18:50 | Recent (72h) | 1 | etcd (3) | General investigation | tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket |
| etcdMembersDown | Warning | 3 | 2026-07-15 18:50 | Recent (72h) | 0 | etcd (3) | None | |
| KubeAggregatedAPIDown | Warning | 3 | 2026-07-15 18:49 | Recent (72h) | 0 | metrics-server (3) | None | |
| KubeJobFailed | Warning | 2 | 2026-07-16 05:25 | Recent (72h) | 0 | send-codes (2) | None | |
| KubePodCrashLooping | Warning | 2 | 2026-07-16 01:12 | Recent (72h) | 1 | library-api (2)accommodations-api (1)rooms-api (1) | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
| KubeHpaMaxedOut | Warning | 2 | 2026-07-16 01:12 | Recent (72h) | 0 | grafana (2) | None | |
| PlatformLatencyP95Warning400ms | Warning | 1 | 2026-07-16 01:06 | Recent (72h) | 0 | grafana (1) | None | |
| KubeAPIErrorBudgetBurn | Warning | 5 | 2026-07-14 11:53 | Seen this week | 0 | grafana (5) | None | |
| AlertmanagerFailedToSendAlerts | Warning | 1 | 2026-07-13 07:42 | Seen this week | 0 | grafana (1) | None | |
| KubeDeploymentReplicasMismatch | Warning | 1 | 2026-07-13 07:41 | Seen this week | 1 | accommodations-api (1)library-api (1)rooms-api (1) | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
Status is heuristic. Slack rarely posts explicit resolutions, so “Seen today” or “Recent” means the alert family still appeared in production recently, not that it is definitely unresolved.
| AWS Alarm | Emails | ALARM | OK | State Flips | First Seen | Last Seen | Latest State | Status |
|---|
“Flapping, latest OK” means the most recent email was an OK, but the alarm toggled repeatedly and is still a reliability concern.
| Thread Date | Alert | Severity | Services | Signal | Key Notes |
|---|---|---|---|---|---|
| 2026-07-16 12:21 | TraefikServiceHighErrorRate | Critical | core-grafana-80 | General investigation | This should go away caused due codex agent running queries |
| 2026-07-15 14:45 | TargetDown | Warning | etcd | General investigation | tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket |
| 2026-07-15 10:40 | etcdInsufficientMembers | Critical | etcd | Release / migration issue | Looks good i did a helm deployed on tuiasi that rolled out alert manger pod and these alerts are triggered all good on tuiasi cluster . | false positive alerts. |
| 2026-07-13 07:41 | KubePodCrashLooping | Warning | accommodations-api, library-api, rooms-api | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |
| 2026-07-13 07:41 | KubeDeploymentReplicasMismatch | Warning | accommodations-api, library-api, rooms-api | General investigation | Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a… |