Write-up published
Resolved
During the overnight hours of June 30, 2026, an incident caused all alerts on our primary server to fail during the image generation process. A cascading failure then occurred, causing new alerts to be skipped as image generation workers repeatedly crashed.
Our investigation identified two root causes, both of which have been addressed:
A GPU driver entered an unhealthy state during a routine refresh of the long-term processor. That invalid driver state was mistakenly placed back into the worker queue, causing every subsequent image generation task to fail immediately.
Our automated recovery system failed to detect the unhealthy driver state and did not trigger a forced refresh. As a result, what should have been a brief interruption became a prolonged outage.
We are confident both issues have been resolved. We have also implemented changes to improve our recovery logic and monitoring so similar failures are automatically detected and corrected much more quickly.
This outage did not meet the uptime commitment outlined in our SLA, and we take full responsibility for that. Affected customers are eligible for SLA service credits. Please contact support@forecastix.xyz to request your credit.
We fell short of the reliability our customers expect. We are committed to learning from this incident and strengthening our systems to ensure they recover correctly and notify our team immediately if a similar issue ever occurs.
Identified
We identified a bug in how our processors handle failed map loads. Under certain conditions, this created a cascading failure across both processors by placing an unhealthy driver state back into the worker queue. As additional workers attempted to load maps, they would crash one by one, eventually stopping image generation.
A fix is currently being developed. In the meantime, all missed alerts have been processed, and alert posting has returned to normal. We will continue investigating and testing until we are confident this driver issue has been permanently resolved.
Investigating
Our long-term alert processor is recovering and working through a backlog of alerts after a GPU driver issue on the server caused image generation to fail. All missed alerts are now being processed and will continue to be delivered as the backlog clears. We apologize for the inconvenience and will be investigating the root cause to prevent this from happening again. Short-term convective alerts were not affected.