During the overnight hours of June 30, 2026, an incident caused all alerts on our primary server to fail during the image generation process. A cascading failure then occurred, causing new alerts to be skipped as image generation workers repeatedly crashed.
Our investigation identified two root causes, both of which have been addressed:
A GPU driver entered an unhealthy state during a routine refresh of the long-term processor. That invalid driver state was mistakenly placed back into the worker queue, causing every subsequent image generation task to fail immediately.
Our automated recovery system failed to detect the unhealthy driver state and did not trigger a forced refresh. As a result, what should have been a brief interruption became a prolonged outage.
We are confident both issues have been resolved. We have also implemented changes to improve our recovery logic and monitoring so similar failures are automatically detected and corrected much more quickly.
This outage did not meet the uptime commitment outlined in our SLA, and we take full responsibility for that. Affected customers are eligible for SLA service credits. Please contact support@forecastix.xyz to request your credit.
We fell short of the reliability our customers expect. We are committed to learning from this incident and strengthening our systems to ensure they recover correctly and notify our team immediately if a similar issue ever occurs.