Processing is delayed
Resolved
Jul 1, 2026 at 8:49pm UTC
Note: This is the final incident report. We will update this page if anything changes. Last updated: 6 July, 2026.
Incident: Temporary Degradation to View Processing
On Monday–Wednesday, 29 June–1 July, Enterspeed experienced a temporary degradation to view processing following an autoscaling configuration change to Enterspeed Processing.
What Happened
On Monday, 29 June at 14:30 CEST, an edit made through the Azure Portal to manage our processing infrastructure triggered a UI glitch that reset internal scaling rules for one of our processing components. The Portal saves configuration changes based only on what's visible on screen, and some of our scaling settings were not shown in the interface — so the save silently dropped them. At the time, this appeared as a cosmetic display issue rather than an actual configuration change, and it was not recognized as a real change until its effects surfaced later.
Without these scaling rules in place, the affected component could no longer scale to handle incoming processing volume, and a backlog began building later that evening. The issue was reported the next day, and an investigation began. Checking processing volume was the natural first step — but volume was low precisely because of the reduced capacity from the earlier configuration change. This was read as the system being under no unusual pressure, rather than recognized as a symptom of that underlying problem, so the root cause wasn't identified at that point.
Separately, we found a gap in our monitoring coverage. An additional processing layer, which fairly distributes jobs across tenants before they reach our processing queues, had been added upstream of our existing monitoring. This layer did not have its own dedicated alerting, so a backlog could build up there without being detected — even though monitoring on the queue itself continued to work correctly. A separate, unrelated alert did fire on the morning of 1 July, but it arrived shortly after three routine alerts for scheduled SQL server maintenance and was not acted on as a result.
The following morning, 1 July, the effects became fully visible as a backlog in our processing queue.
Impact
- Processing of new content updates was delayed starting Monday, 29 June at 14:30 CEST.
- No data was lost.
- All affected jobs were re-queued and successfully processed.
- The full backlog was cleared by 19:07 CEST on Wednesday, 1 July.
- Throughout the incident, there was no downtime to the Delivery API — your applications continued to receive responses as normal.
How We Responded
Once the root cause was confirmed on Wednesday morning, we corrected the autoscaling configuration and increased processing capacity to clear the backlog. Increasing capacity created a secondary, smaller backlog downstream; resolving it fully required an additional scaling adjustment, completed at 19:07 CEST the same day. In parallel, we manually deduplicated queued jobs for tenants with the largest backlogs to speed up recovery.
What We Are Doing to Prevent This
- We are changing how this processing infrastructure is configured, removing the ability for this type of change to be made directly through the Azure Portal, in favor of reviewed, automated deployment processes.
- When investigating a similar issue in the future, we will explicitly check whether low processing volume is itself a symptom of a configuration problem, rather than treating it as a sign the system is healthy.
- We are adding dedicated monitoring to our fairness processing layer, so a backlog building up there is caught directly, rather than only being visible once it reaches a downstream queue.
- We are reviewing how routine, expected alerts (such as scheduled maintenance notifications) are distinguished from alerts that require immediate action, so a genuine signal isn't missed amid expected noise.
Questions or Concerns?
If you have any questions about this incident or its impact, we're here to help — please reach out in our Slack support channel or email us at support@enterspeed.com.
We apologise for the disruption this incident caused and appreciate your understanding.
Affected services
Updated
Jul 1, 2026 at 5:01pm UTC
The downstream residual issue has been resolved.
Affected services
Updated
Jul 1, 2026 at 3:47pm UTC
We identified a downstream residual issue that for the next 15-20 minutes causes a delay in the processing of the last jobs.
Affected services
Updated
Jul 1, 2026 at 1:57pm UTC
All tenants are now done processing and all delays have cleared.
Please reach out via your Slack support channel or support@enterspeed.com if you have any questions.
Affected services
Updated
Jul 1, 2026 at 12:47pm UTC
We have monitored the system through the day, and we can now provide more information.
We learned that the issue originated Monday late in the afternoon and it slowly built up from a small, unnoticeable delay to a big multi-minute delay Wednesday morning.
Shortly after the fix was implemented this morning, most tenants were back with normal processing rates. A couple of tenants have been suffering with delays throughout the day. We have been in direct contact with the affected tenants.
We sincerely apologize for the delays. We will update with further information when the processing is back to normal for all tenants.
Affected services
Updated
Jul 1, 2026 at 8:07am UTC
The fix has been verified in production, and we are monitoring the platform. The initial results are looking as expected and are showing good results.
We expect the backlog to be processed during the next hours. Reach out for a more specific estimate for your tenant.
Affected services
Created
Jul 1, 2026 at 7:05am UTC
We have identified an issue causing processing delays. We are in the process of applying a fix and are investigating the cause.
Affected services