On October 5, 2026, between 12:18 and 12:31 PM ET, Qlty stopped completing analysis results for 13 minutes. About 120 builds across roughly 60 workspaces never completed. On the affected pull requests, the Qlty check on GitHub showed "Qlty encountered a timeout error at 15mins" while the build page on Qlty showed the build as passed. Pushing a new commit to an affected pull request produces a fresh, correct result.
We are sorry for the confusion this caused. Below is what happened, who was affected, why, and what we are changing.
Timeline (US Eastern on October 5, 2026, UTC in parentheses)
12:17 PM (16:17): A routine data cleanup in one workspace started a large maintenance operation on our main database.
12:18 PM (16:18): Analysis results stopped completing. Builds kept running and finishing normally, but their results were not written.
12:30 PM (16:30): Our queue-delay alarm fired.
12:31 PM (16:31): The operation finished and results started completing again.
12:33 to 12:50 PM (16:33 to 16:50): The Qlty check on affected pull requests changed from pending to the 15-minute timeout error.
12:35 PM (16:35): Result processing was back to normal.
About 1:00 PM (17:00): A customer reported a pull request stuck on the timeout error. We advised pushing a new commit.
4:00 PM (20:00): We confirmed the root cause.
Who was affected
During the 13-minute window, every analysis result was delayed by up to 12 minutes. Most builds recovered on their own: our queue retries each result three times, and 73 builds completed on a later attempt.
118 builds did not complete after three attempts and were dropped. They span 88 pull requests and 26 default-branch pushes, across about 80 projects in roughly 60 workspaces. For each of these builds:
The Qlty check on GitHub showed the timeout error instead of a pass or fail, and the check linked to our troubleshooting docs rather than to the build.
No pull request comment or review was posted.
The build page on Qlty showed the build as passed, with no indication that its results were missing.
For the default-branch pushes, that commit is missing from the project's quality history until the next push.
A smaller group of builds that did complete on a retry recorded some issues more than once, which inflated their issue counts. This affects 42 builds. We are deciding whether to correct these or leave them, since a later push replaces them.
If you had a pull request open during this window and the Qlty check still shows the timeout error, push a new commit and the check will run again.
Root cause
A customer ran a routine cleanup of older data in their project. That is a normal, supported action. Behind it, Qlty removes that project's data from one of our largest tables. Because of how that table is laid out, removing one project's data required reading a large share of the whole table to find the small fraction that matched.
That read did not slow the database itself as much as it slowed ClickHouse Keeper, the coordination service that tracks what data each database replica has. Every request to Keeper became about six times slower for the 13 minutes the operation ran.
The final step of processing an analysis result, which computes a build's metrics, asks Keeper to confirm the replica has the latest data before the query starts. Under normal conditions that check takes under a second. With Keeper slowed, and about 40 of these steps waiting at once, each check queued behind the others and the wait grew to several minutes. Our 30-second client timeout cut each attempt short, the queue retried three times over about six minutes, and builds whose three attempts all fell inside the 13-minute window were dropped.
What we are doing
Making data removal a bounded operation. Removing a project's data will cost in proportion to the data being removed, not the size of the table it lives in, so one workspace's cleanup cannot slow the service for everyone.
Reducing the final processing step's dependence on Keeper. We are cutting the number of coordination checks each metrics query makes, and we are designing a version of that step that does not need them at all.
Alerting sooner. We are adding alerts on the database timeouts that preceded the queue delay, so we see an event like this within a couple of minutes instead of twelve.
If you have questions about a specific pull request or build from this window, contact us through qlty.sh/contact/support and include the pull request or build link.
No components marked as affected
Resolved
On October 5, 2026, between 12:18 and 12:31 PM ET, Qlty stopped completing analysis results for 13 minutes. About 120 builds across roughly 60 workspaces never completed. On the affected pull requests, the Qlty check on GitHub showed "Qlty encountered a timeout error at 15mins" while the build page on Qlty showed the build as passed. Pushing a new commit to an affected pull request produces a fresh, correct result.
We are sorry for the confusion this caused. Below is what happened, who was affected, why, and what we are changing.
Timeline (US Eastern on October 5, 2026, UTC in parentheses)
12:17 PM (16:17): A routine data cleanup in one workspace started a large maintenance operation on our main database.
12:18 PM (16:18): Analysis results stopped completing. Builds kept running and finishing normally, but their results were not written.
12:30 PM (16:30): Our queue-delay alarm fired.
12:31 PM (16:31): The operation finished and results started completing again.
12:33 to 12:50 PM (16:33 to 16:50): The Qlty check on affected pull requests changed from pending to the 15-minute timeout error.
12:35 PM (16:35): Result processing was back to normal.
About 1:00 PM (17:00): A customer reported a pull request stuck on the timeout error. We advised pushing a new commit.
4:00 PM (20:00): We confirmed the root cause.
Who was affected
During the 13-minute window, every analysis result was delayed by up to 12 minutes. Most builds recovered on their own: our queue retries each result three times, and 73 builds completed on a later attempt.
118 builds did not complete after three attempts and were dropped. They span 88 pull requests and 26 default-branch pushes, across about 80 projects in roughly 60 workspaces. For each of these builds:
The Qlty check on GitHub showed the timeout error instead of a pass or fail, and the check linked to our troubleshooting docs rather than to the build.
No pull request comment or review was posted.
The build page on Qlty showed the build as passed, with no indication that its results were missing.
For the default-branch pushes, that commit is missing from the project's quality history until the next push.
A smaller group of builds that did complete on a retry recorded some issues more than once, which inflated their issue counts. This affects 42 builds. We are deciding whether to correct these or leave them, since a later push replaces them.
If you had a pull request open during this window and the Qlty check still shows the timeout error, push a new commit and the check will run again.
Root cause
A customer ran a routine cleanup of older data in their project. That is a normal, supported action. Behind it, Qlty removes that project's data from one of our largest tables. Because of how that table is laid out, removing one project's data required reading a large share of the whole table to find the small fraction that matched.
That read did not slow the database itself as much as it slowed ClickHouse Keeper, the coordination service that tracks what data each database replica has. Every request to Keeper became about six times slower for the 13 minutes the operation ran.
The final step of processing an analysis result, which computes a build's metrics, asks Keeper to confirm the replica has the latest data before the query starts. Under normal conditions that check takes under a second. With Keeper slowed, and about 40 of these steps waiting at once, each check queued behind the others and the wait grew to several minutes. Our 30-second client timeout cut each attempt short, the queue retried three times over about six minutes, and builds whose three attempts all fell inside the 13-minute window were dropped.
What we are doing
Making data removal a bounded operation. Removing a project's data will cost in proportion to the data being removed, not the size of the table it lives in, so one workspace's cleanup cannot slow the service for everyone.
Reducing the final processing step's dependence on Keeper. We are cutting the number of coordination checks each metrics query makes, and we are designing a version of that step that does not need them at all.
Alerting sooner. We are adding alerts on the database timeouts that preceded the queue delay, so we see an event like this within a couple of minutes instead of twelve.
If you have questions about a specific pull request or build from this window, contact us through qlty.sh/contact/support and include the pull request or build link.