A flaky test passes and fails on the same code. Here is how to find the real cause, contain the damage and stop it coming back.
A pull request fails CI. The failing test has nothing to do with the change. Someone clicks "Re-run", the build goes green and the change merges. Two weeks later it happens again, and this time nobody opens the failure log at all.
That is the real cost of flaky tests. The rerun costs a few minutes. The loss of trust takes months to rebuild: a red build stops meaning "something is broken" and starts meaning "try again". Once that happens, the next real failure gets the same treatment.
The problem is common even at the biggest engineering organisations. In a 2016 post, Google's testing team reported that about 1.5% of all test runs gave a flaky result and almost 16% of their tests had some level of flakiness. That is one very large codebase, so treat the numbers as an illustration rather than a benchmark. The pattern matters more than the figures: flakiness is not rare, and it does not fix itself.
This guide covers what counts as a flaky test, seven root causes, a six-step method for diagnosing one, a quarantine policy that does not turn into a graveyard, fixes by cause, and the metrics worth tracking. It is about intermittent failures from timing, data, environment and concurrency. If your tests break because a locator changed after a UI update, that is a different problem, and we cover it in our guide to self-healing test automation.
What Is a Flaky Test?
A flaky test is an automated test that gives different results, pass or fail, when it runs repeatedly against the same code, the same configuration and the same data.
It helps to separate three situations that all look like "the test failed":
- A consistently failing test. It fails every time. Either there is a real bug or the test is out of date. This is not flaky.
- A flaky test caused by the test or its environment. The product is fine. The test depends on timing, leftover data, a slow runner or an external service.
- A flaky test caused by the product. The test is correct and has caught a real race condition that users can also hit. This is the dangerous one, because "it's just flaky" is exactly the wrong conclusion.
Most of the work in handling flakiness is deciding which of the three you have.
Why Flaky Tests Cost More Than a Rerun
- Trust erodes. Engineers who learn that red often means nothing stop investigating red.
- Real bugs hide. An intermittent product defect looks identical to an intermittent test. Without investigation, it ships.
- Delivery slows. A suite that takes 40 minutes costs 40 minutes every time someone reruns it, plus the queue for the runner. Multiply that by every developer and every day.
- Release decisions get noisy. A QA lead who cannot tell whether a red run is a blocker ends up making the call on instinct.
The Seven Root Causes of Flaky Tests

| Root cause | What it looks like | Typical fix |
| Timing and asynchronous waits | Passes locally, fails on a slower CI runner. "Element not found" or stale-element errors. The assertion runs before the page or API has finished. | Wait for a condition, not for a fixed number of seconds. |
| Shared or leftover test data | Fails only on a second run, or after another test ran. "Record already exists". | Each test creates and cleans up its own data. |
| Order dependence and shared state | Passes alone, fails inside the suite, or fails only when tests run in parallel. | Isolate state per test and run in random order. |
| Unstable environments | Failures cluster on certain runners, at certain times, or after a deployment. Timeouts and connection resets. | Ephemeral, version-pinned environments with resource limits. |
| External dependencies | Third-party API rate limits, slow email or SMS delivery, a payment sandbox that is down. | Stub third parties in most tests and test the real integration separately. |
| Time, locale and randomness | Fails near midnight, at month end, on daylight-saving change days, or in another time zone. Results in a different order. | Inject a clock, seed random values, sort before comparing. |
| Real concurrency bugs in the product | Intermittent failures that continue with clean data and a stable environment. Duplicate orders, lost updates. | Fix the product, not the test. |
A few of these deserve more detail.
Timing (cause 1) is the most common source of flakiness in UI and API tests. A fixed sleep is a guess about how long something takes. On a developer's laptop the guess is right. On a busy CI runner it is wrong one run in twenty.
Shared data and state (causes 2 and 3) are hard to spot because the test is correct in isolation. They show up when you add parallel execution, which is why a suite that was stable at 50 tests becomes flaky at 500. If your data setup is a problem, our guide to test data management covers synthetic data and masking.
Time and locale (cause 6) is the one teams forget. Tests that assume a 24-hour day, a single time zone or a fixed month length pass for most of the year and fail on specific dates. Our timezone testing service exists because these failures are predictable and still catch teams out.
Real product bugs (cause 7) are why you should never skip the diagnosis step. A test that fails one run in fifty on "duplicate order created" is not noise.

How to Diagnose a Flaky Test: Six Steps
1. Confirm it is flaky and measure the rate
Run the single test many times against the same commit. Do not trust a single rerun. The goal is a number such as "failed 4 of 30 runs", which tells you how hard the problem is and gives you a baseline to prove a fix.
# Playwright: run one test 30 times
npx playwright test tests/checkout.spec.ts --repeat-each=30
2. Capture evidence on every failure
A failure you cannot inspect cannot be fixed. Collect the screenshot, the console and network logs, the application logs for the same time window, and the runner details such as machine, parallel worker and start time. Most frameworks can record a trace automatically on failure.
// playwright.config.tsimport { defineConfig } from '@playwright/test'; export default defineConfig({retries: process.env.CI ? 2 : 0, // limited retries in CI onlyuse: { trace: 'on-first-retry' }, // keep a trace when a retry happens});
Playwright reports a test that fails and then passes on retry as "flaky", which gives you a list to work from.
3. Group failures by signature
Collect the failure messages and the step where each one failed. Failures with the same message at the same step usually share one cause. Failures with different messages in the same test often mean more than one problem is hiding there.
4. Look for correlations
Compare failing and passing runs. Ask whether the failures cluster on a specific runner, time of day, parallel worker count, branch, or on the test that ran just before. A strong correlation usually points to the cause category before you read any code.
5. Reproduce it under pressure
Make the conditions worse on purpose: throttle the CPU, run with more parallel workers, run the suite in random order, repeat with the same random seed. If the failure appears only under load, suspect timing or environment. If it appears only in a certain order, suspect shared state.
6. Classify it: test, environment or product
Decide which of the three situations from earlier you have. This sets the owner. A test problem goes to the test author, an environment problem to the platform team, and a product problem goes to the developers as a defect, not as a "flaky test" ticket.

A Quarantine Strategy That Does Not Become a Graveyard
Quarantine means moving a test out of the blocking path of the pipeline while keeping it running and visible. Google describes the same idea: its tooling monitors the flakiness of tests and, when it is too high, quarantines the test and files a bug.
Quarantine becomes a graveyard when tests go in and never come out. These rules prevent that.
- 1Define "flaky" in numbers. For example, a test that failed and later passed on the same commit at least twice in its last 20 runs on the main branch. These are example starting points, so tune them to your suite.
- 2Quarantine, do not delete or silently skip. The test keeps running in a non-blocking job and keeps reporting.
- 3Give every quarantined test an owner, a ticket and a deadline. For example, 14 days.
- 4Protect critical paths. If the test guards payments, login or data deletion, add a compensating check, manual or automated, until it is fixed.
- 5Write exit criteria. For example, a test returns to the blocking suite after 50 consecutive green runs in the non-blocking job.
- 6Make the quarantine visible. Show the count, the age of each test and the owner on a dashboard, and review it weekly. Agree a ceiling, for example 5% of the suite, above which fixing takes priority over new tests.
Retries, quarantine, deletion or a fix?
| Option | When it helps | The risk |
| Retry (once or twice in CI) | A single network blip should not block a merge. | Retries hide flakiness. Track every "passed on retry" as a flake signal. |
| Quarantine | The test is flaky and the cause will take time to find. | Becomes permanent without owners and deadlines. |
| Delete | The behaviour is covered elsewhere or no longer exists. | Losing coverage nobody noticed. Check first. |
| Fix now | The cause is known and small. | None, as long as you re-run the 30-run check from step 1 to prove it. |
Fixes That Last, by Cause
Timing. Replace fixed sleeps with waits on a condition: an element visible, a network call complete, a status changed.
// Flaky: guesses how long the page needs
await page.waitForTimeout(3000);
await expect(page.getByText('Order confirmed')).toBeVisible();
// Stable: waits for the condition, up to a limit
await expect(page.getByText('Order confirmed')).toBeVisible({ timeout: 10_000 });
Data. Create the data a test needs at the start of that test, through an API where possible, with a unique identifier such as a UUID. Remove it afterwards, or run the test in a transaction that is rolled back.
Shared state. Give each test a fresh browser context or a clean session. Do not let tests read what other tests wrote. Run in random order in CI so hidden dependencies show up (for example, the pytest-randomly plugin for Python).
Environment. Use ephemeral environments created per run, pin the versions of browsers, drivers and base images, set CPU and memory limits, and add a health check that must pass before the first test starts.
External services. Stub the third party in most tests. Keep a small, separate set of tests against the real service, and run them on a schedule instead of on every commit.
Time and randomness. Inject a controllable clock, set random seeds, and sort collections before comparing them. Add tests for the days that break assumptions: month end, leap day, daylight-saving change.
Product concurrency. Reproduce it with the failing run's evidence, write it up as a defect, and fix the code. Keep the test.
Keep New Flakiness Out
- Run new or changed tests many times in the pull request before they merge, for example 20 runs.
- Run tests in random order in CI.
- Add a lint rule that rejects fixed sleeps.
- Favour lower-level tests. API and unit tests flake less than long browser journeys, so keep end-to-end tests for the journeys that matter. Our guide to building a test automation framework from scratch covers where each layer belongs.
- Treat test code as production code, with review and an owner. See how to write maintainable test scripts.
- Pin tool and image versions and upgrade them deliberately.
Metrics Worth Tracking
| Metric | Why it matters |
| Flake rate per test | Shows which tests do the most damage, and proves a fix worked. |
| Pass-after-retry rate | Retries hide flakiness. This number shows how much is hidden. |
| Quarantine size and age | Shows whether quarantine is working or filling up. |
| Time from first flake to quarantine or fix | Shows how fast the team reacts. |
| CI minutes lost to reruns | Turns a vague annoyance into a cost the business can see. |
Use these to improve the system. Do not use them to rank individual engineers, because that encourages people to hide flakes instead of reporting them.
Flaky Test Triage Checklist
Copy this into your team wiki.
- Ran the test at least 30 times on the same commit and recorded the failure rate
- Captured a trace, screenshot, logs and runner details for each failure
- Grouped failures by error message and step
- Checked runner, time, parallelism and preceding test for a pattern
- Tried to reproduce under CPU throttling, parallel load and random order
- Classified the cause as test, environment or product
- If product: filed a defect and kept the test
- If quarantined: assigned an owner, ticket and deadline
- If critical path: added a compensating check
- After the fix: re-ran the 30-run check and recorded the new rate
Where Flaky-Test Work Fits in a QA Strategy
Stable tests are the base that regression suites, CI/CD gates and release decisions sit on. A suite nobody trusts does not protect a release, no matter how many tests it has. Fixing flakiness raises the value of the automation you already own.
Testriq builds automation in Selenium, Cypress and Playwright as part of its automation testing services and regression testing. For pipeline integration, see continuous testing in CI/CD.
Frequently Asked Questions
What is a flaky test?
A flaky test is an automated test that sometimes passes and sometimes fails on the same code, configuration and data. The cause is usually timing, shared data, an unstable environment, an external service, or a real concurrency bug in the product.
Are automatic retries a good fix for flaky tests?
Retries are a containment tool, not a fix. One or two retries in CI stop a single network blip from blocking a merge. But retries also hide flakiness, so record every test that passes only on retry and treat that list as work to do.
How many times should I rerun a test to confirm it is flaky?
There is no universal number. Running the test 20 to 50 times on the same commit is a practical range. The aim is to measure a failure rate, such as 3 failures in 30 runs, so that you can show the rate dropping after a fix.
Should we delete flaky tests?
Only after checking that the behaviour is covered elsewhere or no longer exists. Deleting a flaky test that guards a real risk removes the warning without removing the risk. Quarantine it with an owner and deadline first.
Can a flaky test be a real bug?
Yes. If a correct test fails intermittently with clean data and a stable environment, the product may have a race condition. Classify every flake as a test, environment or product problem before you decide what to do.
Does switching to Playwright or Cypress remove flakiness?
Modern tools that wait automatically for elements remove one common cause, fixed-sleep timing. They do not remove shared data, unstable environments, external dependencies or product bugs, so you still need the diagnosis and quarantine process.
Conclusion
Flaky tests are a symptom with several different causes, and each cause has a different fix. Measure the failure rate first, capture evidence on every failure, classify the cause as test, environment or product, and quarantine with an owner and a deadline while you fix it. Add guardrails so new tests are proven stable before they merge.
If your pipeline is full of reruns and your team has stopped trusting red builds, start with the failure history of your CI runs and the six-step method above. If you want a second pair of eyes on your test automation, talk to a QA specialist.
