A retry is a clue: investigate flaky Playwright tests
Prepared with AI assistance and linked primary sources. Examples are illustrative unless stated otherwise.
A test that fails and then passes on retry still needs investigation. Preserve the first failure, check timing and shared-state assumptions, and wait for the behaviour the requirement actually promises. Validate a targeted correction with repeated runs and a deliberate failure case; adding retries alone does not explain the defect.
Begin with a controlled timing problem
This October 9 evergreen lesson examines flaky-test investigation using established Playwright behaviour. It is not a release announcement. Use a disposable local page with a Save button and a status element. Clicking Save should eventually change that status to Saved after a simulated delay; no real backend or customer data is needed.
Define two practice delays, 100 milliseconds and 1,200 milliseconds, and select them deliberately instead of using an unrecorded random value. Write the requirement: after a successful save action, the page shows Saved within an agreed limit. Record the chosen limit before testing. The delay variation is part of the experiment, not evidence that a real product has the same performance.
Preserve the first failed attempt
Playwright distinguishes a first-attempt pass from a test that fails initially and passes on retry. Its retry documentation calls the latter flaky. Keep the first failure's assertion message and relevant evidence, then compare the repeated attempt. A passing retry is useful information about variability, but it does not identify the original cause.
Record the app build, test version, selected delay and starting state for every run. If the first attempt leaves data behind, a later attempt may encounter a different situation. For this lab, reset the page before each run and keep the fixture local. In a larger suite, investigate shared accounts, records and server state separately.
Recognise a one-time read of changing state
An assertion such as expect(await page.getByTestId('save-status').textContent()).toBe('Saved') reads text once and compares that returned value. If the status is still Saving at that moment, the assertion fails even when the simulated operation completes shortly afterwards. Save the actual and expected values to make that timing assumption visible.
Do not respond by removing the assertion or accepting either Saving or Saved. That would weaken the requirement. Also avoid choosing a long fixed sleep merely because one machine happens to pass afterwards. First decide which condition represents completion and whether the application provides a stable way for the test to observe it.
Wait for the promised condition
Use await expect(page.getByTestId('save-status')).toHaveText('Saved', { timeout: 3000 }) for this practice requirement. Playwright's web assertions repeatedly check the matching element until the condition is satisfied or the timeout expires. The example's three-second limit is a deliberate lab choice, not a universal recommendation for production response times.
Run the same test with the 100-millisecond and 1,200-millisecond delays. Both should complete when the status becomes Saved within the chosen limit. Record the outcomes without claiming every timing issue is solved. A selector that identifies the wrong control or a save operation that never succeeds requires a different correction.
Prove that the check still detects a problem
Add a controlled mode that leaves the status at Saving forever. The condition-based assertion should fail after its timeout, demonstrating that the revised test has not become unconditional. Keep this intentional failure separate from the normal passing run. Name the mode clearly so a reviewer can understand why its expected result is a failure.
Then repeat the normal cases several times and report the exact number of attempts and observed outcomes. If any fail, investigate their evidence before declaring success. Passing ten repetitions supports those ten observations; it does not prove a zero failure probability or cover different browsers, network conditions and shared environments you did not exercise.
Write a diagnosis someone can review
Summarise the original timing assumption, the observed failure, the changed assertion and the passing and intentionally failing cases. State that the delay was simulated. Explain why waiting for Saved matches the requirement more directly than waiting an arbitrary number of milliseconds. This makes the reasoning reviewable without requiring someone to trust your final green screenshot.
In a real flaky test, timing may be only one candidate among data collisions, environment differences, unstable selectors and genuine application defects. Keep those possibilities open until evidence narrows them. Continue through the related Playwright + GenAI learning path by applying the same investigation record to one authorized practice test and preserving the first failure.
Sources and further reading
- Playwright documentation: Retries · checked 2026-10-09
- Playwright documentation: Assertions · checked 2026-10-09
Spot an error? Email info@spothub.in with the article link and correction.
