Blog / How to Test Feature Flags End-to-End Without Flaky Tests
How-To

How to Test Feature Flags End-to-End Without Flaky Tests

Feature flag tests flake because the test does not own the flag state. Pin the flag per test, cover the states that ship, and skip the combinatorial matrix.

Ten boolean feature flags put your app in 1,024 possible states. LaunchDarkly’s own testing guidance says covering every one would take 1,024 test cases (LaunchDarkly, 2022). Nobody writes those. The teams who try end up with a suite that goes red on a combination no human ever intended to ship.

Wiring flag-aware tests across dozens of flows taught us that the flake is almost never the flag. The test just does not own the flag’s value, so whatever the flag service returns the moment it runs becomes a silent, uncontrolled input. Fix that and flag tests stop being a coin flip. The sections below cover pinning the state, choosing which states to test, and keeping flags from becoming test debt.

What you’ll learn

  • Why a feature flag makes an end-to-end flow nondeterministic
  • Three ways to pin a flag to a known value inside the test
  • A decision rule for which flag states deserve a test, instead of the full matrix
  • How to keep stale flags from turning into silent test debt

The Flake Is Not the Flag

A feature flag adds a hidden input to every flow it touches. The same build renders a new checkout or the old one depending on a value your test never set, so an end-to-end test that ignores the flag is really testing whichever state the flag service returned that run. The nondeterminism does not come from your framework or your selectors. It comes from an uncontrolled input sitting in the middle of the flow.

The first time this bit us, the tell was that the failures had no pattern. The same checkout test failed maybe one run in five, always on a different assertion, never reproducible locally. We chased selectors and waits for two days before we noticed the staging flag environment was shared, and a percentage rollout on an unrelated flag was sending different runs down different branches. The test was not flaky. It was faithfully testing a different app each time.

The problem grows as teams ship more code. Teams gate more of it behind flags to ship it safely, and more flags means more hidden inputs, and more flows whose behavior depends on a switch nobody pinned. Testing AI-generated code is partly a feature-flag problem now, because flags are how much of that extra code ships.

Why Feature Flags Break End-to-End Tests

A flag breaks an end-to-end test for the same reasons it is useful in production. Every property that makes a flag good for progressive delivery is precisely what makes an untamed test flake:

  • The value is external. The flag lives in a service or a config store your test does not own. Because the flag controls live outside the test, the test can never assume which value it will get.
  • The value changes under you. Flags are meant to be flipped at runtime. A teammate toggles a rollout, or a percentage rule puts the test’s identity in the other bucket, and a test that passed an hour ago now hits the other branch.
  • The environment is shared. When several suites point at one flag configuration, a targeting change made for one team flips a flag mid-run for another. That shared-config flakiness is a predictable failure mode, not bad luck.

None of these are edge cases you can harden away at the UI layer, because they are the design of the mechanism. A smarter wait or a tougher selector will not fix them. The fix is taking the flag’s value out of the service’s hands for the duration of the test, the same pattern that recurs across the hardest-to-stabilize flaky tests: stop observing an external system you do not control, and set the input yourself.

Why Reading the Flag From the Dashboard Keeps Flaking

The instinct is to let the test read whatever the flag platform serves, the same way production does, and just assert on the result. Teams reach for the vendor’s dashboard, set the flag “on” for staging, and run. It flakes in the ways you would predict:

  1. A shared environment gets retargeted mid-run. Someone changes a rule for their own testing, your flag flips, and half your assertions invert with no code change on your side.
  2. A percentage rollout buckets whoever the test happens to be. Platforms commonly hash a stable identity, LaunchDarkly from the context key and kind and Unleash from the user ID or, failing that, the session ID, so the same identity gets the same answer on every run. The flake starts when the identity changes between runs, such as a freshly generated test user or a shared account, when someone changes the rollout percentage, or, on Unleash, when neither ID is set and it picks a random number each time. Then a 50% flag sends some runs down one branch and some down the other, with no change on your side.
  3. The service call itself can fail. A slow or unreachable flag lookup falls back to a default, so your “on” test quietly runs the “off” path and asserts against the wrong screen.

Unleash describes two common approaches, a “mock” approach and a “platform” approach (Unleash, 2023). Both beat reading a live dashboard value at runtime. The deeper fix is to decide the flag’s value in the test and make the app honor that decision.

The mental model

Your test is the source of truth for the flag, not the flag service. Set the value before the app boots, run the flow against that fixed value, and never let a runtime lookup decide which branch you are testing. Everything below is how to make the app honor the value your test picked.

Pin the Flag State at the Source

Every reliable flag test follows one rule. Fix the flag’s value before the flow starts, at a layer the flag service cannot override mid-run. There are three places to do it, ranked by how much of the real flag machinery they still exercise. Pick per test based on what you need to verify.

  1. Override at the request level so the running app reads your value instead of calling out. Most of the real app runs, the flag lookup is short-circuited.
  2. Bootstrap the SDK with fixed values so the client resolves flags from an in-memory set with no network at all. Fastest and most deterministic, furthest from the real platform.
  3. Target a dedicated test user on the real platform so the actual SDK and network path resolve the flag. Closest to production, slower, and only as stable as the environment you point at.

The next three sections wire up each one, then a table maps them to when each is the right call.

1. Override the Flag at the Request Level

The cleanest option for an end-to-end test is to let the app read the flag from a request-scoped override in non-production, such as a cookie, header, or query parameter that production ignores. Your test writes the override before it navigates, the app honors it for that session only, and the flag service is never consulted for that flow. Almost the entire app stays in the loop, with only the one nondeterministic call removed.

In Playwright, set the override cookie on the context before the first navigation so the value is present when the app boots:

// The app reads this override map only when NODE_ENV !== 'production'.
// Production ignores it, so a real user's cookie does nothing.
await context.addCookies([{
  name: 'ff-overrides',
  value: encodeURIComponent(JSON.stringify({ 'new-checkout': true })), // URL-encoded so strict cookie parsers accept it; the app decodes it
  domain: 'staging.example.com',
  path: '/',
}]);
await page.goto('/checkout'); // relative paths need baseURL in the Playwright config

The one requirement is a small, non-production hook in the app that decodes and merges this override on top of whatever the flag client would return. Build it once, gate it exactly like any test-only bypass, and it serves every flag test you write. It is the same shape of narrow, environment-scoped contract that keeps authentication and OTP flows deterministic: a door the test environment honors and production does not.

2. Bootstrap the SDK With Fixed Values

When you want zero network in the loop, initialize the flag SDK from a fixed, in-memory set of values instead of the live service. The client resolves every flag from data your test supplies, so there is nothing to time out, retarget, or roll the dice on. Vendor-neutral tooling makes this a first-class mode. OpenFeature, a CNCF incubating project that standardizes flag evaluation, ships an in-memory provider built for exactly this:

import { OpenFeature, TypedInMemoryProvider } from '@openfeature/server-sdk';

// Deterministic flags for the test run. No flag service, no network.
await OpenFeature.setProviderAndWait(new TypedInMemoryProvider({
  'new-checkout': { variants: { on: true, off: false }, defaultVariant: 'on', disabled: false },
}));

Other vendor SDKs expose the same idea under a different name, such as LaunchDarkly’s TestData source or Unleash’s bootstrap option. Turn off network fetching where the SDK lets you, so the fixed values are the only source. The trade-off is that you test every code path the flag gates, but not the flag service’s targeting rules or its network behavior. Use this for the bulk of your flow coverage, where the question is whether the “on” checkout and the “off” checkout both work.

Skip the Flag Plumbing

Pin the state, hand the flow to Pie, and run each variant without scripting selectors.

Book a Demo

3. Target a Test User on the Real Platform

Sometimes the platform itself needs verifying: does this targeting rule serve the new experience to the right users, and does the app fall back correctly when the service is unreachable? A feature flag integration test answers that, and it wants the real SDK and network path instead of an in-memory stand-in. Create a dedicated test user or a dedicated test environment on your flag platform, set the flag deterministically for that identity, and drive the flow as that user.

The discipline that keeps this from flaking is isolation. Give the test its own targeting rule keyed to an identity nothing else uses, so no other suite or teammate can retarget it out from under you. Never run these against a percentage rollout, and never share one flag environment across parallel workers.

This path is slower and only as reliable as the environment you point at, so reserve it for the handful of tests that need to prove the platform resolves correctly, and lean on overrides and bootstrapped SDKs for everything else. It fits the integration tier of a continuous testing pipeline, where a broken targeting rule should fail fast before the flow tests run.

Which Approach to Use

Match the technique to what the test needs to prove and how much of the real flag machinery you want in the loop. Use overrides or a bootstrapped SDK for the bulk of your coverage, plus a few platform-targeting tests to keep the flag service honest.

ApproachBest forReal flag path exercisedSpeed & determinismSetup cost
Request-level overrideMost end-to-end flow testsApp runs, flag lookup skippedFast, deterministicLow (one non-prod hook)
SDK bootstrapUnit and integration, both branchesNone, in-memory valuesFastest, deterministicLow (SDK test mode)
Platform targetingProving targeting and fallbackFull SDK plus networkSlower, environment-boundMedium (test user or env)

Which Flag Combinations Are Worth Testing

Pinning the state solves reliability. It does not tell you which states to bother with, and this is where most guidance stops short. The combinatorial math is a trap: those 1,024 states for ten flags are almost all combinations that never ship together and carry no user value. LaunchDarkly’s own advice is that you should never test all combinations of all flags. The useful question is not how many states exist, but which states reach real users. Four do:

  1. The current production configuration. Every flag at its live value. This is the state your users are actually in right now, and it is the one most suites forget to assert because it feels like the default.
  2. The next state you are rolling out. The flag under test flipped to its target value, with everything else at production. This is the change you are about to ship, so it is the state the release depends on.
  3. The off or fallback state. What users get when the flag is off or the flag service is unreachable and the client returns its default. A missing fallback test is how a flag-service outage quietly becomes a customer-facing one.
  4. Known interactions. The small set of flags that touch the same flow, such as a new checkout flag and a new payment-provider flag that both alter the purchase path. Test those together, because their interaction is where the real bugs hide.

That makes four states per flagged flow instead of a thousand. The rule scales because it is anchored to what ships, so the size of the flag set stops mattering. A new flag adds a handful of meaningful states to a flow, with no exponential blowup. Getting coverage right on the states that matter is the same discipline as building any regression suite that stays fast: test what can break in a way users feel, and skip the rest on purpose.

Should You Remove Feature Flags?

Yes, and the cleanup is a testing concern as much as housekeeping. A flag that has fully rolled out and stabilized is dead weight the moment its job is done. Pete Hodgson’s feature-toggle guidance on martinfowler.com treats flags as inventory that carries a cost. Every abandoned flag is also one more hidden input in a flow you assumed was simple, plus an untested off-path that quietly rots. Stale flags are how a codebase accumulates states nobody remembers, which is precisely the combinatorial mess the previous section works to avoid.

The habit that keeps it under control is cheap: open the removal ticket in the same pull request that introduces the flag, and give short-lived release flags an expiry the way you would a TODO. When a flag is retired, delete the flag, the dead branch, and the tests that pinned it. The point of a flag is to be temporary. Treating it as permanent is how the test surface silently doubles while everyone swears the app got no more complex.

How Pie Tests Flows Behind Feature Flags

Pinning the flag is the easy half. The hard half is everything the flag changes downstream. A release flag does not just gate one line. It swaps a component, reflows a page, moves a button, renames a field. Every one of those is a selector your scripted test wrote against the old layout, and it breaks the sprint the flag flips. Selector churn like that is what rots a suite.

Pie, an autonomous QA platform, tests web, iOS, and Android apps by looking at the screen, as a user does. You pin the flag state as you would for any test, through a test account or tenant, a control your app supports, or a configured script that calls a flag endpoint you control. Then you point Pie at the flow on a staging URL with a test login. Two things follow:

  • The old layout is not baked into the test. Pie identifies elements visually rather than by a hard-coded locator, so a flag that reshapes the UI can leave the test as written. If the flag changes what the flow should do, update the expected result.
  • Each variant is a run against a pinned state. Covering the “on” checkout and the “off” checkout is two runs against two pinned states, each with its own expected result, not two hand-maintained scripts. Pie does not infer which flag sits behind a screen, so pin the state and say what you expect.

Keeping a script alive through every variant a flag introduces is the expensive part of flag testing, and Pie takes it off your plate. You decide which states matter, and Pie runs the flows inside them without anyone on selector duty. For the wider picture of how this fits a modern suite, the end-to-end testing guide covers where flag tests sit alongside the rest of your coverage.

Pin the Flag, Then Test the Flow

Feature flag tests flake because the flag’s value is an input your test does not own, sitting in a service that can change it mid-run. Take that input back. Override the flag at the request level, bootstrap the SDK with fixed values, or target a dedicated test user, and the flow becomes reproducible.

Then test the states that ship (current, next, off, and the flags that interact) and skip the thousand-state matrix. Retire flags the moment their job is done so the surface stops growing behind your back. Pie can adapt to the UI churn each flag introduces, so the flow can stay under test even after the flag reshapes the screen it runs on.

Stop Rewriting Tests Every Time a Flag Flips

Pin the state, point Pie at the flow, and it can adapt when the UI changes.

Book a Demo

Frequently Asked Questions

A feature flag is a runtime switch that toggles a code path without a deploy, so one build can behave two ways.

In testing, that adds a hidden input to every flow the flag touches.

A test that does not control the flag's value tests whichever state the flag service returned that run.

Pin the flag to a known value before the app boots. Use a request-level override the app honors only in non-production, an SDK bootstrapped with fixed values, or a dedicated test user on the real platform.

Whichever you pick, the test sets the flag state and the flag service never surprises it.

A feature flag integration test runs the real SDK and network call against a test environment instead of mocking the flag. It asks whether the flag resolves to the value you set, targeting rules and fallback included.

Use it to catch broken targeting rules, and end-to-end tests to check each resulting flow.

No. Ten boolean flags produce 1,024 combinations, and LaunchDarkly's guidance says covering them all takes 1,024 test cases (LaunchDarkly, 2022). Almost none of those states ever ship together.

Test the states real users reach, such as production today, the state you are rolling out, the fallback state, and any flags that interact inside one flow.

Yes. Once a flag is fully rolled out and stable, delete it and its dead branch.

Pete Hodgson's guidance on martinfowler.com treats flags as inventory with a carrying cost. Every abandoned flag also leaves an untested off-path and one more hidden input. Put a removal ticket in the same PR that introduces the flag.

Stop letting the shared flag environment decide your test's inputs. When several suites share one flag configuration, a targeting change for one team flips a flag under another team's tests mid-run.

Pin the flag inside the test with an override or bootstrapped SDK, and never read the live dashboard value at runtime.

Yes, once the flag state is pinned through a test account or tenant, a control your app supports, or a configured script. Then point Pie at the flow on a staging URL with a test login.

Each variant is a run against a pinned state with its own expected result, not a separately scripted test.

Adithya Aggarwal
Adithya Aggarwal
CTO & Co-founder at Pie

Eight years building search and delivery systems at Amazon. The kind of scale where flaky tests block billion-dollar releases. Now CTO at Pie, building AI agents that adapt when your UI changes. LinkedIn →