Blog / How to Test OAuth and SSO Login Flows End-to-End in CI
How-To

How to Test OAuth and SSO Login Flows End-to-End in CI

Mocks and saved sessions skip the Google or Okta handoff, where logins break in production. Test the real flow in CI and know what to do when it is blocked.

The tests were green. The login broke anyway.

That happened often enough, across enough apps we wired up “Sign in with Google” and enterprise SSO for, that we stopped calling it a fluke and started calling it the tell. An OAuth suite can pass for months and still miss the exact failure that takes down real sign-ins, and once we traced why, the reason was the same every time: nothing in the suite ever touched the real handoff to the provider.

Most OAuth testing guides tell you to route around the real login. Mock the provider, reuse a saved session, or inject a signed token. All three work, and all three skip the handoff your users take. The fix is one real login test, run reliably, plus mocks for everything else.

What you’ll learn

  • Why the standard advice tells you to bypass the real login
  • Exactly which failures mocking the provider hides
  • How to run the real OAuth flow in CI reliably
  • What the common “Access blocked” errors actually mean

Why Most Guides Tell You to Fake the Login

The advice to fake the login is correct, and it comes from a real constraint. Public identity providers can block automated sign-ins.

Point Selenium, Cypress, or Playwright at a live Google login page and you are driving a browser that announces itself as automation. The provider treats it as what it looks like, a bot trying to get into someone’s account. Google’s own OAuth policies bar sending sign-in requests to embedded user agents, and its sign-in help says Google might stop sign-ins from browsers controlled by software automation, so the scripted approach can fail intermittently.

Because of that wall, the ecosystem converged on three ways around the real login, and you should use all three:

  1. Reuse a captured session. Log in once, save the cookies and tokens, and inject them into later runs. This is what Playwright’s storageState and Cypress’s cy.session() are built for.
  2. Mock the identity provider. Intercept the redirect and hand your app’s callback a pre-signed authorization code or token, so you test your own route handling without touching Google at all.
  3. Use a dedicated test account on a controlled or staging authorization server, ideally with multi-factor authentication turned off, so the sign-in has fewer moving parts.

None of this is wrong. The trouble starts when it becomes the only strategy. A suite built entirely on mocks and injected sessions never exercises the redirect to the provider and back, and that handoff is precisely where OAuth logins break.

What Mocking the Provider Actually Hides

Mocking the provider tests your callback handler. It does not test the flow your users take. The moment you replace the real redirect with a fake token, you stop covering the entire round-trip between your app and the identity provider, and that round-trip has its own failure modes that never show up in a green mocked suite.

Four failures live in the part you mocked away:

  • Redirect URI drift. The authorization-code flow requires the callback URL to match a registered redirect URI exactly. Change a domain, add a path, or promote to a new environment, and the provider rejects the authorization request with redirect_uri_mismatch. A mock never sends you back to a real redirect URI, so it cannot catch this.
  • Consent-screen and account-picker changes. Real users pass through an account chooser and, on first authorization, a consent screen. Providers redesign these, add scopes-review steps, and insert interstitials. A mock renders none of it.
  • Provider policy walls. An external app in “testing” status that asks for more than basic sign-in scopes only admits accounts on its test-user list, and an unverified app requesting sensitive scopes gets an “unverified app” interstitial or an outright block. These are policy states on the provider side, invisible to a mock.
  • The state parameter round-trip. OAuth’s CSRF protection depends on your app generating a state value and verifying the identical value on the callback. Mocking both ends of that exchange means you are validating your own fake against your own fake.

The ranking guides leave a gap here. Autonoma’s OAuth testing guide gets closest, making the case for a layered approach that keeps real authorization-code round-trip tests for at least the critical happy path, and it is right. But many teams read the headline as “mock it,” wire up the mock, and never build the one real test that would catch a redirect_uri_mismatch before a customer does.

Test the Real Flow Once, Reuse the Session Everywhere

The whole strategy fits in one rule. Drive the real provider login exactly once, in a single dedicated test, and give every other test its session for free. The same discipline makes end-to-end auth testing survive scale, and it is the move we recommend for 2FA and OTP flows: prove the human path works one time, then stop re-running it.

The mental model

One test drives the real OAuth handoff end to end, catching redirect drift, consent changes, and policy walls. Every other test starts from a saved session and never touches the provider. The first is your early-warning system. The rest are fast, deterministic, and unblockable.

The reason this matters is not just speed. Driving a live provider hundreds of times per suite can trip rate limits and suspicious-activity checks, and those read as flakiness in your dashboard. Re-run the real login once, cache the result, and you remove both the flake and the block in a single decision.

How to Run the Real OAuth Login in CI

Running the real login reliably comes down to reducing the reasons a provider flags you: an account it does not recognize, a second factor it cannot verify, and, on Google, a browser it may treat as automated. Google may stop automated sign-ins, so keep this to one test and have a fallback. Handle those, drive the flow once, and save the session. Here is the sequence we use.

  1. Provision a dedicated test account on a test tenant. Never automate a real user’s Google or Okta account. Create an account that exists only for tests, and where you control the provider (Okta, Auth0, a staging Keycloak), point tests at a non-production tenant so you are not rate-limited against your live directory.
  2. Register the account and the callback. For Google, add the test account under the OAuth consent screen’s test users while the app is in testing status and asks for more than basic sign-in scopes, and register every callback (including http://localhost ports) as an authorized redirect URI. Exact-match matters, so a wrong port or trailing slash is the usual failure.
  3. Turn off the second factor for that one account, or generate the code in the test. A dedicated test account with MFA disabled is the simplest path. If you must keep MFA on, compute the TOTP yourself using the technique in the 2FA guide.
  4. Drive the flow in a real browser context, not an embedded one. Google explicitly disallows embedded user agents. Use a standard, non-embedded browser, click through the account picker and consent, and let your app complete the exchange. Google may stop sign-ins from automated browsers, so have a fallback.
  5. Capture the session and reuse it. Once you land back on your app authenticated, save storageState and hand it to every other test.

The capture step is the payoff, and it is a few lines. In Playwright, a setup project runs the real login once and writes the session to disk:

// auth.setup.js: runs the real provider login a single time.
import { test as setup } from '@playwright/test';

setup('authenticate via real OAuth', async ({ page }) => {
  await page.goto('/login');
  await page.getByRole('button', { name: 'Sign in with Google' }).click();

  // Dedicated test account, MFA off, added as a Google test user.
  // Google changes these labels, so check them against the live page.
  await page.getByLabel('Email or phone').fill(process.env.TEST_GOOGLE_EMAIL);
  await page.getByRole('button', { name: 'Next' }).click();
  await page.getByLabel('Enter your password').fill(process.env.TEST_GOOGLE_PASSWORD);
  await page.getByRole('button', { name: 'Next' }).click();

  // A first sign-in may add a consent screen here. Click through it, then land back on the app and persist the session.
  await page.waitForURL('**/dashboard');
  await page.context().storageState({ path: 'playwright/.auth/user.json' });
});

Every other test declares storageState: 'playwright/.auth/user.json' and starts signed in, having never re-touched the provider. The saved file holds live session cookies, so add playwright/.auth to .gitignore. The setup test is where redirect drift, a changed consent screen, or a policy block surfaces first, on your schedule instead of a customer’s.

Skip the OAuth Plumbing

Hand your login flows to Pie and spend less time maintaining the auth setup.

Book a Demo

When to Mock and When to Drive the Real Flow

The layered strategy is not mock-or-real. It is mock for volume, real for coverage of the handoff. Choose the technique by what each test needs to prove, and keep at least one test on the real path so the mocked majority never drifts away from reality.

ApproachBest forCatches redirect/consent/policy bugs?Speed & reliabilityBlocking risk
Mock the providerUnit tests, most CI runsNo, skips the round-tripFastest, fully deterministicNone
Reuse saved sessionThe bulk of authenticated E2E testsNo, login already happenedFast, deterministicLow; refresh the saved session when it expires
Real login, test tenantOne critical happy-path testYes, the whole pointSlower, needs careLower with a test account; Google may block automated sign-ins
Real login, public providerFinal pre-release smoke checkYes, including provider policySlowest, most fragileHighest, run it rarely

The shape most teams land on: mocks and a reused session for the hundreds of tests that only need an authenticated user, one real-login test against a test tenant on every CI run, and an occasional real login against the public provider before a release to catch policy-level surprises.

The Errors You Will Hit and What They Mean

When the real login does get blocked, the error is usually telling you something specific. The errors below are the ones we hit most, straight from the practitioner threads (the Supabase Cypress discussion from 2022 is a good example of a developer stuck on exactly this):

  • “Access blocked” with a note that the app has not completed Google verification or is still being tested. Your OAuth app is in testing status with more than basic sign-in scopes and the account is not on the test-user list, or it is requesting sensitive scopes as an unverified app. Add the account under the consent screen’s test users, or verify the app.
  • “Access blocked: authorization error” with disallowed_useragent. You are signing in from an embedded webview. Switch to a standard, non-embedded browser, which is the documented requirement. Google may also stop sign-ins from browsers it detects as automated.
  • redirect_uri_mismatch. The redirect URI in the request does not exactly match a registered redirect URI. Check the scheme, host, port, and path, including whether localhost and your CI hostname are both registered.
  • Intermittent CAPTCHA or “verify it’s you” prompts. The provider flagged suspicious activity, which can happen when the suite drives the login too many times. Logging in once and reusing the session is the strongest fix here.

None of these are framework bugs. They are the provider doing its job, and every one of them is invisible to a mocked test, which is why they surprise teams in staging or production instead of in CI.

How Pie Runs the Real Login Flow

The setup above works, and it remains a standing maintenance cost. The account picker gets a new layout, the consent screen adds a step, the “Sign in with Google” button moves, and the selectors in your setup test break the sprint after you write them. Scripted suites are worst at exactly this kind of churn.

Pie is autonomous QA for web and mobile apps. For a login flow, two parts of that matter most:

  • It navigates the real handoff visually. Pie identifies the account picker, the password field, and the consent button by what they look like, not by a hard-coded selector, so it can adapt when a provider redesigns the screen. A real browser can still be challenged or blocked by the provider.
  • It works from a setup you control. You give Pie a staging URL and a dedicated test login. That means a usable test tenant or account, the federation redirects registered, automation the provider permits, and either an MFA policy the test can satisfy or an approved session setup. Pie does not provision the identity provider for you.

You describe the login flow once, and the plumbing underneath it takes less of your time. The redirect, the consent screen, and the account picker that moved overnight stop being your selectors to fix, and the one test that exercises the OAuth handoff can adapt when the provider changes the screen.

Test the Handoff, Not Just the Callback

Mocking OAuth is the right call for most of your suite, and reusing a session is how you keep it fast. The mistake is stopping there. If nothing in your suite ever drives the real redirect to Google or Okta and back, the redirect URI, the consent screen, and the provider’s own policy walls are untested, and those are where OAuth logins break.

Run the real flow once against a test tenant, in a real browser, with a dedicated account, and reuse the session everywhere else. One test like that is cheap insurance against the failures no mock can see. And when maintaining even that one flow becomes its own tax, let an agent that reads the screen like a user carry it.

Stop Babysitting Your Login Tests

Give your OAuth and SSO flows to Pie. It can adapt when the provider changes the screen.

Book a Demo

Frequently Asked Questions

Yes, but exactly once. Keep one end-to-end test that drives the full authorization-code flow through the real provider against a test tenant, so you catch redirect drift, consent-screen changes, and provider policy walls.

Every other test should reuse the session that login captured.

Two common causes. Your app is in testing mode, asks for more than basic sign-in scopes, and the account is not on its test-user list. Or Google stopped the sign-in because the browser is embedded or automated.

Add the test account to the consent screen's test users, and use a real, non-embedded browser.

Register your local callback (your localhost address, port, and callback path) as an authorized redirect URI in the Google Cloud console. While the app is in testing status, add your test account under the consent screen's test users.

Redirect URIs must match exactly, so a trailing slash or wrong port is the usual culprit.

Mock it for the bulk of your suite. Mocking is fast, deterministic, and correct for unit tests and most CI runs that only need an authenticated user.

Drive the real provider for at least the critical happy path, because that is the only way to catch the failures mocking hides at the redirect-to-callback boundary.

Do not re-drive the identity provider on every test. Run the real login once, save the session with Playwright storageState or Cypress cy.session, and start every other test already signed in.

Driving the live provider hundreds of times per run can trigger rate limits and bot-detection checks that read as flakiness.

Use a dedicated test account with MFA disabled for the sign-in path, or generate the second-factor code yourself instead of waiting on a real device.

The mechanics of computing a TOTP or reading an emailed code in a test are covered in our guide to testing 2FA and OTP flows.

Yes, on compatible flows. Pie navigates the account picker and consent screen visually, so it can adapt to a provider UI change without a hard-coded selector.

It works from a dedicated test account you control. Google and SAML controls can still block automated sign-in, so set up the tenant to the provider's testing rules.

Adithya Aggarwal
Adithya Aggarwal
CTO & Co-founder at Pie

Eight years building search and delivery systems at Amazon. The kind of scale where flaky tests block billion-dollar releases. Now CTO at Pie, building AI agents that adapt when your UI changes. LinkedIn →