Article · Flaky Test Detection & Quarantine Engineering

Sending Slack Alerts for Flakiness SLO Breaches

Routing a flakiness SLO breach to Slack puts the alert in front of the on-call engineer the moment a suite crosses its threshold, without anyone watching a dashboard. This guide extends the parent section on reliability dashboards for QA teams by showing how to fire a Slack incoming webhook from CI when the SLO is breached, how to write threshold logic that distinguishes a real regression from a single bad run, and how to dedupe so the channel does not become noise people mute.

13 sections URL: /flaky-test-detection-quarantine-engineering/reliability-dashboards-for-qa-teams/sending-slack-alerts-for-flakiness-slo-breaches/
Flakiness SLO breach decision flow to Slack A computed flake rate is compared to the SLO; only a sustained breach that is not already open posts a Slack message. flake rate from CI rate > SLO? yes already open? no post to Slack no → skip yes → dedupe
Only a sustained breach that is not already open results in a Slack post, suppressing duplicate noise.

Root cause #

An alert that fires on every flaky run trains people to ignore it. Flakiness is, by definition, intermittent — a suite at a 3% flake rate will produce a failing run roughly one time in thirty even when nothing has regressed. If the alert condition is “this run had a flaky test,” the channel fills with single-run noise and the signal of a genuine SLO regression is lost. The mechanism that fixes this has two parts: evaluate the breach against a rolling aggregate rather than a single run, and dedupe so the same open breach does not re-post on every subsequent pipeline.

Slack incoming webhooks are the right transport because they are a single authenticated URL with no OAuth scopes to manage and no per-message rate ceiling that CI will realistically hit. The webhook URL is a secret — anyone holding it can post to the channel — so it lives in repository secrets as SLACK_WEBHOOK_URL, never inlined. The hard part is not sending the message; it is deciding whether to send it at all.

Alert on a sustained breach, not a run At any non-zero flake rate a single run fails occasionally; the alert must key on a rolling aggregate. per-run alertchannel muted rolling-window breachreal signal post once
Evaluating against a 7-day rate stops single-run noise from training people to ignore the channel.

Step-by-step fix #

1. Compute the breach against a rolling window #

Compare the rolling flake rate to the SLO, not the latest run. A single run is too noisy to gate an alert.

// scripts/evaluate-slo.js
// Trade-off: a longer window is more stable but slower to react to a real
// regression; 7 days balances responsiveness against single-run noise.
import { readFileSync } from 'node:fs';

const SLO = 2.0; // percent of executions allowed to be flaky
const history = JSON.parse(readFileSync('flake-history.json', 'utf8')); // last 7d of runs
const flaky = history.reduce((n, r) => n + r.flaky, 0);
const total = history.reduce((n, r) => n + r.total, 0);
const rate = (flaky / total) * 100;

const breached = rate > SLO;
console.log(JSON.stringify({ rate: rate.toFixed(2), slo: SLO, breached }));
process.exit(breached ? 1 : 0);

2. Dedupe so an open breach does not re-fire #

Persist a marker (a cache key or a tiny state file in an artifact) so a breach that is already open is not re-announced every pipeline. Only post on the transition into breach.

// scripts/should-alert.js — true only on the no→yes transition
// Trade-off: state in CI cache can be lost on eviction, causing a re-alert;
// acceptable because a re-alert on an open breach is rarely harmful.
import { existsSync, writeFileSync, rmSync } from 'node:fs';

const breached = process.argv[2] === 'true';
const MARKER = '.flaky-breach-open';
const wasOpen = existsSync(MARKER);

if (breached && !wasOpen) { writeFileSync(MARKER, '1'); console.log('alert'); }
else if (!breached && wasOpen) { rmSync(MARKER); console.log('recovered'); }
else { console.log('skip'); }

3. Post the Slack webhook from GitHub Actions #

Send a Block Kit payload only when step 2 decided to alert, reading the webhook from the SLACK_WEBHOOK_URL secret.

# .github/workflows/flaky-slo.yml
- name: Alert Slack on SLO breach
  if: steps.dedupe.outputs.decision == 'alert'
  env:
    SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK_URL }}
    RATE: ${{ steps.evaluate.outputs.rate }}
    RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
  run: |
    # Trade-off: a rich Block Kit payload is more actionable but harder to
    # template; keep it to one section + one button to stay maintainable.
    curl -sS -X POST "$SLACK_WEBHOOK_URL" \
      -H 'Content-Type: application/json' \
      -d "{\"blocks\":[
        {\"type\":\"section\",\"text\":{\"type\":\"mrkdwn\",\"text\":\":rotating_light: *Flakiness SLO breached* — rate *${RATE}%* exceeds 2% on \`${GITHUB_REF_NAME}\`\"}},
        {\"type\":\"actions\",\"elements\":[{\"type\":\"button\",\"text\":{\"type\":\"plain_text\",\"text\":\"View run\"},\"url\":\"${RUN_URL}\"}]}
      ]}"

4. Send a recovery notice #

Closing the loop matters as much as opening it. Post a recovery message on the yes→no transition so the channel knows the breach cleared, and route persistent offenders into building auto-quarantine workflows rather than alerting forever.

- name: Post recovery
  if: steps.dedupe.outputs.decision == 'recovered'
  env:
    SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK_URL }}
  run: |
    # Trade-off: recovery notices add volume but prevent stale "is it fixed?"
    # pings; keep them terse.
    curl -sS -X POST "$SLACK_WEBHOOK_URL" -H 'Content-Type: application/json' \
      -d '{"text":":white_check_mark: Flakiness SLO recovered — rate back under 2%."}'
Dedupe on the state transition A marker file makes the alert fire only on the no→yes transition and recover on yes→no. not breached breached post once (no→yes)recover on yes→no
Posting only on the transition keeps an open breach from re-firing every pipeline.

Pitfalls #

  • Alerting per run. Single-run failures are expected at any non-zero flake rate. Mitigation: evaluate against a rolling window, not the latest run.
  • No deduping. Re-posting an open breach every pipeline gets the channel muted. Mitigation: alert only on the no→yes transition.
  • No recovery message. People keep asking “is it fixed?”. Mitigation: post on the yes→no transition too.
  • Webhook URL leaking into logs. Echoing the env var exposes a postable secret. Mitigation: never echo it; keep it in repository secrets only.
  • Threshold set at the current baseline. An SLO equal to today’s rate fires constantly. Mitigation: set the SLO above the current baseline, then ratchet down.
Close the loop with recovery A recovery notice on yes→no tells the channel the breach cleared; chronic offenders go to quarantine. breach open recovered notice chronic → quarantinenot a recurring page
A chronically flaky suite is a backlog item, not a recurring page — route it to quarantine.

Reliability targets #

Metric Target Notes
Flakiness SLO ≤ 2% over 7d Rolling-window breach condition
Alert dedupe 1 post per breach Only on no→yes transition
Time to alert < 5 min after run Webhook POST is near-instant
False-alert rate < 1 / week Achieved via rolling window
Recovery notice 100% of closed breaches On yes→no transition
Slack-alert scorecard Targets for the SLO, dedupe, time to alert, and false-alert rate. ≤ 2%SLO (7d) 1post / breach < 5 mintime to alert < 1/wkfalse alerts
One post per breach with a recovery notice keeps the channel meaningful.

Frequently Asked Questions #

Q: How do I stop the alert from firing on every flaky run? A: Evaluate the breach against a rolling 7-day flake rate rather than a single run, and post only on the transition into breach. A test suite at any non-zero rate will occasionally fail a run without that being an SLO regression.

Q: Where should the Slack webhook URL live? A: In repository secrets, as SLACK_WEBHOOK_URL. The URL is itself the credential — anyone with it can post to the channel — so it is never inlined or echoed to logs.

Q: What should happen to a suite that breaches repeatedly? A: Stop alerting and start quarantining. A chronically flaky suite is a backlog item, not a recurring page; route it into an auto-quarantine workflow so the channel stays meaningful.

Routing Determines Whether Anyone Acts #

An alert delivered to a general engineering channel is delivered to nobody in particular, and for any given alert the reasonable assumption for each reader is that it concerns someone else. That assumption is correct often enough that it becomes a habit, and the habit outlives whatever prompted the alert.

Routing to the owning team’s own channel changes the economics: the audience is small enough that responsibility is unambiguous, and the message arrives where that team already works rather than requiring them to visit somewhere. The mapping needed is one that most repositories already maintain for review routing, so deriving the destination from ownership rather than from a hand-kept list keeps it current for free.

Two details make routed alerts materially more useful. Including the owning team’s handle in the message body makes the responsibility explicit even when the channel is shared. And including the specific tests driving the breach, rather than only the aggregate, converts the alert from a status update into a task — a rate over a threshold prompts discussion, while three named tests prompt a fix.

Where an alert genuinely concerns everyone — the whole pipeline is degraded, the runners are failing — a broad channel is correct, and that rarity is what makes it noticed.

Choosing What Deserves an Alert at All #

Alerting is a claim on attention, and a reliability programme can exhaust that budget quickly. Three conditions are worth alerting on; most others are better as a weekly digest.

A sustained SLO breach. The rolling rate has crossed the threshold and stayed there, which means something changed rather than that a run was unlucky. This is the primary alert and it should fire once per breach, with a recovery notice when it clears.

An infrastructure failure rate spike. Distinct from test flakiness, owned by a different team, and usually actionable immediately — a runner image problem or a service outage affects everything and is worth interrupting for.

A quarantine expiry that has passed. Time-based rather than metric-based, and it needs a nudge because nothing else will produce one.

Everything else — a single flaky run, a test crossing a per-test threshold, a small movement in the aggregate — belongs in the weekly digest where it can be read alongside the ranked worklist. The test of whether an alert deserves to exist is simple: if the expected response is “yes, we know”, it should not be an alert.

Silence Is a Feature #

An alerting setup that stays quiet for weeks and fires once, accurately, is working exactly as intended.

Including the specific tests driving a breach, not just the aggregate, is what converts an alert into a task. A rate over a threshold prompts discussion; three named tests with owners prompt a fix.