Hive Hive
Sign in

Stress-testing newly added tests before they merge

#88 · Tuist · Public · Created directly

CLI Testing
Draft
Proposal

Important

Verified against the codebase and the prior art on 2026-08-31, with the duration curve and both guard thresholds measured against production analytics the same day. New-test detection already exists, on the wrong side of the run: check_new_test_cases/3 stamps is_new at ingestion, after the tests have run. The client seam already exists twice: both clients fetch a server-computed set of test identifiers before a run (quarantine) and both execute a server-issued plan of identifiers (sharding). Selective testing is on by default and decides what runs, at test-target granularity, which invalidates both the guard the prior art uses and the window is_new relies on. A run whose repetitions disagree already means “flaky” everywhere, which is why this needs almost no new rendering and why it is dangerous by default.

Summary

This RFC lands a stress gate: on a merge request, the tests that branch adds are rerun several times each, and the job fails if one disagrees with itself. Tuist detects flakiness only after the fact today, so the first person to pay for a flaky test is someone who did not write it, days later, in an unrelated pull request. The gate moves detection to the moment the fix is cheap.

Tuist is already at parity with the market on the passive half of flaky-test management. The gap is the active half, which one product ships, and it ships it without the block. Four decisions are expensive to reverse once shipped:

  1. The server owns the newness verdict, the client owns the loop, so Gradle and Bazel cost a filter translation rather than a reimplementation.
  2. The client activates the gate, the server parameterises it. Enforcement is the CLI’s exit code, so what an enforcer does belongs in the repository next to the command.
  3. The candidate set comes from the run that just happened, inside the same invocation. Selective testing forces the second half.
  4. The gate blocks rather than tags, through an exit code with no VCS dependency, and its evidence is recorded so the machinery already listening for flakiness does not mistake it for organic evidence.

Motivation

A customer with a very large iOS suite already does this by hand: CI reruns every newly added test and blocks the merge request if the runs disagree. They want Tuist to own it, and the instinct is right, because the expensive half is the half they have to approximate and we do not.

Anyone can write the loop. What a homegrown version cannot do is decide what “new” means. It has to infer it from the diff, and a diff does not know about XCTest cases inherited from a base class, Swift Testing’s @Test display names, per-argument identity on parameterised cases, or annotation-driven discovery on the JVM. Each is a place where the homegrown gate stresses nothing and reports green. Tuist holds the default branch’s test history, which answers the question directly.

The retrospective path is not a substitute. first_run events already say a test is new and the flakiness monitors already say it is unreliable, but they arrive in that order, separated by however long the test takes to misbehave on the default branch. Datadog reports this approach catches up to 75% of flaky tests before they merge, which is their population and not ours, but the right order of magnitude to stop arguing about whether repetition finds flakes.

What a user gets

One opt-in option, and the off state is not passing it

--stress-new-tests <report|enforce> on tuist test and tuist xcodebuild test, with the paired TUIST_TEST_STRESS_NEW_TESTS env key every other test flag has. The three states are off, report and enforce. Off is what you get by not passing the option, so the gate is opt-in and a project that never adds it never runs a stress pass, however many tests its branches add. Upgrading a CLI or a plugin therefore cannot hand anyone a merge gate, and adoption is a per-pipeline edit rather than a switch somebody flips. This is the part hardest to change once shipped, so the shape matters more than the name.

It takes a value rather than being a boolean, because a gate that blocks and a gate that reports must not look the same in a CI file: anyone reading the pipeline can see whether this job can fail on flakiness. It also removes the question of a default mode: there is no default, so whoever adds the option decides in the same edit whether the job can fail.

On tuist xcodebuild test it sits with the other Tuist options ahead of the passthrough arguments, not after the terminator. Passing -test-iterations yourself after the terminator is not this feature: it repeats everything rather than the new tests, and nothing blocks on the result.

On Gradle it is a stressNewTests block on the existing tuist extension in settings.gradle.kts, beside buildCache and testQuarantine, with the same environment variable as the per-lane override. The plugin has no Gradle-property convention and reads System.getenv throughout, which is why the env form rather than the declaration is what varies behaviour per matrix lane on both clients.

Off-by-default must not follow the local idiom. TestQuarantineExtension.enabled is a nullable boolean meaning “on for CI, off locally”; copying it here would be a mistake, because quarantine only ever makes a run greener while a gate makes runs redder. Absent means off, on every client, in every environment.

Bazel has no surface yet and this RFC does not invent one. With no test ingestion there is no verdict to configure, so a surface designed now is a surface to unbuild. The contract it will have to satisfy is fixed: the same three states with off as the unconfigured one, --runs_per_test as the loop, expressed through .bazelrc and a per-lane --config.

report is enforce minus the exit code, and that is a requirement

The verdict query, the candidate set, all three guards, the retry curve, the slow exclusion, the upload, and the way repetitions are recorded are identical in both modes. If any differed, the soak would measure something that does not predict enforcement and the ramp would be worthless.

The only divergence is what happens when a candidate disagrees with itself. report emits a warning and a takeaway naming the mode to switch to, and the process exits on the first pass’s own result. enforce fails the run with the same non-zero exit the CLI already returns for a failed test run; a distinct exit code would only matter to a script nobody has written.

report is the adoption ramp, and it matters more than which pipeline goes first: a team watches what the gate would have blocked for a couple of weeks, then changes one word, without editing a workflow twice and without a fortnight of blocked merges to discover the curve was wrong for their suite. The first customer starts here regardless, because the catch rate and the added wall clock are only knowable from their own data.

What it prints

Output is prose, because the CLI’s output is prose: a .section line, one line per candidate with its repetition count and outcome, and AlertController warnings and takeaways at the end. Roughly, in report:

Stress-testing new tests
3 new test cases, 3 stressed, 0 excluded
AppFeatureTests/CheckoutTests/testAppliesDiscount 10 repetitions, 2 failed
AppFeatureTests/CheckoutTests/testClearsCart 10 repetitions, passed
AppFeatureTests/CartTests/testTotal 10 repetitions, passed
! testAppliesDiscount failed 2 of 10 repetitions and would have blocked this merge
request. Pass --stress-new-tests enforce to make it blocking.

enforce prints the same up to the blank line, then throws with the same sentence minus “would have”. When a guard fires it replaces the per-candidate list with its reason and the count that triggered it. A branch that adds nothing prints nothing and runs no pass, which is the common case.

When it does nothing, which is most of the time

  • The option was not passed, or the account is not entitled: nothing runs. This is every project by default, and the account entitlement is what keeps the feature dark before it is generally available.
  • A branch that adds no tests: no candidates, no pass, no added wall clock. Costs one request.
  • The first pass already failed: no pass, reported as skipped. The run is red either way and the budget is better spent after the visible failure is fixed. A useful property falls out: every candidate reaching the stress pass passed on its first execution, so it can only pass every repetition or disagree, and “failed 2 of 10” is unambiguous.
  • No default branch, or no default-branch history: the premise guard fires visibly and the run exits zero. Same answer check_new_test_cases/3 already gives.
  • A bulk identity change, such as renaming a test module: the bulk-change guard fires and names the count, so the cause is legible.
  • The server is unreachable: no verdict, no pass, the run’s own result stands. The gate never blocks a merge because Tuist was down, matching what the CLI already does when the quarantine fetch fails.

Every bound reports when it bites. A gate that quietly stresses twelve of forty candidates and reports green is worse than one that refuses.

What it costs

Only the tests the branch added are rerun, and each is priced by its own measured duration: at or under 5s gets 10 repetitions, under 10s gets 5, under 30s gets 3, under 5 minutes gets 2, slower is excluded and reported as excluded. Two guards bound the rest, a candidate cap (starting value: 200 cases) and a wall-clock ceiling on the pass (starting value: 10 minutes). No rebuild happens; the pass reuses what the first pass built. Fleet-wide 94% of CI test cases land in the top bucket and the median averages 4ms, so for the ordinary merge request adding two or three tests the cost is seconds against a suite of thousands. The ceiling exists for the tail: two hundred candidates at the p90 of one second would run to about half an hour without it.

All of these are delivered from the server and tuned from telemetry rather than re-litigated here. Two behaviours are decisions rather than knobs: all repetitions run rather than stopping at the first failure, because “failed 3 of 10” and “failed 1 of 10” are different conversations for the author; and each repetition gets a fresh process, because in CI every run is a fresh process, and same-process repetition lets a test pass ten times and still flake once per real run.

Where the result shows up

The evidence needs no new rendering. A candidate that disagreed is already badged New in the dashboard, already draws its per-repetition pass/fail strip on the test case and test case run pages, and already reaches the flaky section of the “Tuist Run Report” comment.

That it renders for free is the hazard, not the win. test_case_is_flaky?/1 keys on mixed repetition statuses, so a stress finding is indistinguishable from organic flakiness, and being marked flaky is upstream of auto_mark_flaky_tests (default on, threshold 1), the cooldown, the alerts and the comment. All of that would fire in report mode, where the author was told nothing would happen. A preview that silently marks tests flaky project-wide is not a preview, so the pass must be legible to those consumers as one gated observation. This is hard to undo once the data has landed, so the distinction has to exist before the first stress pass is ever ingested.

The line runs between consumers that attribute and consumers that describe. The cross-run aggregates, auto-marking, the cooldown and the alerts all attribute flakiness to the test, so stress repetitions must be invisible to them. The dashboard run view and the run report comment describe one run, so they show the finding.

In the comment it belongs in the flaky section, marked as a different kind of flaky test. A reviewer asking what is unreliable about this branch should find one list rather than two, and a section of its own would split one concern across two places in the same comment. What has to be visible is the provenance: this test was caught by the stress gate under deliberate repetition, not observed flaking in the wild. Same section, distinct label, because they are the same concern reached by different routes and only the route differs.

What no surface has is the verdict. The evidence says a test disagreed; nothing says the gate ran, how many candidates it examined, which guard fired, or that this would have blocked. That belongs on the test run in the dashboard, and it is not optional polish: the comment is GitHub-only, so for the customer driving this the dashboard is the only place the verdict can live outside CI scrollback. Rendering it in the comment is the same upgrade as a check run, and waits on the same thing.

How it interacts with quarantine

A new test that is already muted cannot fail the gate. Muting means the test runs and its failure is masked, and the gate inherits that mask rather than overriding it, so a team that decided to tolerate a test does not have that reversed by a feature enabled for another reason. Skipped tests never run, so they never become candidates. Muted candidates are still stressed and recorded, because a new muted test failing four of ten is what a reviewer wants before the mute becomes permanent.

A test that fails the gate blocks; it is not auto-quarantined, even where auto_quarantine_flaky_tests is on. test_case_ids_with_successful_default_branch_run/3 already keeps cases with no successful default-branch run out of automated quarantine, and a gate failure is by definition in that set. Auto-quarantining would mute a test nobody has ever trusted and let it land muted, converting “fix this before it lands” into “land it broken and forget”. --skip-quarantine is unaffected and independent.

Current state

  • is_new is decided at ingestion, over a trailing 90-day window, from test_case_branch_presence, and marks nothing new when the default branch is unset. It runs after the tests do, because the row it decorates is the one being written.
  • test_case_branch_presence retains one row per case per branch, forever. A ReplacingMergeTree(ran_at) keyed on (project_id, git_branch, is_ci, test_case_id), whose migration notes the ran_at comparison is a filter applied after the primary-key search rather than part of the key.
  • Selective testing is on by default and decides what runs. SelectiveTestingGraph hashes each test target with its transitive dependencies, SelectiveTestingServicing.cachedTests returns already-tested targets as -skip-testing, and a successful run reports each target’s hash back through the command event’s selective_testing_metadata, which is what makes the next run skip it. Granularity is the test target, not the case.
  • Quarantine is already a pre-run server fetch on both clients, with a post-run mask. Sharding is already a server-issued plan executed by both clients, applied through -only-testing and includeTestsMatching, with an inventory read from branch history so a suite the server has never seen lands in the catch-all shard.
  • The client already holds the per-case result set locally, with durations, from parsing the result bundle before any upload. Repetitions already ingest into test_case_run_repetitions, and xcodebuild produces them via -test-iterations and -test-repetition-relaunch-enabled.
  • A disagreeing run already means flaky, and that starts machinery. test_case_is_flaky?/1 keys on mixed repetition statuses; auto_mark_flaky_tests defaults to true at threshold 1, alongside flaky_cooldown_days and the flaky alerts, all on the project record. test_case_ids_with_successful_default_branch_run/3 keeps never-validated cases out of automated quarantine.
  • The only merge gate today is server-enforced and GitHub-only. Bundle-size thresholds are evaluated by BundleThresholdWorker and delivered as a check run; the CLI uploads and never fails on a breach. The only VCS API client is GitHub’s, and Tuist.VCS knows GitLab only well enough to build a pipeline URL.

Prior art

Buildkite, Trunk and Develocity are all passive: Buildkite detects from one test disagreeing on one commit SHA and has a first-execution monitor, which is feature for feature what Tuist already has; Trunk explicitly reruns nothing; Develocity detects via retry-on-failure and matching input fingerprints. None reruns a new test on purpose, and none blocks a merge.

Datadog Early Flake Detection is the only shipped implementation of the active loop. From the ddtrace sources: settings and a known-tests list are fetched from the backend before the session, anything absent is new, and new tests are retried on a duration-keyed curve (5s: 10, 10s: 5, 30s: 3, 5m: 2, excluded above five minutes), with a faulty_session_threshold of 30% disabling the mechanism when too much of a session reads as new. The curve and the server-delivered parameters are adopted; the session-proportion threshold is not, for reasons below; and EFD deliberately does not block, since blocking is a separate Quality Gates product. That split is coherent for a platform that sells gates and incoherent for us, because the customer asking is already blocking by hand.

Measured against the fleet (2026-08-31)

Three of the questions this RFC had left open were answerable from production analytics rather than argument. All figures are fleet-wide aggregates over projects with CI test runs, not one customer’s suite, and avg_duration is a per-case average rather than a first-attempt measurement.

  • The inherited duration curve fits, and the worry that it would not was wrong. Across the 3.72M test cases that have run in CI, 94.4% average at or under five seconds and so receive the full ten repetitions. 1.9% fall in the ten-second bucket, 1.8% in the thirty-second bucket, 1.9% in the five-minute bucket, and 0.013% are excluded for exceeding five minutes. The median case averages 4ms and the p90 is about one second. The curve’s top bucket is where almost everything lands, so the gate is not quietly weakened by test durations, and the slow exclusion is a rounding error rather than a hole.
  • The bulk-change guard needs an absolute floor, because most projects are small. The median project has 209 test cases that have run in CI, the p25 has 57, and 151 of 441 projects have fewer than 100. A bare 30%-of-inventory ratio fires at 63 new cases on a median project and under 30 on a small one, which a single feature branch can legitimately produce.
  • Per-project parameters are right, because almost nobody has many projects. Of 1,071 accounts with test runs in the last ninety days, 967 have exactly one project and only 27 have five or more. The “forty projects to configure” case is real but rare, which makes an account-level default an upgrade rather than a v1 requirement.

Proposal

The client activates the gate, the server parameterises it

The existing merge gate is server-configured because it is server-enforced. The bundle gate evaluates in a worker and delivers a check run, so the client is a data source and its cooperation is not required. This gate is the opposite: enforcement is the CLI’s exit code, in the client’s process, on CI the server never observes. Where an enforcer takes its instruction from is a different question with a different answer.

A remote switch that silently changes every pipeline is the failure mode to avoid. As a project setting, someone flips it and every job starts failing merge requests with no diff to review and nothing to revert, and two runs of the same commit stop agreeing. A flag beside the command is reviewable, revertable and bisectable, and it lets adoption go one pipeline at a time rather than all of them at once on a merge gate. The counter-argument, that a branch can switch the gate off in the merge request it would have blocked, proves too much: the same merge request can delete the test job, and every in-repo check has that property, answered by required status checks and ownership rules.

The parameters still have to be server-side, or the tuning story in this RFC is not achievable: in a flag or in ProjectDescription every adjustment is a CLI release plus a bump in every project, while from the server it is a settings write. They ride in the verdict response, so there is no second round trip and no way for client and server to disagree.

The parameters are per project, beside the flaky-test policy already on the project record, and not per account. That is where the guards have to live anyway, since the inventory they measure is a property of one project, and splitting the curve away from them would fracture one settings surface across two scopes. Nine in ten accounts have exactly one project, so the account-level default is an upgrade for the few who would notice.

Central enforcement is a deliberate non-goal for v1 and a clean upgrade. An organisation that later wants branches unable to opt out gets a project setting that forces the gate on regardless of the flag, joining the flaky-test policy block already on the project record. Entitlement stays FunWithFlags per account, the pattern runners_enabled?/1 already establishes. Entitlement, activation and parameters are three questions and collapsing any two costs more than keeping them apart.

The candidate set comes from the run that just happened, in the same invocation

Diff-based detection is rejected because it fails silently: parsing test declarations out of source is per language and per framework, every framework has a construct that defeats it, and the output of getting it wrong is an empty candidate set and a green gate.

The two-pass run wins and its first pass is not a cost, because CI runs the suite anyway. What falls out is the exact set of cases that executed, at case granularity, already in the client’s memory. The client sends those identifiers, gets back the subset with no default-branch history, and reruns exactly that subset. The set is observed rather than inferred, and the same pass measures each candidate’s duration, which is what prices its repetitions; a single-pass design has to guess.

The stress pass runs inside the same invocation, and this is forced. A successful run reports its test-target hashes as selective-testing hits, so a second tuist test invocation on the same tree would find every target it just tested to be a hit, skip all of them, and report success having stressed nothing.

The gate must not read its verdict back out of ingestion. is_new is stamped during ingestion, which on the result-bundle path is an asynchronous job, and that queue has wedged in production. The verdict is answered at request time from the same branch-presence data.

Selective testing decides what runs, so no guard is a proportion of what ran

The prior art’s session-proportion threshold would switch the gate off in the case it exists for. A new module is added with two tests and nothing else changes. Selective testing skips every unchanged target, so the session is those two tests, 100% new, and a 30% session threshold declares it faulty and stresses nothing. The better the cache works, the more reliably the gate refuses. That is the modal case inverted, not an edge case.

A session’s composition states which targets changed, not whether the project’s history is trustworthy. Splitting the two gives the guards a user sees fire: a premise guard on whether the server has any default-branch history for this project at all, which selective testing cannot move; a bulk-change guard keeping the ratio but with the inventory as its denominator (starting value: 30%), which catches identity changes such as a renamed test module; and the candidate cap, because cost genuinely is a session property.

The bulk-change guard applies only above an absolute floor of new candidates (starting value: 50). The median project has around two hundred test cases that have run in CI and a third of projects have fewer than a hundred, so a bare ratio would fire on a branch that legitimately added twenty. Below the floor the candidate cap is the only bound, which is safe because the cost there is small by definition.

Sharding needs no fan-in. A shard matrix partitions the test set, so a new test runs in exactly one shard, is stressed by that shard, and the union of the per-shard candidate sets is exactly the set of new tests that ran. Per-shard gating is complete rather than partial. The shards also merge into one test run server-side, so the dashboard verdict aggregates across them without an extra step, and the only thing that ever needed care here was the guards, which moving them off session proportions already fixed.

“New” means ever, not recently

The gate asks whether a case has ever run in CI on the default branch, dropping the trailing window. Selective testing is again the reason: a stable module is not retested on the default branch while its dependency closure is unchanged, so after ninety days its rows leave the window and its whole test target reads as new, an error that grows with cache effectiveness. Dropping the window is free, since the rows are retained indefinitely and the comparison is a post-primary-key filter.

The window is correct where it is, because check_new_test_cases/3 feeds a trailing-window aggregate whose own comment notes a rename “heals on its own once the window has moved past it”. A gate is read once, so self-healing in three months is not a property it can use. The consequence to document is that the new-test badge and the gate can disagree about a long-dormant test, and both are right.

The verdict is a subtraction, and the retry primitives cannot be the gate

A server-issued plan of what to stress cannot be built: the shard plan derives its inventory from branch history, and history is what a new test lacks. So the candidate set originates on the client and the server subtracts, keyed on the (name, suite_name, module_name) identity the tests tables already share, which is what makes that one query serve every client that writes into them.

-retry-tests-on-failure, --flaky_test_attempts and the Test Retry Gradle plugin all rerun only what already failed, which excludes the entire population the gate examines: tests that passed first and would fail fourth. Their purpose is to hide flakiness. The primitive wanted is unconditional repetition.

What each build system contributes, stated asymmetrically because it is asymmetric

  • Xcode: -only-testing for the set from identifier lists the CLI already assembles, -test-iterations with relaunch for the loop, results already at home in test_case_run_repetitions. Both entry points route through TestService, so it lands once.
  • Gradle: Test.filter.includeTestsMatching for the set, as TuistTestSharding.kt already does; the loop needs explicit re-execution, because up-to-date checks otherwise decline to run the task twice on unchanged inputs. Same hazard as selective testing, same resolution.
  • Bazel: our value is the verdict and the gate, not the loop. --runs_per_test already repeats better than anything we would write. What Bazel lacks is history to subtract against. It is blocked on ingestion, not on this, and the obligation is to build nothing that must be unbuilt: keep the verdict on the shared identity and the loop behind the client. See spec #74.

What this gate cannot catch

Running a new test alone, repeatedly, catches self-contained nondeterminism and nothing else. A test that passes alone and fails after another has polluted shared state passes this gate, as does one that pollutes state others depend on. The prior art shares the limitation, so it is not a reason to prefer anything else, but the gate would be oversold without saying it. Catching it means running the new test inside its module’s ordering, or shuffling, which multiplies cost by the module rather than the test.

Scope

In scope: the --stress-new-tests option and its env key in both modes on tuist test and tuist xcodebuild test, plus the equivalent Gradle block; the request-time verdict on the server, answering “ever run in CI on the default branch” and returning the parameters and guard signals in one response; the parameters stored beside the flaky-test policy already on the project record; the per-account FunWithFlags entitlement; the second pass inside the existing TestService invocation and its Gradle counterpart; the retry curve, both premise guards with the bulk-change floor, the candidate cap, the slow exclusion and the reporting of each; recording the pass so its consumers can tell a gated observation from an organic one, covering the aggregates, auto-marking, the cooldown and the alerts; and the gate’s verdict on the test run in the dashboard.

Out of scope, but aligned: Bazel and its configuration surface, both blocked on test ingestion (spec #74); the check run and any rendering of the verdict in the run report comment, both blocked on a GitLab client; order-dependent flakiness; the ingestion-time window in check_new_test_cases/3; any expression of the parameters in ProjectDescription; account-level parameter defaults; the project setting that forces the gate on regardless of the flag; and auto-quarantine of gate failures.

Trade-offs

Advantages

  • Catches a flaky test while its author still holds the context that produced it.
  • Ships the active half the market mostly lacks, plus the block the one product with the loop sells separately.
  • Adds no new seam: the verdict rides the pre-run fetch quarantine already uses, the identifier contract sharding already uses, and surfaces that already render the evidence.
  • Adoption is incremental and reviewable, with a report mode that differs from enforcing only in the exit code.
  • A CI result stays reproducible from the repository, while tuning stays a settings write rather than a CLI release.
  • Works with selective testing rather than against it: the case it reduces to a two-test session is the case the gate most wants.
  • The inherited curve is measured to fit rather than assumed to: 94% of CI test cases get the full ten repetitions and a thousandth are excluded as too slow.
  • Ships to a GitLab customer on day one, and every other CI provider at the same time.
  • Costs nothing on the overwhelming majority of merge requests, which add no tests.

Disadvantages

  • Catches only self-contained flakiness; order-dependent flakes pass.
  • Repetition is a probabilistic filter, weakest exactly where tests are slowest: a four-minute test gets two repetitions and a rare flake in it usually lands. That is about 2% of cases fleet-wide, but nothing says the tests a given branch adds follow the fleet distribution.
  • Activation is per invocation, so a team with many pipelines adds the option to each and a new pipeline does not inherit the gate. That is the cost of the property that makes adoption gradual.
  • Until the central-enforcement override exists, nothing stops a branch removing the option in the merge request the gate would have blocked.
  • The Xcode and Gradle spellings differ, so the documentation carries two forms of one contract with the env variable as the only thing common to both.
  • The gated-observation distinction has to reach several existing consumers, and every one missed turns solicited evidence into an organic-looking flaky signal.
  • Job duration now varies with how many tests the branch adds. Bounded, but no longer constant.
  • The gate and the new-test badge can disagree about a long-dormant test. Correct in both places, confusing the first time it is noticed.

Alternatives considered

A project setting as the primary switch

This RFC’s earlier position. Rejected once the enforcement shape was taken seriously: the bundle gate can be configured server-side because the server enforces it, and this gate is enforced by the client’s exit code. A remote toggle that changes every pipeline with no diff and no revert is the wrong shape for something that blocks merges, and it forecloses adopting one workflow at a time. It returns as the optional override.

A Tuist.swift option

Rejected because ProjectDescription is versioned with the CLI, so every tuning change becomes a release plus a manifest bump everywhere, and because the manifest is not where CI behaviour is expressed. The Gradle extension is the closest thing to a manifest declaration here, and only because the plugin has no other configuration surface.

A boolean flag with a separate report-only flag

Rejected because whether a job can fail on flakiness should be readable at the call site, and because two booleans admit a meaningless fourth state.

Inheriting the quarantine extension’s CI-aware default on Gradle

Rejected because the two features carry opposite risk: quarantine only ever makes a run greener, while a gate defaulting on turns a plugin upgrade into failed merge requests nobody asked for.

Let the stress pass be flaky evidence like any other

What happens if nothing is done, since test_case_is_flaky?/1 already keys on exactly that shape. Rejected because it makes report mode a lie: auto-marking, the cooldown, the alerts and the run report comment would all fire on evidence the gate solicited, in the mode whose entire promise is that nothing happens yet.

A section of its own in the run report comment

Give the gate’s findings their own heading rather than marking them inside the flaky section. Rejected because a reviewer asking what is unreliable about this branch would then have two lists to read for one concern. The provenance is what needs to be visible, not the separation, so the finding sits in the flaky section under its own label.

Stress even when the first pass failed

Rejected because the run is already blocked, so the verdict changes no outcome while costing the full wall-clock budget, and the second finding lands on someone who cannot act on it until the first is fixed.

A faulty-session threshold as a proportion of the session

The prior art’s guard. Rejected because selective testing controls the denominator, so the ratio measures which targets changed rather than whether history is trustworthy. Survives with the denominator moved to the project’s inventory and an absolute floor under it.

A fan-in step after a shard matrix

Aggregate the per-shard verdicts before deciding. Rejected as unnecessary: sharding partitions, so every new test is stressed by exactly one shard and the union is complete, and the shards already merge into one test run for reporting.

Run the stress pass as a separate CI step

Simpler to explain and keeps the first job’s duration predictable. Rejected because the first pass reports its target hashes as selective-testing hits, so the second invocation would skip everything it just tested and report success having stressed nothing.

Push the known-test set to the client and diff locally

Datadog’s model, and single-pass. Rejected for v1 on client shape rather than merit: it needs an in-process hook into the test framework, which their tracer has and our CLI does not, and it moves a set proportional to the whole suite rather than to the branch. Worth revisiting on Gradle, where the plugin does have test listeners.

Tag rather than block, as the prior art does

Rejected because tagging is the feature Tuist already has. The customer is not asking for a better label, they are asking for the merge to stop.

References

  • Spec #74, Kura REAPI Cache Instrumentation for Bazel: the ingestion prerequisite for the verdict half on Bazel.
  • Datadog Early Flake Detection, with parameters taken from the ddtrace sources (_api_client.py, retry_handlers.py, api/_session.py) rather than the documentation, which omits them.
  • Buildkite Test Engine, Trunk Flaky Tests, Develocity: the passive half Tuist already matches.
  • Production ClickHouse, queried 2026-08-31 for the Measured section: test_cases for the duration distribution and per-project inventory sizes, test_runs for projects per account.
  • server/lib/tuist/tests.ex: check_new_test_cases/3, test_case_is_flaky?/1, test_case_ids_with_successful_default_branch_run/3.
  • server/lib/tuist/vcs.ex: the run report comment, its tests table and its failed and flaky sections.
  • server/lib/tuist_web/live/test_case_live.html.heex, test_case_run_live.html.heex: the is_new badge and the per-repetition strip that already render the evidence.
  • server/lib/tuist/bundles/workers/bundle_threshold_worker.ex: the one existing merge gate, and why its configuration can live entirely on the server.
  • cli/Sources/TuistKit/Services/TestService.swift, TestQuarantineService.swift, Sharding/ShardPlanService.swift, Commands/XcodeBuild/XcodeBuildTestCommand.swift: the pre-run fetch, the post-run mask, the server-issued plan, and the passthrough shape the option sits ahead of.
  • cli/Sources/TuistCache/SelectiveTestingServicing.swift, cli/Sources/TuistAlert/AlertController.swift, cli/Sources/TuistEnvKey/EnvKey.swift: selection, the alert vocabulary, and the TUIST_TEST_* convention.
  • gradle/src/main/kotlin/dev/tuist/gradle/TuistPlugin.kt: the tuist settings extension and TestQuarantineExtension.enabled’s CI-aware default.
Draft history
Revision Status Edited
Revision 14 Edited by marek@tuist.dev
Draft
Revision 13 Edited by marek@tuist.dev
Draft
Revision 12 Edited by marek@tuist.dev
Draft
Revision 11 Edited by marek@tuist.dev
Draft
Revision 10 Edited by marek@tuist.dev
Draft
Revision 9 Edited by marek@tuist.dev
Draft
Revision 8 Edited by marek@tuist.dev
Draft
Revision 7 Edited by marek@tuist.dev
Draft
Revision 6 Edited by marek@tuist.dev
Draft
Revision 5 Edited by marek@tuist.dev
Draft
Revision 4 Edited by marek@tuist.dev
Draft
Revision 3 Edited by marek@tuist.dev
Draft
Revision 2 Edited by marek@tuist.dev
Draft
Revision 1 Edited by marek@tuist.dev
Draft
Comments

No comments yet

Comments from contributors and members will appear here.

Sign in to comment

Comments are available to authenticated users.