Training to Failure vs Two Reps in Reserve: Reading the Single-Set Trial

The useful decision is not “failure or never failure.” It is which stopping point fits the job:
- For strength or local muscular endurance in the tested single-set format, estimated two RIR performed similarly to failure.
- For muscle growth, several measures leaned toward failure, but differences were generally modest.
- For a multi-set routine, this trial cannot tell us whether the extra effort on set one improves the session or reduces later output.
That framework keeps the eight-week randomized trial in its actual setting rather than turning one comparison into a universal rule.
The single-set design
The randomized controlled trial assigned 42 young, resistance-trained men and women to one of two parallel groups. Both groups used the same nine exercises for all major muscle groups, trained twice per week, and continued for eight weeks.
The failure group took each set to momentary failure. The 2-RIR group stopped when participants estimated that they could complete two more repetitions. This made effort and stopping point the primary contrast while keeping the broad program structure aligned.
Researchers measured muscle thickness at the biceps brachii, triceps brachii, and quadriceps femoris. They also tested muscular strength, power, local muscular endurance, and the participants' ability to estimate repetitions in reserve in the bench press and squat.
Because participants were already resistance trained, the experiment speaks more directly to experienced lifters than a novice intervention would. Yet “young and trained” is still a population boundary. The results do not automatically cover older adults, beginners, rehabilitation settings, or competitive sport programs with much higher weekly workloads.
Outcome scoreboard
| Outcome | Trial reading |
|---|---|
| Strength | Similar gains between failure and 2-RIR |
| Local muscular endurance | Similar gains |
| Muscle thickness | Several measures tended to favor failure; absolute differences were generally modest |
| Countermovement jump | Numerically favored failure, without clear statistical support |
The scoreboard is more informative than naming a winner.
Both groups made appreciable gains in most measured outcomes. Strength improvements were similar between failure and 2-RIR, as were changes in local muscular endurance. That directly challenges the idea that a set must reach failure to improve either quality under this program.
Hypertrophy was less tidy. Several measures tended to favor failure, but the paper's report describes the absolute differences between conditions as generally modest. “Tended to favor” is not the same as a consistent, statistically decisive victory across every site. It identifies a possible tradeoff worth studying, not a license to announce that two RIR leaves muscle on the table in every routine.
Countermovement-jump changes also numerically favored failure. The statistical evidence did not clearly support either the null or alternative hypothesis, so the direction should not be sold as proof. Numerical separation can be interesting while remaining too uncertain for a firm conclusion.
The combined picture is therefore outcome-specific: similar strength and endurance, modest hypertrophy tendencies toward failure, and an uncertain jump result. Compressing those findings into a single winner loses the information lifters actually need.
The hidden variable: estimating RIR
The 2-RIR condition depended on participants estimating how many repetitions remained. That estimate cannot be observed directly in the moment; its accuracy becomes known only if a lifter continues and tests the limit, which would defeat the point of stopping early during every training set.
Participants were more accurate at estimating RIR in the bench press than in the squat. Accuracy improved during the intervention, particularly for the bench press. This suggests that RIR judgment can become more calibrated with practice, but also that accuracy varies by exercise.
That exercise difference matters in the real gym. An honest two-RIR estimate on a familiar machine movement may not have the same reliability as an estimate during a technically demanding squat. Load, rep range, discomfort tolerance, and exercise familiarity can all affect how the effort feels, even though this trial was not designed to isolate every one of those factors.
The result also cautions against defining the groups by intention alone. If a person assigned to two RIR routinely stops five repetitions early, the program no longer resembles the tested condition. Conversely, a purported two-RIR set that repeatedly ends at failure belongs functionally to the other side of the comparison.
Where this evidence applies
A single-set routine can be time efficient, and the authors concluded that it can promote adaptations even among trained people moving from higher-volume programs. That is a useful result for lifters whose main constraint is time.
It is not evidence that one set is the best volume for muscle gain. The study held volume low to compare two stopping strategies within that format. A multi-set program introduces accumulated fatigue: taking the first set to failure may reduce repetitions or load on later sets, while stopping short may preserve output. The new volume dose-response analysis addresses weekly set exposure from a different evidence base and should not be collapsed into this eight-week trial.
Duration creates another limit. Eight weeks can reveal measurable changes, but it cannot tell us whether repeated failure work remains equally tolerable across a long training year. It also cannot settle how fatigue interacts with sport practice, dieting, or a dense exercise schedule.
For strength and local muscular endurance under this program, the trial gives no reason to insist on failure: two RIR produced similar adaptations. A lifter who values technique consistency or wants to limit acute fatigue has evidence that stopping short can still work when the estimate is reasonably accurate.
For hypertrophy, the modest tendencies toward failure are worth acknowledging without upgrading them into certainty. When a routine contains only one set per exercise, pushing that set harder may be a reasonable way to ensure a strong stimulus. In a multi-set program, the costs and benefits can shift because subsequent work must still be completed.
Occasional calibration can help RIR estimates. That might mean observing whether load and repetitions progress as expected, using consistent exercises, and reviewing sets where technique remains stable. It does not require failing every set. The trial itself showed estimation improving with practice, especially on the bench press.
A future multi-set comparison would need to measure the first set's effort alongside the repetitions, load, and fatigue carried through the rest of the session. Until then, using this single-set result to dictate every set in a high-volume block answers a question the study never asked.
More Training
TrainingVelocity Loss in Concurrent Training: The 0%, 15%, and 40% Study
An eight-week trial tested 0%, 15%, and 40% squat velocity-loss thresholds before running, revealing muscle, strength, and endurance tradeoffs.
TrainingMenstrual-Cycle Phase Training: What the New Strength Study Found
A within-participant trial found no advantage to concentrating resistance-training volume in the follicular or luteal phase over balanced training.
TrainingAccentuated Eccentric Training Once, Twice, or Three Times Weekly
A 12-week squat study compared accentuated eccentric training one, two, or three days per week in trained athletes. Most outcomes were similar.
