Key Takeaways
- •The same trial reports different numbers under different estimands. STEP 1 published −14.9% and −16.9% for the same patients: one figure counts everyone regardless of what they did, the other estimates the effect if the regimen was followed as intended.
- •Comparing across trials compounds the error. The naive semaglutide-versus-tirzepatide gap taken from separate trials is 7.6 points. The head-to-head trial measured 6.5.
- •A head-to-head trial exists, and it is the only comparison that settles anything. SURMOUNT-5 randomised 751 participants and reported 20.2% against 13.7%.
- •Duration, population and comparator all move the number independently of how well a compound works.
- •A percentage without a trial name, a duration and an estimand is not a result, it is a marketing figure that happens to be numeric.
- •None of this transfers to a research vial. Trials are run with pharmaceutical-grade material released against a specification.
Tirzepatide reduced body weight by 22.5%. It also reduced body weight by 20.9%. It also reduced body weight by 20.2%.
All three figures are correct. All three come from properly conducted, peer-reviewed Phase 3 trials. None of them contradicts the others. The first two are the same trial, the same patients, the same data, analysed under two different pre-specified rules. The third is a different trial against a different comparator.
Anyone quoting one of those numbers at you has chosen it, and the range spans 2.3 percentage points before a single question about study design has been asked. Once you start comparing across trials, the manufactured gap gets considerably larger than the real one, and this article shows exactly how much larger using published figures you can check.
This matters here because research compounds are sold on the strength of a clinical literature that describes a different product entirely. Understanding what a trial percentage means is the difference between reading that literature and being marketed at with it. All Volta material is supplied for laboratory research use only.
The estimand: the choice that moves the number without touching the data
An estimand is the precise definition of what a trial is estimating. It sounds like statistical housekeeping. It is worth one to two percentage points on exactly the headline figures people quote.
The problem it solves is real. In a long trial, people stop taking the drug, drop out, or start something else. What should the analysis do with them? Two defensible answers:
| Estimand | The question it answers | Effect on the number |
|---|---|---|
| Treatment policy (also: treatment regimen) | What happens to everyone assigned to this arm, regardless of whether they stayed on it? | Lower. Discontinuers are counted, and they typically do worse |
| Trial product (also: efficacy) | What happens if the regimen is followed as intended? | Higher. Estimates the effect of the drug taken as designed |
Neither is a trick. The first reflects real-world use; the second isolates the drug's effect. Regulators generally lead with the first. Marketing materials generally lead with the second.
Both appear in STEP 1, the pivotal semaglutide trial published in the New England Journal of Medicine in 2021: 1,961 adults, randomised 2:1, 68 weeks.
| STEP 1, week 68 | Semaglutide | Placebo | Difference |
|---|---|---|---|
| Treatment policy estimand | −14.9% | −2.4% | −12.4 points |
| Trial product estimand | −16.9% | −2.4% | −14.4 points |
Same trial. Same participants. Two percentage points apart.
SURMOUNT-1, the tirzepatide equivalent at 72 weeks, does the same thing: 20.9% under the treatment-regimen estimand and 22.5% under the efficacy estimand at the highest strength studied, with the published range running from 16.0% at the lowest strength to 22.5% at the highest.
So before any comparison begins, each compound already has two legitimate headline numbers and a range across strengths. The single figure that reaches a product page is a selection.
What else differs between two trials
Estimand is one axis. There are at least four more, and every one of them moves a percentage independently of how well the molecule works.
| Axis | Why it changes the number |
|---|---|
| Duration | STEP 1 ran 68 weeks, SURMOUNT-1 ran 72. Curves are still descending at both points, so a longer trial reports a larger reduction for the same drug |
| Population | Participants with type 2 diabetes consistently show smaller body-weight reductions than those without. A trial that enrols one, both, or neither is not measuring the same thing |
| Comparator | Against placebo the number is a difference from near-zero. Against an active drug it is a difference from something that works |
| Background care | Both trials ran alongside lifestyle intervention, and the intensity of that support differs between programmes |
| Strength studied | A range gets reported. A single number gets quoted |
None of this is hidden. It is in the methods section of every paper. It is simply absent from every summary.
The arithmetic, done properly
Here is the demonstration, using only published figures.
The naive comparison, which is the one in circulation: take tirzepatide's best number from SURMOUNT-1 and semaglutide's headline from STEP 1.
22.5% − 14.9% = 7.6 points
That comparison mixes estimands. 22.5% is the efficacy estimand; 14.9% is the treatment policy estimand. It compares a drug taken as intended against a drug taken as people actually take it.
Match the estimands and compare like with like:
| Comparison | Tirzepatide | Semaglutide | Gap |
|---|---|---|---|
| Naive, mixed estimands | 22.5% | 14.9% | 7.6 points |
| Both "as intended" | 22.5% | 16.9% | 5.6 points |
| Both "regardless of adherence" | 20.9% | 14.9% | 6.0 points |
Three different answers from the same four published numbers, spanning two full percentage points, purely from which pairing you choose.
Now the trial that actually settles it. SURMOUNT-5 randomised 751 participants 1:1 to tirzepatide or semaglutide, head to head, over 72 weeks:
Tirzepatide 20.2%, semaglutide 13.7%. Gap: 6.5 points. In absolute terms, 22.8 kg against 15.0 kg.
The naive cross-trial comparison overstated the real gap by more than a full percentage point. The like-for-like cross-trial comparisons bracketed it, one above and one below. Only the randomised head-to-head measured it.
The lesson generalises past this example: cross-trial arithmetic produces a number, and the number is not an estimate of anything in particular. It can overstate or understate. Its error is not predictable, which is precisely why it is not a substitute for a comparison that was actually run.
In Stock and Shipping from British Columbia
Every vial below ships domestically within Canada with a batch-specific Certificate of Analysis. Supplied for laboratory research use only.
>99% HPLC purity standard. Read the Certificates of Analysis.
See the full Canadian catalogueA checklist for any percentage
Six questions. If a source cannot answer the first three, the figure is not usable.
- Which trial? A name or a registry number. "Studies show" is not a source.
- Over how long? Curves that have not plateaued make duration a lever.
- Which estimand? If unstated, assume the more flattering one was chosen, because it usually was.
- In whom? Diabetes status, baseline characteristics, and the entry criteria.
- Against what? Placebo or an active comparator, and at what strength.
- Primary or secondary endpoint? A missed primary with an encouraging secondary is a missed trial.
The evidence grade tool scores a claim against this shape, the claim checker tests marketing wording, and the research literacy guide covers trial design in more depth.
Why none of it transfers to a vial
This is the part that cuts against a supplier's interest, so it belongs in plain sight.
Every figure above was produced with material manufactured to pharmaceutical standards, characterised exhaustively, released against a specification, and administered under supervision in a controlled protocol. The percentage is a property of that whole system, not of a molecule in the abstract.
A research-grade vial shares the sequence and none of the rest. So a trial result is evidence that a molecule does something under specified conditions. It is not evidence about the contents of any particular vial, and a page that places a clinical percentage next to a product is inviting a transfer that the data does not support.
Which is the same substitution described in what the human evidence actually shows, and the reason the identity and supplier verification questions are answered separately from the literature question.
Frequently Asked Questions
What is an estimand, in plain terms?
The exact definition of what a trial is measuring, including how it handles participants who stop taking the drug. A treatment policy estimand counts everyone as assigned, so discontinuations drag the number down. A trial product or efficacy estimand estimates the effect if the regimen was followed as intended, which gives a higher figure from identical data.
Which estimand should I trust?
Neither exclusively. They answer different questions and a good paper reports both. If you want to know what a drug does when taken as designed, the trial product estimand is the right one. If you want to know what happens to a population offered it, the treatment policy estimand is. The error is not choosing one, it is comparing one drug's efficacy estimand against another's treatment policy estimand.
Why can't I just compare percentages from two trials?
Because the two trials differ in duration, population, comparator, background care, strength and estimand, and each of those moves the number independently of the drug. The comparison in this article was off by more than a percentage point against the head-to-head trial, and the direction of that error was not predictable in advance.
Does a bigger percentage mean a better compound?
Not on its own. It means a larger measured effect on one endpoint, in one population, over one duration, under one analysis rule. Adverse events, discontinuation rates, durability after stopping and the size of the trial all belong in the same judgement, and none of them appear in a single headline figure.
Do these trial results apply to research-grade material?
No. Trials use pharmaceutical-grade material with controlled identity, content, impurity profile and sterility. Research-grade material has the same sequence and none of that assurance, which is why its own documentation has to be assessed separately. The trial tells you about the molecule; the certificate tells you about the vial.
How do I find the estimand in a paper?
It is in the statistical analysis section of the methods, and in well-reported trials it appears in the results tables as separate rows or columns. In STEP 1 both estimands are reported explicitly with their confidence intervals. If a summary quotes a number without saying which one it is, go back to the primary paper rather than trusting the summary.




























