peptide-evidence

An Academic Study Estimated 41.6%–71.1% Failure in Grey-Market Peptides. Our 64,787-Record Reanalysis Finds a Different Pattern.

Reviewed against independent lab data · last reviewed · methodology

TitrateLab Research Desk · Evidence audit

64,787checkable laboratory records in the defined reanalysis cohort

The academic failure rate reproduces. Its market-wide interpretation does not. A larger, multi-source record shows improving purity, concentrated serious failures and a community surveillance system worth measuring.

TitrateLab aggregates third-party peptide laboratory reports and turns fragmented certificates into attributable, dated evidence. Our August 8 snapshot contains 73,048 laboratory records, including 72,186 published records. To our knowledge, it is the largest assembled dataset for analyzing grey-market peptide identity, purity, labeled quantity, endotoxin results, temporal trends and vendor-level failure patterns.

This analysis uses a defined 64,787-record cohort, extracted August 8, 2026 UTC, to evaluate the methods, external validity and conclusions of a recent academic paper released as a preprint before peer review.

The paper reported that 41.6% to 71.1% of grey-market peptide samples failed its composite quality criteria. We can reproduce that arithmetic. The study design does not support interpreting those percentages as estimates of failure prevalence across the grey market.

73,048raw laboratory records in the August 8 snapshot
72,186published records
64,787defined checkable evidence cohort
6,441samples in the academic analytic cohort
Disclosure and boundary

I operate TitrateLab. We take no vendor money or paid placement. This article analyzes published laboratory evidence; it is not medical advice, a vendor endorsement or a claim that any unapproved injectable is safe. TitrateLab does not run the assays. Independent laboratories do. We collect, normalize and analyze their reports.

What the larger record shows

Key findings
  • The academic result is a report-weighted rate from one voluntary testing program. The preprint screened 6,487 Finnrick reports and analyzed 6,441 samples. Titrate's defined cohort contains 64,787 checkable records from multiple testing programs.
  • Several safety-relevant indicators improved within the largest evidence sources. Within Freedom, below-95% purity fell from 1.49% to 0.37%, material underfill fell from 9.78% to 3.99%, and severe underfill fell from 1.30% to 0.61%. Finnrick's below-95% purity rate fell from 5.97% to 1.73%.
  • Not every indicator improved. Finnrick's identity-failure sentinel rate increased from 3.28% to 4.05%. Across label-referenced results, overfill above 10% increased from 38.73% to 49.00%.
  • Serious observed failures were concentrated. Depending on the outcome, 0.53% to 2.43% of vendors with relevant evidence accounted for half of identity failures, below-95% purity results and severe label-referenced underfills.
  • The grey market is not one quality profile. Positive-purity failure rates vary sharply by compound, and the observed records of the best-covered market segments differ materially from the serious-failure tail.
  • Community testing is a real quality-control layer. It does not provide pharmaceutical manufacturing assurance, but it creates public, sample-level evidence that communities can inspect and act on.

The paper’s composite endpoint combines failed identity, low purity, underfill and overfill. That compression hides both the improvement in safety-relevant indicators and the strong performance of the best-observed market segments. Failure is not one measurement.

Why the academic percentage is not market prevalence

The Mendias and Awan preprint screened 6,487 Finnrick reports and analyzed 6,441 samples covering fourteen peptides. The authors explicitly acknowledge that their dataset is “not a random or representative sample.” They note that vendors confident in their products may be more likely to submit samples and that poor batches may be underrepresented.

Voluntary testing can bias results in either direction. Frequent testing by quality-focused vendors can lower observed failure rates. Complaint-driven or suspicious-sample testing can increase them. Vendor-sponsored, community-sponsored and complaint-driven submissions may have different selection mechanisms.

The study:

Report count is not equivalent to market share, unique batch count or unique purchase count. The twenty most-tested companies contributed 49.3% of the paper’s analytic cohort. Vendors with more submitted reports therefore receive greater weight regardless of sales volume or batch independence.

The resulting statistic is a report-weighted failure rate within one self-selected testing program. The study design provides no basis for converting it into a market-wide prevalence estimate.

We reproduced the composite rates

The paper applied two composite standards:

Under those models, 41.6% and 71.1% of reports failed. We applied comparable rules to 9,212 current Finnrick records carrying both required measurements. The resulting failure rates were 42.2% under the first model and 70.8% under the stricter model.

The arithmetic is reproducible. The disagreement concerns sampling, generalization and the information lost when identity, purity, underfill and overfill become one binary endpoint.

A composite failure is not one kind of failure

Under the paper’s first model, a vial measuring 111% of its label fails. Under the stricter model, a result 6% above label fails. The label “compounding” does not make these intervals universal standards. The paper states that its thresholds “do not correspond to formally codified specifications for any specific product.” It describes the 98% purity cutoff as pragmatic and the 99.5% cutoff as a conservative comparator.

A result containing the identified compound at 111% of label and a result in which the labeled compound was not identified both receive the same binary failure outcome. Both indicate quality-control concerns, but their meaning differs.

372 / 54,317identity-failure sentinels among identity or purity-assessable records · 0.68%
553 / 53,945positive numerical purity results below 95% · 1.03%
233 / 34,084non-Finnrick label-referenced results at least 40% under label · 0.68%
1,460 / 34,084non-Finnrick label-referenced results more than 10% under label · 4.28%
16,184 / 34,084non-Finnrick label-referenced results more than 10% above label · 47.48%
48.23%valid label-referenced results within ±10% of labeled quantity

Finnrick quantity rows are excluded from these label-conformance figures because Finnrick’s deviation field is relative to its batch claim, not the vial label. Finnrick remains included in the purity and identity analysis and in the direct reproduction of the paper’s Finnrick-based models.

Only 48.23% of valid label-referenced quantity results were within ±10% of label. That indicates poor aggregate fill consistency relative to the interval. It does not follow that the remaining results were empty, impure or contaminated. Most quantity results outside the interval were overfills.

Overfill may provide more identified material than labeled while simultaneously indicating weak manufacturing consistency. It should be reported separately from underfill and failed identity.

Compound mix changes the pooled result

The paper pooled fourteen compounds, including retatrutide, tirzepatide, semaglutide, CJC-1295, ipamorelin, tesamorelin, BPC-157, GHK-Cu and TB-500. Our reconstruction of the contemporaneous Finnrick feed indicates that retatrutide and tirzepatide alone represented approximately 64% of the cohort. The broader metabolic and GLP-family group represented roughly 71%.

Observed failure profiles differ materially by compound.

CompoundPositive purity below 95%More than 10% under label¹More than 10% over label¹
Retatrutide56/9,669 (0.58%)91/5,333 (1.71%)2,819/5,333 (52.86%)
Tirzepatide20/7,073 (0.28%)72/4,018 (1.79%)1,715/4,018 (42.68%)
CJC-129589/873 (10.19%)30/432 (6.94%)210/432 (48.61%)
Ipamorelin12/1,116 (1.08%)54/658 (8.21%)295/658 (44.83%)
Tesamorelin44/2,283 (1.93%)39/1,549 (2.52%)868/1,549 (56.04%)

¹ Label-conformance columns exclude Finnrick because its deviation field is relative to the batch claim rather than the vial label.

CJC-1295’s observed below-95% purity rate is approximately eighteen times retatrutide’s. Its material-underfill rate is approximately four times as high. Pooling the compounds produces a cohort-level rate that does not describe any individual compound.

Retatrutide’s observed profile consists of high positive purity, comparatively low underfill and frequent overfill. This indicates a manufacturing-consistency problem, but not the same failure pattern observed for CJC-1295.

Quality changed as community testing expanded

A fixed observation period does not establish that later cohorts have the same failure profile. We compared the latest twelve months ending August 6, 2026 with the preceding twelve months.

Across the full cohort, positive-purity results below 95% fell from 2.88% to 0.67%. The source mix changed substantially, so that aggregate is not a controlled market trend.

Across non-Finnrick label-referenced results, material underfill fell from 5.47% to 4.02%, and severe underfill fell from 0.91% to 0.62%. Results within ±10% of label fell from 55.80% to 46.98%, while overfill above 10% increased from 38.73% to 49.00%.

Among 44 vendors with at least ten non-Finnrick label-referenced quantity results in both periods, the weighted material-underfill rate decreased from 4.79% to 2.56%. Twenty-four vendors improved, fifteen worsened and five were unchanged. Vendors with larger report counts contributed disproportionately to the aggregate change.

These observed changes occurred while community testing and public report sharing expanded. They do not establish causation. The data support testing as a mechanism for identifying deficient submitted samples, distributing warnings and creating pressure for correction; they do not measure buyer response or prove that testing caused the trends.

Testing-source composition can imitate a compound trend

Aggregate CJC-1295 results changed substantially between periods:

During the same interval, Freedom’s share of CJC-1295 reports increased from 11.58% to 70.19%, while Finnrick’s share decreased from 72.59% to 14.47%. Part of the apparent compound improvement was a change in which testing program supplied the evidence.

Within Finnrick alone, the below-95% purity rate decreased from 41.52% to 13.95%, while the identity-failure sentinel rate increased from 8.56% to 14.00%. The aggregate purity trend reflects both improvement within Finnrick and a major change in source composition. The aggregate identity trend is not evidence of uniform improvement.

Rankings and trend estimates should report laboratory mix, evidence source, observation period and denominator.

Serious observed failures were concentrated

The composite failure rate is not strongly concentrated because overfill is common. Identity failure, low purity and severe underfill show substantially greater vendor concentration.

Identity-failure sentinels
19 / 3,583
Positive purity below 95%
57 / 3,581
At least 40% under label
69 / 2,840
More than 10% under label
140 / 2,840

Each row shows how many attributed vendors account for half of that observed failure mode:

Identity-failure detection is source-dependent: 369 of 372 identity-failure sentinels come from Finnrick, while other sources encode identity outcomes differently. These concentration figures describe the observed corpus, not latent failures that were never tested or encoded.

The results do not establish fraud, intent or a particular sourcing channel. They show substantial heterogeneity among attributed vendors. A market-wide binary rate does not preserve that distribution.

Identity failure and numerical purity are separate metrics

An identity failure encoded as a numerical zero should not be included in mean positive purity. A vendor may have successfully identified samples averaging 99.9% purity and separate samples in which the labeled compound was not identified.

Encoding those identity failures as zero-percent purity could reduce a reported average to approximately 30%. That calculation corrupts both signals. The appropriate representation is:

A high mean positive purity does not offset an identity failure. An identity failure does not change the measured purity of successfully identified samples.

Endotoxin coverage remains inadequate

The academic study had endotoxin data for 243 of 6,441 analyzed samples, less than 4%. It reported 74 with no detectable endotoxin, 133 below the lower quantification limit and 36 with measurable endotoxin between 0.5 and 40 EU/mL.

The 36 measurable results constitute 14.8% of the endotoxin-tested subset. This is 36 of 243 endotoxin-tested reports, not 15% of the full cohort and not a population estimate for the grey market.

Measurable endotoxin is not necessarily an out-of-specification result. Interpretation requires an applicable limit based on route and maximum dose. Grouping results from 0.5 to 40 EU/mL does not establish that all 36 would exceed the same product-specific limit. The FDA explains the dose-based calculation, and no single cross-product threshold resolves it.

Titrate currently has 4,846 nonblank, non-n/a endotoxin-result fields in the conservative cohort. They contain heterogeneous categorical labels and numerical units, so they are not pooled into a single cross-product failure rate here.

The larger denominator still does not establish a market-wide contamination rate. Testing is voluntary, selection is nonrandom and laboratory reporting conventions differ. The supported conclusion is that endotoxin-testing coverage remains inadequate and that purity cannot substitute for endotoxin testing.

Manufacturing assurance and public evidence are different controls

Approved manufacturers must test each batch for conformity with final specifications before release and use documented sampling and testing plans. That requirement is codified in 21 CFR 211.165. Retail pharmacy customers generally are not provided the manufacturer’s lot-release assay for the product they receive.

Organized community testing can provide public certificates documenting identity, purity, tested quantity, date, laboratory and batch. That creates direct public visibility into a submitted sample. It is not equivalent to pharmaceutical manufacturing control.

A community test usually establishes what one submitted sample contained. It does not prove every vial in the batch is identical, and it does not replace sterility assurance, stability testing or a recall system. Approved manufacturing provides stronger mandatory quality assurance. Community testing can provide greater public access to sample-level analytical evidence.

Our companion analysis, Pharma’s Quality Bar Is Lower Than You Think, examines the product-specific and regulatory limits behind dose, purity and endotoxin comparisons. The FDA’s allowable excess-volume guidance also makes clear why overfill is a manufacturing-consistency signal rather than free material with no downside.

Two systems, different evidence

Pharmaceutical manufacturing: mandatory batch release, validated processes, stability programs and formal recall authority.

Community testing: voluntary, public, sample-level certificates that can reveal deficient material, recurring failure patterns and changes over time.

Neither should be described as the other. The value of community testing is visibility and feedback, not a claim of pharmaceutical equivalence.

Conclusion: public testing changed what the market can know

The growth of community-funded testing and public report sharing has created a decentralized, sample-level quality-surveillance system in a market that previously depended largely on vendor claims.

The strongest evidence of improvement is source-specific. Within Freedom, below-95% purity fell from 1.49% to 0.37%, material underfill fell from 9.78% to 3.99%, and severe underfill fell from 1.30% to 0.61%. Within Finnrick, below-95% purity fell from 5.97% to 1.73%. Among 44 vendors with sufficient label-referenced quantity data in both periods, material underfill fell from 4.79% to 2.56%.

These are observational changes, not proof that community testing caused them. The proposed mechanism is testable: a sample is analyzed, the certificate is distributed, potential batch problems and recurring vendor failures become visible, and market participants can respond. This analysis does not measure buyer response or vendor causality.

Titrate supports that process by aggregating reports, normalizing their fields, preserving source and date information and making failure patterns searchable.

The same evidence shows that the grey market is not analytically uniform. A small fraction of attributed vendors produced half of the serious observed failures, while the best-observed segments repeatedly produced highly pure, correctly identified material. Pooling those groups into one market-wide failure rate removes the information people need to distinguish them.

The preprint’s abstract and conclusions extend beyond what its nonrandom sampling frame can establish. The same type of testing evidence also shows source-specific improvement, concentrated serious failures and high-performing market segments.

Approved manufacturers retain stronger mandatory quality systems. The grey market has developed a different and valuable control: public laboratory evidence that communities can inspect and act on. Its coverage is increasing. This analysis does not estimate its causal effect on market quality.

Titrate’s position is favorable to this evidence-driven grey market because the record supports differentiation rather than blanket condemnation. Community testing has converted quality from an unverifiable claim into a measurable and contestable record. Titrate’s role is to make that feedback loop faster, broader and more technically reliable.


Sources

Titrate cohort definition: Snapshot extracted August 8, 2026 UTC. Eligible records were published, had identified attribution to a visible nonmerged manufacturer, and carried a verification URL or an approved public evidence-source classification. Positive purity required purity_pct > 0; exact zero was counted separately as an identity-failure sentinel. Label-conformance analysis required -99 < deviation_vs_label_pct <= 200 and excluded Finnrick. “More than 10% under” used < -10; “at least 40% under” used <= -40. Compound rows excluded blends. Vendor counts used canonical manufacturer IDs.

This is an analysis of published laboratory evidence, not medical advice, a clinical-safety determination or a vendor endorsement.

Vendor and manufacturer names are used descriptively to identify parties in the documentary record; inclusion is not endorsement. Think a passage misrepresents the record? [email protected].
NOW OPEN

TitrateLab is open.

Create a free account to read the full corpus and run the COA tools. Go Pro for $13/mo, or $33 once — lifetime.

Pro · $13/mo or $33 lifetime — see what's included. Free accounts read the full corpus and run the COA tools.