In short: Start by identifying the study type: cell (in vitro), animal (in vivo) or human. Evidence levels set by the National Health and Medical Research Council (NHMRC) rank human study designs from level I, a systematic review of randomised controlled trials, to level IV, case studies. Cell and animal work sits outside that scale as preclinical evidence.
Peptide papers range from experiments on cells in a dish to randomised trials in people, and the type of study limits what its results can show. This guide explains how to place a study on the evidence scale used in Australia, what to check in its methods and results, and how to confirm it has not been retracted. It is the starting point for our articles on reading the research, including why most peptide research is preclinical and the anti-doping status of research peptides. If peptide chemistry is new to you, begin with what is a peptide?.
What types of study appear in peptide research?
Three broad kinds, plus the reviews that summarise them: cell studies, animal studies and studies in people.
| Study type | What it is | What it can and cannot show |
|---|---|---|
| Cell study (in vitro) | Cells or tissue grown outside a living organism | How cells respond under controlled conditions; not how a whole organism responds |
| Animal study (in vivo) | Experiments in living animals, most often rodents | Responses in a whole organism of another species; not responses in people |
| Case study or case series | Measurements in one or more people, with no comparison group | That something was observed; not what caused it |
| Cohort or case-control study | Observational comparison of groups of people | Associations, which a third factor linked to both groups (confounding) may explain |
| Randomised controlled trial | People allocated by chance to the intervention being tested or a comparison | The highest level for a single study in the NHMRC scale 1 |
| Narrative review | An expert summary of selected papers | Context; the selection may be partial |
| Systematic review | A pre-specified search and appraisal of all eligible studies, sometimes with a pooled analysis (meta-analysis) | Only as much as the included studies can show |
What are the NHMRC levels of evidence?
For questions about interventions, the National Health and Medical Research Council (NHMRC) hierarchy runs from level I, a systematic review of randomised trials, down to level IV, case studies. The table uses the wording of the NHMRC levels as reproduced in a 2009 publication of NHMRC’s National Institute of Clinical Studies 1.
| NHMRC level | Study design (intervention questions) | In plain terms |
|---|---|---|
| I | A systematic review of level II studies | A systematic review of randomised controlled trials |
| II | A randomised controlled trial | Participants allocated to groups by chance |
| III-1 | A pseudorandomised controlled trial | Allocation by a rule such as alternation, not true chance |
| III-2 | A comparative study with concurrent controls: non-randomised experimental trial, cohort study, case-control study, or interrupted time series with a control group | Comparison groups studied at the same time, without randomisation |
| III-3 | A comparative study without concurrent controls: historical control study, two or more single-arm studies, or interrupted time series without a parallel control group | Comparison against past data or separate studies |
| IV | Case studies with either post-test or pre-test/post-test outcomes | Measurements after, or before and after, with no comparison group |
NHMRC has since updated its approach: its Guidelines for Guidelines handbook states that it replaces the NHMRC additional levels of evidence and grades for recommendations for developers of guidelines (2009) 2. NHMRC’s guidance on assessing certainty of evidence describes GRADE (Grading of Recommendations Assessment, Development and Evaluation) as now considered best practice by many international guideline developers 3. GRADE rates the certainty of a body of evidence, outcome by outcome, as high, moderate, low or very low. Randomised trials start high and observational studies low, and ratings move for risk of bias, imprecision, inconsistency, indirectness and publication bias 3.
Illustration: evidence hierarchies rank study designs, with systematic reviews of randomised trials at the top.
Where do cell and animal studies fit?
Outside the NHMRC levels, which describe clinical study designs, so cell and animal work is preclinical evidence that the scale does not grade. It can show a mechanism or a measurable response in a model; on its own, it cannot show what happens in people.
Researchers have adapted GRADE for animal evidence. Hooijmans et al. (2018) set out how GRADE could rate the certainty of evidence from preclinical animal studies of therapeutic interventions, and flagged open problems such as consistency across species and using indirectness to predict translation to humans 4. NHMRC’s certainty guidance lists this work among the applications and variations of GRADE 3.
Translation from animals to people is uneven. In a 2024 umbrella review (a review of systematic reviews) covering 122 articles and 367 therapeutic interventions, Ineichen et al. reported that 50% of interventions progressed from animal studies to any human study, 40% to randomised controlled trials and 5% to regulatory approval, with median times of 5, 7 and 10 years 5. They also reported 86% concordance between positive results in animal and clinical studies 5.
Illustration: cell studies are run on cultured cells in flasks and plates, outside a living organism.
What should you check in the methods section?
The model, the comparison group, how subjects were allocated, whether outcomes were assessed blind, and how the sample size was chosen.
- The model. Which cell line, animal species or group of people was studied. For cell work, check whether the cells’ identity was confirmed: Horbach and Halffman (2017) estimated that 32,755 articles reported research on misidentified cells, cited in turn by about half a million other papers 6.
- The comparison. A control group that differs only in the factor being tested.
- Randomisation and blinding. In a random sample of 146 publications describing in vivo research, Macleod et al. (2015) found that 20% of those where randomisation was appropriate reported it, 3% reported blinded assessment of outcome (assessors not knowing which group each subject was in), and none reported a sample size calculation 7.
- Sample size. How many cells, animals or people per group, and how that number was decided.
- Outcomes. What was measured, and whether the main outcome was named in advance. A change in a laboratory marker is not the same as a change in function.
- The material. Whether the paper reports the source, purity and identity of the peptide used. Our guides to HPLC purity and mass spectrometry identity testing explain those checks.
Reporting checklists set out what a paper should state. ARRIVE 2.0 covers animal research, and its “Essential 10” items include study design, sample size, randomisation, blinding, outcome measures and statistical methods 8. CONSORT 2025 is a 30-item checklist for randomised trials 9, and PRISMA 2020 is a 27-item checklist for systematic reviews 10.
How should you read the results?
Look at the size of an effect and its uncertainty, not only whether a P value fell below 0.05. A P value indicates how unusual the data would be if the tested hypothesis, usually “no difference”, and every other assumption of the analysis were true. A confidence interval gives the range of effect sizes most compatible with the data under the same assumptions.
Greenland et al. (2016) listed 25 common misinterpretations of P values, confidence intervals and statistical power (the chance that a study detects an effect of a given size) 11. Among them: that a P value is the probability that the tested hypothesis is true, that statistical significance means an important effect has been found, and that a non-significant result means the effect is small. They noted that in small studies even large effects can fail to reach significance 11.
Who did the research, and has anyone repeated it?
Check the authors, their funding and conflict-of-interest statements, and whether independent groups have reported similar results. In Macleod et al.’s sample, 10% of in vivo publications included a conflict-of-interest statement 7.
Replication matters. The Reproducibility Project: Cancer Biology repeated 50 experiments from 23 high-impact papers; for positive effects, the median effect size in the replications was 85% smaller than in the originals, and 46% of replications succeeded when positive and null effects were combined 12. When most papers on a compound come from a single laboratory, independent replication is the missing piece, and our compound reviews, such as the BPC-157 research review, say so where it applies.
How can you find out whether human trials exist?
Search the trial registries: ClinicalTrials.gov and, for Australia and New Zealand, the ANZCTR. A registration records what a trial planned to measure, which can be compared with what was later published.
A systematic review can also show the balance of evidence. In a 2025 systematic review of the BPC-157 literature relevant to orthopaedics, Vasireddi et al. included 36 studies, of which 35 were preclinical and 1 was a clinical study 13. Because NHMRC level I is a systematic review of randomised controlled trials 1, a systematic review made up mostly of animal studies does not reach level I, however carefully it was done.
How does this blog label the evidence it cites?
Every finding we report names the study type and model, in the past tense, and every reference ends with a bracketed label such as [rodent study], [cell study], [systematic review] or [phase 2 randomised controlled trial]. We use NHMRC’s level names because they describe a single study’s design in a few words, and every cited paper is checked for retraction notices before publication. Our editorial standards explain how we source, write and correct content. Evidence and regulatory status are separate questions; for anti-doping status by compound, see research peptides and anti-doping.
Frequently asked questions
What is the difference between in vitro and in vivo?
In vitro means “in glass”: experiments on cells, tissues or molecules outside a living organism, such as cells grown in a flask. In vivo means “within the living”: experiments in whole organisms, which in peptide research are usually animals. Both count as preclinical when no people are involved, and neither appears in NHMRC’s evidence levels, which describe clinical study designs 1. Our peptide glossary defines related terms.
What is the highest level of evidence in the NHMRC hierarchy?
Level I: a systematic review of level II studies, which for intervention questions are randomised controlled trials 1. NHMRC’s Guidelines for Guidelines handbook has since replaced its 2009 levels document 2, and NHMRC’s guidance describes GRADE as now considered best practice by many international guideline developers for rating how certain a body of evidence is, from high to very low 3.
Is a review article the same as a systematic review?
No. A narrative review summarises papers its authors chose, while a systematic review follows a pre-specified search and selection method so that eligible studies are found and appraised consistently. The PRISMA 2020 statement sets out a 27-item checklist for reporting systematic reviews 10. A systematic review is only as strong as the studies it includes: one built mostly on animal work is still preclinical evidence.
How can I tell whether a paper has been retracted?
Search its title or DOI in the Retraction Watch Database, and check its record in Europe PMC or PubMed, where retraction notices and errata are linked to the original article. Problems short of retraction matter too: Horbach and Halffman (2017) estimated that 32,755 articles reported research on misidentified cells 6.
References
- National Health and Medical Research Council, National Institute of Clinical Studies. Appendix F: Levels of evidence and recommendation grading. In: Emergency department stroke and transient ischaemic attack care bundle: information and implementation package (2009). https://www.nhmrc.gov.au/sites/default/files/images/appendix-f-levels-of-evidence.pdf. Accessed 28 September 2026. [evidence hierarchy]
- National Health and Medical Research Council. Guidelines for Guidelines: NHMRC and Guidelines. https://www.nhmrc.gov.au/guidelinesforguidelines/nhmrc-and-guidelines. Accessed 28 September 2026. [guideline development handbook]
- National Health and Medical Research Council. Guidelines for Guidelines: Assessing certainty of evidence. https://www.nhmrc.gov.au/guidelinesforguidelines/develop/assessing-certainty-evidence. Accessed 28 September 2026. [guideline development handbook]
- Hooijmans CR, de Vries RBM, Ritskes-Hoitinga M, et al. Facilitating healthcare decisions by assessing the certainty in the evidence from preclinical animal studies. PLoS One. 2018;13(1):e0187271. doi:10.1371/journal.pone.0187271. PMID: 29324741. PMCID: PMC5764235. [methods paper]
- Ineichen BV, Furrer E, Grüninger SL, et al. Analysis of animal-to-human translation shows that only 5% of animal-tested therapeutic interventions obtain regulatory approval for human applications. PLoS Biol. 2024;22(6):e3002667. doi:10.1371/journal.pbio.3002667. PMID: 38870090. PMCID: PMC11175415. [umbrella review]
- Horbach SPJM, Halffman W. The ghosts of HeLa: how cell line misidentification contaminates the scientific literature. PLoS One. 2017;12(10):e0186281. doi:10.1371/journal.pone.0186281. PMID: 29023500. PMCID: PMC5638414. [literature analysis]
- Macleod MR, Lawson McLean A, Kyriakopoulou A, et al. Risk of bias in reports of in vivo research: a focus for improvement. PLoS Biol. 2015;13(10):e1002273. doi:10.1371/journal.pbio.1002273. PMID: 26460723. PMCID: PMC4603955. [meta-research study]
- Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0: updated guidelines for reporting animal research. PLoS Biol. 2020;18(7):e3000410. doi:10.1371/journal.pbio.3000410. PMID: 32663219. PMCID: PMC7360023. [reporting guideline]
- Hopewell S, Chan AW, Collins GS, et al. CONSORT 2025 statement: updated guideline for reporting randomised trials. PLoS Med. 2025;22(4):e1004587. doi:10.1371/journal.pmed.1004587. PMID: 40228477. PMCID: PMC11996237. [reporting guideline]
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71. PMID: 33782057. PMCID: PMC8005924. [reporting guideline]
- Greenland S, Senn SJ, Rothman KJ, et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol. 2016;31(4):337-350. doi:10.1007/s10654-016-0149-3. PMID: 27209009. PMCID: PMC4877414. [methods review]
- Errington TM, Mathur M, Soderberg CK, et al. Investigating the replicability of preclinical cancer biology. eLife. 2021;10:e71601. doi:10.7554/eLife.71601. PMID: 34874005. PMCID: PMC8651293. [replication project]
- Vasireddi N, Hahamyan H, Salata MJ, et al. Emerging use of BPC-157 in orthopaedic sports medicine: a systematic review. HSS J. 2025;21(4):485-495. doi:10.1177/15563316251355551. PMID: 40756949. PMCID: PMC12313605. [systematic review]
Check the lab report
Every published certificate of analysis names its lab, lot and test date. For laboratory research use only.
