Part 1|3: Why Does the Research Keep Saying It Doesn't Work?
Understanding the Problem with Meta-Analysis
"If the research says it doesn't work — check who designed the research."
Pastor Kevin Cicchino
Category: Studies · Wellness Foundations
Tags: meta-analysis, research literacy, human movement, fitness research, rehabilitation, evidence-based practice, training principles, workshop resources, wellness foundations, health foundations
Part 1 of 3
You've probably experienced this.
You're dealing with a physical issue — chronic pain, a movement limitation, a training plateau — and someone hands you a study, or you find one online, and it says something like:
"No significant difference was found between groups."
"The intervention showed no measurable effect."
"Evidence is insufficient to support this approach."
And you walk away thinking — so nothing works? Or worse, you start doubting something that was actually helping you, because a research paper said it shouldn't be.
Here's what most people don't know: the problem is often not the intervention. It's the research method used to evaluate it.
Specifically, it's a method called a meta-analysis¹ — and when it's applied incorrectly, which happens more than most people realize, it doesn't clarify the science. It distorts it.
What Is a Meta-Analysis and Why Do We Use It?
A meta-analysis is a research method that synthesizes multiple existing studies on the same topic and pools their data to produce a single combined conclusion. The idea is sound — if one study has 50 participants and another has 60, combining them gives you a larger sample size and, theoretically, more reliable results.
Done well, a meta-analysis can be one of the most powerful tools in research. When done poorly, it produces conclusions that appear authoritative on the surface but rest on a fundamentally flawed foundation.
And in the fields of fitness, rehabilitation, and human performance, poorly constructed meta-analyses have become increasingly common — leading to widespread conclusions that interventions don't work when the real issue is that the studies were never sufficiently compatible to be combined in the first place.
The Five Core Problems
1. Mixing Studies That Should Never Be Mixed
The most common and damaging error in meta-analysis is pooling studies together that have fundamentally different populations, interventions, and outcomes — and treating the combined result as meaningful.
Imagine someone trying to determine whether "exercise" improves health outcomes. They combine studies on strength training, long-distance running, yoga, and high-intensity interval training² — different populations, different goals, different measures of success — and average the results together.
The conclusion? "Exercise shows mixed results."
That conclusion is not wrong because exercise doesn't work. It's wrong because the question was too broad, the studies were incompatible, and combining them erased the very distinctions that make each approach meaningful.
This is called heterogeneity³ — when the studies being combined are too different from each other to produce a valid pooled result. A well-designed meta-analysis carefully accounts for this. A poorly designed one ignores it entirely.
2. Treating a P-Value⁴ Like It's the Whole Story
Most research conclusions hinge on something called a p-value⁴ — a statistical measure that indicates whether a result is likely due to chance. The conventional threshold is p < 0.05, meaning there's less than a 5% probability that the result occurred by chance.
The problem is that p-values have come to be treated as a binary pass/fail test. Either a result crosses the threshold and is called "significant," or it doesn't and gets dismissed.
What this ignores is effect size⁵ — the actual magnitude of the change that occurred. A study can show a genuinely meaningful improvement in pain, mobility, or strength that doesn't reach statistical significance simply because the sample size was too small. That doesn't mean nothing happened. It means the study wasn't large enough to detect it with statistical confidence.
Dismissing a result because p > 0.05 while ignoring a consistent, meaningful effect size⁵ is not careful science. It's an oversimplification that leads to real interventions being labeled ineffective — not because they failed, but because the measurement tool wasn't applied correctly.
3. Confusing "We Didn't Prove It" With "It Doesn't Work"
This is one of the most important distinctions in all of research literacy — and one of the most commonly misunderstood.
A null result⁶ — a study that finds no significant difference — does not prove that an intervention is ineffective. It proves that this particular study, with this particular design, this particular sample size, and these particular measurement tools, did not detect a significant effect.
Those are very different statements.
Studies fail to detect real effects for many reasons: the sample size was too small to achieve statistical power⁷, the measured outcome wasn't the right one, the intervention wasn't applied long enough, or the measurement tools introduced error. None of these failures is evidence that the intervention doesn't work.
When a meta-analysis pools multiple null results and concludes that an intervention is ineffective without examining why those studies produced null results, it commits a logical error. It is treating the absence of proof as proof of absence — a fundamental mistake that has led to many genuinely effective approaches being dismissed in the literature.
4. Averaging Away the Truth
Even when individual studies show a consistent positive trend — meaning most of them point in the same direction — a meta-analysis can obscure that trend entirely through the process of averaging.
Here is a simplified example. Imagine ten studies on a particular rehabilitation technique. Eight of them show moderate improvement. Two show no change. A meta-analysis pools all ten, averages the effect sizes⁵, and the two null results pull the average down enough that the conclusion reads: "insufficient evidence to support effectiveness."
But the actual pattern — eight out of ten studies trending positive — is meaningful information that just got erased.
This is why directional trends matter. Before any data gets pooled, the first question should always be: which direction are most of these studies pointing? That pattern often tells you more than the pooled average.
5. Starting With the Answer and Working Backwards
Research should follow the evidence. But meta-analyses can be designed — intentionally or unintentionally — in ways that predetermine the outcome through the studies they choose to include or exclude.
Narrow inclusion criteria⁸ — the rules that determine which studies qualify for the meta-analysis — can quietly shape results toward a specific conclusion. If the criteria systematically exclude smaller studies, studies using different methodologies, or studies that don't fit the preferred design, the resulting meta-analysis reflects only a portion of the available evidence.
This is not always deliberate. Sometimes it reflects a genuine methodological preference. But the effect is the same: the conclusion appears to be based on comprehensive evidence when it is actually based on a curated selection of it.
Why This Matters to You
You might be thinking — this is a research problem, not my problem. I'm not writing meta-analyses.
But it absolutely affects you.
When a trainer, a physical therapist, a physician, or a wellness professional tells you that something "isn't supported by the research," there is a very real chance that conclusion came from a meta-analysis that pooled incompatible studies, dismissed meaningful effect sizes⁵ because they didn't cross a statistical threshold, or confused null results with proof of ineffectiveness.
Understanding these problems doesn't mean rejecting research. It means reading it more carefully — and not accepting a headline conclusion without asking how that conclusion was reached.
The goal of research is to get closer to truth. When the methods undermine that goal, the conclusions mislead rather than inform. And in fields as practical and consequential as physical rehabilitation and human performance, that has real consequences for real people.
What Good Research Practice Looks Like
Rather than simply critiquing what goes wrong, it's worth naming what a more reliable approach looks like:
Start with a broad, inclusive search — don't exclude studies prematurely based on design or size. Sort studies carefully by population, intervention type, and outcomes measured before combining anything. Identify directional trends first — which way are most studies pointing — before averaging data. Use meta-analysis selectively, only when studies are genuinely comparable and effect sizes⁵ are consistent. And always weigh statistical results alongside practical, real-world relevance.
That kind of systematic, transparent approach produces conclusions that are actually useful — not just statistically tidy.
What's Coming in Parts 2 and 3
In Part 2, we look at the conjugate system⁹ — a foundational, integrated approach to training and performance that addresses many of the fragmentation problems that make research in this field so difficult to interpret in the first place.
In Part 3, we walk through practical implementation — how to move from understanding these principles to actually applying them in training and recovery in a way that is systematic, measurable, and responsive to the body's real needs.
Connected Resources
For a deeper look at how active and passive movement fit into a complete approach to physical care, read:
Does It Matter Whether Movement Is Active or Passive? What the Research Actually Says
Are You Actually Moving — or Just Being Moved?
Footnotes / Reference Glossary
¹ Meta-Analysis — A research method that combines data from multiple individual studies on the same topic to produce a single pooled conclusion. Powerful when done correctly; misleading when studies are incompatible or methods are flawed.
² High-Intensity Interval Training (HIIT) — A training method alternating short bursts of intense exercise with recovery periods. Used as an example of how different exercise types cannot be meaningfully averaged together.
³ Heterogeneity — In research, the degree to which studies differ from one another in population, design, intervention, or outcomes. High heterogeneity makes pooling data unreliable.
⁴ P-Value — A statistical measure indicating the probability that a result occurred by chance. A p-value below 0.05 is conventionally considered "statistically significant." Often misused as the sole measure of whether a result matters.
⁵ Effect Size — A measure of the magnitude of a result — how large or meaningful a change actually was, independent of whether it crossed a statistical threshold. More practically informative than p-values alone.
⁶ Null Result — A study outcome where no statistically significant difference was found between groups. Does not prove an intervention is ineffective — only that this study did not detect a significant effect.
⁷ Statistical Power — The ability of a study to detect a real effect if one exists. Studies with small sample sizes often lack sufficient power, making null results unreliable as evidence of ineffectiveness.
⁸ Inclusion Criteria — The rules that determine which studies qualify to be included in a meta-analysis or systematic review. Narrow or biased criteria can shape results toward predetermined conclusions.
⁹ Conjugate System — A holistic, integrated approach to training and performance that emphasizes systematic progression, real-time feedback, and the relationship between effort, preparedness, and adaptation. Explored in depth in Part 2 of this series.
This article is for educational purposes. It is not intended as personal medical advice. If you are dealing with pain, injury, or movement limitations, please consult a qualified healthcare provider.