🧵 Why the obsession with p < 0.05 is hurting science. 1/ This meme says it all. p = 0.0501? Pain. p = 0.0499? Pure euphoria. Two numbers. Nearly identical. Yet we treat them like night and day. Why? 2/ The 0.05 p-value threshold is arbitrary. It came from R.A. Fisher in the 1920s. And we’ve been worshipping it like a sacred line ever since. But it’s not magic. It's convention. 3/ What does p = 0.05 actually mean? It means: If the null hypothesis is true, there’s a 5% chance we’d see this extreme of a result by random chance. That’s it. Not: "This is true." Not: "This will replicate." 4/ p = 0.0499 and p = 0.0501 are nearly identical. But one gets you a “significant” label. The other gets dismissed. That’s broken thinking. 5/ Quoting Mike Love: “A smaller p-value is not more interesting.” “We should focus on effect sizes.” He’s right. 6/ What’s an effect size? It tells you how big the difference is. Not just if it’s statistically detectable. A gene with a log2 fold change of 3 matters. Even if p = 0.06. 7/ P-values shrink with more data. Got 10,000 samples? You’ll find “significance” for even the tiniest differences. Statistically significant ≠ Biologically meaningful. 8/ Also, be careful when testing thousands of genes. Even with a p < 0.05 threshold, false positives will sneak in. Use multiple testing correction: FDR, Bonferroni. Always. 9/ Let’s reframe: Instead of: “Did I beat the p < 0.05 line?” Ask: Is the effect meaningful? Is it reproducible? Does it make biological sense? 10/ Want a better practice? Look at the distribution of p-values. Report adjusted p-values (FDR). Highlight effect sizes. Don’t cherry-pick. 11/ And don’t forget confidence intervals. They show the range of plausible effect sizes—not just a binary yes/no. More context, more truth. 12/ Key takeaways: 0.05 is a line in sand, not a cliff p-values ≠ effect size Focus on biological meaning Always correct for multiple testing Use p-values as part of the story—not the whole story 13/ If you're making big decisions based on p = 0.0499 vs 0.0501... You're not doing science. You're doing stats theater. Look deeper. Think harder. Go beyond the stars. 14/ And please—share this with a friend still chasing tiny p-values. Let’s stop celebrating noise and start celebrating insight. I hope you've found this post helpful. Follow me for more. Subscribe to my FREE newsletter chatomics to learn bioinformatics https://lnkd.in/erw83Svn
Mastering Statistical Analysis
Explore top LinkedIn content from expert professionals.
-
-
STATISTICAL V. SUBSTANTIVE SIGNIFICANCE Take a moment to consider the following scenario. One study with n = 100 reports a focal effect with an associated p-value of 0.02. Another study with n = 1000 reports a focal effect with an associated p-value of 0.02. Which study presents the strongest evidence the effect is really there? This scenario is adapted from Bakan (Psych. Bull. 1966). Many scholars chose the second scenario. They are wrong. I quote Bakan (p. 429): “The rejection of the null hypothesis when the number of cases is small speaks for a more dramatic effect in the population [larger effect size]; and if the p-value is the same, the probability of committing a Type I error remains the same.” Many papers equate implicitly or explicitly statistical significance with substantive significance. Yet, a p-value does not inform you whether the effect has any real world meaning. Ralph Tyler (Educ. Res. Bulletin, 1931) already wrote that a statistically significant difference is not necessarily an important difference, and a difference that is not statistically significant may be an important difference. Unfortunately, we are still making the same mistake 90 years later. A statistically significant result may be substantively nonsignificant (trivial). But also, a statistically nonsignificant result may be substantively significant. I see so many studies in our field reporting regression coefficients with *** and I have no idea how large the effect is. This tendency to equate statistical with substantive significance persists to the extent that the prestigious American Statistical Association (not exactly an organization afraid of advanced statistics) came out with a formal statement on p-values—the “ASA Statement on Statistical Significance and P-Values" cautioning researchers: "Statistical significance is not equivalent to scientific, human, or economic significance. Smaller p-values do not necessarily imply the presence of larger or more important effects, and larger p-values do not imply a lack of importance or even lack of effect. Any effect, no matter how tiny, can produce a small p-value if the sample size or measurement precision is high enough, and large effects may produce unimpressive p-values if the sample size is small or measurements are imprecise." I am not arguing against statistical significance. Rather that articles should report statistical AND substantive significance. In my view, our primary task is uncovering factors that make a meaningful difference. That calls for effect sizes. As a nice “bonus,” substantive significance is less amenable to p-hacking than statistical significance. If you enjoyed this, share it with others and follow me, Jan-Benedict Steenkamp, for more writing. Journal of Marketing
-
How can an A/B test be “statistically significant” but not be totally trustworthy? I’ve been wrestling with this question for over a decade. Through extensive research, I've now come to understand that a big part of the answer lies in confidence intervals. Here’s the simplest way to explain it: Imagine you run an A/B test with very small numbers: 🚥 Version A: 82 visitors, 3 conversions 🚦 Version B: 75 visitors, 12 conversions The math shows the result is statistically significant. ⚡ The p-value is 0.0088, well below the common p < 0.05 threshold ⚡ Observed power is reported as 95.09%, well above the standard 80% rate The results look convincing. Statistical significance, long treated as the gold standard, has been achieved! But, here's the problem. Statistical significance can be "gamed" with low traffic tests because it only answers one narrow question: 🔦 If there were truly no difference between versions, how likely is it this result would happen by chance? That’s it. That's all statistical significance answers. It doesn't tell you whether the result is stable or repeatable. And, as you can imagine, with tiny samples, like 3 vs. 12 conversions, you get exaggerated effects. Every single conversion has an outsized influence. One or two people behaving differently can completely flip the outcome. ➡️ This is where confidence intervals come in. A confidence interval is the range of outcomes that could reasonably be true given the data. In small tests, that range is really wide. So the actual conversion effect might be smaller or larger than what you achieved in the test. You can't know with precision. So you don't have a dependable estimate of how big the improvement really is, or whether the result would hold if you ran the test again. It's important to realize, a confidence interval is not the same as a confidence level. Remember: a confidence interval is the range of values that could reasonably be true given the data. A label of “95% confidence” describes how that range was constructed, not how certain or correct the result is. Which means, a 95% confidence interval can still be very wide, creating substantial uncertainty around the estimate. When there's such uncertainty, the numbers may appear exaggerated. That's where Twymann’s Law comes in. It states, anything that looks interesting or unusual is usually wrong. In small samples, results are extreme because the noise does most of the work. So while a statistically significant difference can be measured in a small-sample study, you can't reliably measure how big that difference actually is. That's why 3 vs. 12 conversions often fail to replicate once more data is collected. 📣 Call to action for 2026: Run tests that are not only statistically significant, but also have a large enough sample size to produce narrow confidence intervals, so you can not only detect effects, but also estimate them precisely enough to make accurate, trustworthy test decisions.
-
𝐓𝐡𝐞 𝐩-𝐯𝐚𝐥𝐮𝐞 𝐢𝐬 𝐨𝐧𝐞 𝐨𝐟 𝐭𝐡𝐞 𝐦𝐨𝐬𝐭 𝐦𝐢𝐬𝐮𝐬𝐞𝐝 𝐧𝐮𝐦𝐛𝐞𝐫𝐬 𝐢𝐧 𝐫𝐞𝐬𝐞𝐚𝐫𝐜𝐡. 𝐃𝐨 𝐲𝐨𝐮 𝐤𝐧𝐨𝐰 𝐰𝐡𝐲? Because many people treat it as a stamp of truth. It is not. A p-value does not tell you: ➤ The probability that your hypothesis is true ➤ The probability that the null hypothesis is true ➤ How large the effect is ➤ Whether the result is clinically useful ➤ Whether the study was well designed What it tells you is narrower: ➤ How surprising are these data if there is truly no effect? ↳ Smaller p-values mean the data would be more unexpected under the null hypothesis. That is useful. But it is not enough. Before you trust a p-value, ask: ➤ What is the effect size? ↳ A tiny effect can become statistically significant in a large study. ➤ What is the confidence interval? ↳ Precision matters when judging an estimate. ➤ What is the study quality? ↳ Bias can make a small p-value look more convincing than it should. ➤ What is the context? ↳ Statistical significance does not automatically mean practical importance. The p-value is a clue. Not a verdict. Good research interpretation does not stop at “p < 0.05.”
-
How to Determine Statistical Significance in Quantitative Research: Its Core Conditions, Statistical Laws, and Key Tools for Hypothesis Testing Statistical significance is a fundamental concept in quantitative research, used to determine whether the results obtained are due to chance or reflect a true relationship between variables. It is commonly expressed through the p-value, a numerical indicator used to assess the strength of the results. To determine statistical significance, researchers begin by formulating two hypotheses: the null hypothesis (H₀), which assumes no relationship or difference, and the alternative hypothesis (H₁), which suggests the presence of a relationship or difference. After collecting data, an appropriate statistical test is applied to calculate the p-value, which is then compared to a pre-defined significance level (usually 0.05 or 0.01). If the p-value is lower than the significance level, the null hypothesis is rejected, and the result is considered statistically significant. Key conditions for statistical significance include: - Selecting an appropriate significance level (α) before analysis. - Using a suitable statistical test based on the nature of the data (quantitative or categorical). - Ensuring a sufficient sample size to maintain statistical power. - Meeting the assumptions of the chosen test, such as normal distribution or homogeneity of variance. Statistical laws associated with significance include: - The law of probability: used to assess the likelihood of an outcome. - The normal distribution law: foundational for many tests like the t-test. - Laws of variance and standard deviation: used to measure data dispersion. Common tools for testing statistical significance vary depending on the data type and hypothesis structure, including: - The t-test for independent or paired samples. - ANOVA for comparing differences across multiple groups. - Chi-square test for categorical data. - Pearson correlation for relationships between two quantitative variables. It is important to note that statistical significance does not necessarily imply practical importance. Therefore, researchers are encouraged to accompany statistical analysis with scientific and contextual interpretation of the results. Additionally, using measures of statistical power helps evaluate the study’s ability to detect true differences or effects. So, understanding statistical significance—its conditions, laws, and tools—enhances the quality of scientific research and supports data-driven decision-making through rigorous and meaningful analysis.
-
Do you really need a p-value below 0.05? Maybe not. In most research, a p-value below 0.05 is treated as the finish line. If it's under that threshold, the result is called significant. If it's above, it often gets ignored. But that 0.05 cutoff isn't a law. It’s just a habit, and it may not match the goals of your study. That p-value threshold comes from something called the significance level, often referred to as alpha. Alpha represents the amount of risk you're willing to take when deciding whether an observed effect is real or just due to random chance. More specifically, it’s the probability of making a Type I error, which means concluding there's an effect when there actually isn't one. A standard alpha of 0.05 means you're accepting a five percent chance of making that mistake. But that level of risk isn’t always appropriate. If your decision could lead to serious consequences, such as launching a product change that affects user trust or safety, you may want to lower your alpha to 0.01 or even 0.001. If you’re exploring early signals or looking for ideas to investigate further, you might choose a more relaxed threshold like 0.10. The key point is that the significance level should reflect the cost of being wrong. Not every study deserves the same cutoff. This choice also affects how we interpret confidence intervals. A 95 percent confidence interval is tied directly to an alpha of 0.05. If you change your alpha to 0.01, your confidence interval becomes wider to reflect a 99 percent level of confidence. The width of the interval tells you how precise your estimate is, and whether zero is included in the range. If zero is within the interval, the effect may not be statistically meaningful at your chosen alpha. But more importantly, confidence intervals show the range of plausible values for the effect size, which often tells a more complete story than a single p-value. Rather than just asking "is it significant," you can ask "how large might the effect be, and how uncertain are we?" This becomes even more important when you are running multiple tests. If you apply the same alpha repeatedly across many comparisons, your chances of a false positive increase. That’s where adjustments come in. You may use methods like the Bonferroni correction or control the false discovery rate to keep your conclusions meaningful and trustworthy. All of this points to a larger truth. Statistical significance is not a final answer. It is a decision-making tool, and its value depends on how you use it. Choosing the right alpha level, interpreting confidence intervals properly, and understanding the cost of different types of errors are all part of responsible research. So before accepting the 0.05 rule by default, take a step back and ask yourself: what are the risks, what is the cost of being wrong, and what level of uncertainty are you actually willing to live with?
-
A p-value of 0.001 does not mean a treatment works well. It means the result is unlikely to be due to chance. That's it. It says nothing about whether the effect is large enough to matter to a patient. Statistical significance tells you whether an effect is likely real. It's a mathematical threshold (usually p < 0.05). With a large enough sample, even a tiny difference can be statistically significant. Clinical significance tells you whether that effect is meaningful in practice. Does it improve symptoms? Reduce hospital stays? Change how a patient feels day to day? Here's a concrete example. Imagine a trial of 20,000 patients testing a new blood pressure drug. The drug lowers systolic blood pressure by 1 mmHg compared with placebo. With that sample size, the p-value might be 0.001. Statistically significant. But 1 mmHg? That's not going to change anyone's clinical outcome. The reverse happens too. A small trial might show a 15-point improvement in quality of life, but with a p-value of 0.08. Not statistically significant. Yet that difference, if real, would be very meaningful to patients. As medical writers, we deal with both concepts constantly. Being precise about which one we're reporting (and which one a study actually demonstrated) matters. Next time you read "the difference was statistically significant," ask: was it clinically significant too? #MedicalWriting #ClinicalResearch #Biostatistics #MedComms #SciComm
-
𝐄-𝐯𝐚𝐥𝐮𝐞𝐬: 𝐀 𝐌𝐨𝐝𝐞𝐫𝐧 𝐀𝐥𝐭𝐞𝐫𝐧𝐚𝐭𝐢𝐯𝐞 𝐭𝐨 𝐏-𝐯𝐚𝐥𝐮𝐞𝐬 𝐢𝐧 𝐒𝐭𝐚𝐭𝐢𝐬𝐭𝐢𝐜𝐚𝐥 𝐓𝐞𝐬𝐭𝐢𝐧𝐠 In data science and statistics, most practitioners used to rely on p-values to make decisions. However, p-values are often criticized for being difficult to interpret and prone to misuse. To address these concerns, e-values have emerged as a powerful tool that simplifies hypothesis testing, particularly in the era of big data and optional stopping. 🔎 𝐖𝐡𝐚𝐭 𝐚𝐧 𝐄-𝐯𝐚𝐥𝐮𝐞 𝐢𝐬 An e-value is a measure of evidence against a null hypothesis. Unlike a p-value, which is a probability, an e-value is an expectation. If the null hypothesis is true, the expected value of an e-variable is at most 1. The larger the e-value, the stronger the evidence against the null. 💡 𝐓𝐡𝐞 𝐦𝐚𝐢𝐧 𝐢𝐝𝐞𝐚 The concept is rooted in game theory, and based on "betting" against the null hypothesis: ➡️ You start with a "budget" of 1. ➡️ The e-value represents your capital after observing the data. ➡️ If your e-value is 20, you have 20 times more evidence than you started with. ➡️ If the e-value is close to 0, the evidence for the alternative is weak. 💪 𝐖𝐡𝐲 𝐄-𝐯𝐚𝐥𝐮𝐞𝐬 𝐦𝐚𝐭𝐭𝐞𝐫 E-values solve several traditional problems found in classical p-value testing: 🛠️ Handling optional stopping: ▶️ With p-values, you cannot stop an experiment early just because you like the result (it inflates Type I error). ▶️ E-values remain valid even if you look at the data and decide to stop or continue the test. 🛠️ Combining evidence: ▶️ It is notoriously difficult to combine p-values from different studies. ▶️ E-values from independent tests can be multiplied to get a combined measure of evidence. 🛠️ Intuitive scaling: ➡️ An e-value of 10 means the data is 10 times more likely under the alternative than the null. ➡️ This "likelihood ratio" approach is much easier to explain to stakeholders than the "probability of seeing more extreme results". 🛠️ Safe testing: ▶️ E-values helps to design experiments that are robust to "p-hacking". ⚠️ 𝐋𝐢𝐦𝐢𝐭𝐚𝐭𝐢𝐨𝐧𝐬 ➡️ Can be more "conservative" than p-values in some scenarios. ➡️ Not yet as widely adopted in standard software libraries. ➡️ Requires a shift in mindset from "significance thresholds" to "evidence accumulation". 👉 𝐄-𝐯𝐚𝐥𝐮𝐞𝐬 represent a significant step forward in making statistics more flexible and reliable. If you are tired of the limitations of p-values in dynamic environments like A/B testing or real-time monitoring, e-values are definitely worth exploring. #Statistics #DataScience #HypothesisTesting #ABTesting #DataDriven #DecisionMaking
-
A few years ago, during a public health project in a rural district, we were studying the impact of a nutrition awareness program on childhood anemia. After months of fieldwork, data collection, and sleepless nights cleaning spreadsheets, we ran the analysis. The p-value came out to 0.049. The team celebrated “It’s significant!” someone said: Funding bodies were happy. But something didn’t sit right with me. Yes, the p-value was below 0.05, but when we looked closer, the actual reduction in anemia was barely 1.5%. Statistically significant? Maybe. Practically meaningful? Not really. We realized we were about to report a result that looked “positive” on paper but would not change lives on the ground. That moment changed how I understood research. In public health, p-values are tools, not truths. They tell us about the role of chance, not the size or importance of an effect. A "significant" p-value doesn’t always mean our intervention works in real life. And a "non-significant" result doesn’t always mean failure. Sometimes, it just means we need a bigger sample or a different lens. If we want to build real impact, not just good-looking papers, we have to stop worshipping the p-value and start asking deeper questions: Is the effect real? Is it useful? Will it matter to the people we serve? These are the reflections I share as someone passionate about evidence-based public health. If this resonates with you, follow along, I write often about data, decisions, and the human stories behind statistics. #PublicHealth #PValue #RealImpact #FieldResearch #HealthData #Epidemiology #BeyondStatistics #EvidenceBasedPolicy
-
Do you believe a significant p-value means the results aren’t due to chance? You’re wrong! A p-value below 0.05 doesn’t prove an effect is real, and it definitely doesn’t mean there’s only a 5% chance the result happened by luck. It simply tells us how often we’d expect to see data this extreme if there were no real effect. That’s a big difference. The problem is that p-values don’t measure what most people think they do. They don’t tell us how likely the hypothesis is true, they don’t measure the strength of an effect, and they don’t account for prior knowledge. In fact, even with p < 0.05, there’s often a 20–30% chance the result is a false positive. If the original hypothesis was unlikely to begin with, that number shoots even higher. This is a huge issue, especially in UX research, where we deal with complex user behaviors, context-dependent interactions, and subtle design changes. A rigid threshold like 0.05 doesn’t capture the full picture. That’s where Bayesian analysis comes in. Instead of giving a yes-or-no answer, Bayesian methods update our confidence as we collect more data. They factor in prior knowledge and tell us how likely an effect is real, given the data we have. This is critical in UX, where decisions are based on continuous learning rather than one-time experiments. Rather than asking if a new feature is "statistically significant," we should be asking: how likely is this change to improve the user experience? Bayesian methods allow us to weigh all available evidence and make better, more nuanced decisions. Science isn't about proving things with a single number. It’s about reducing uncertainty and making informed choices. P-values alone can’t do that, but Bayesian thinking can.