Why do scientists insist on calling things “significant” even when they don’t sound that important at all? For instance, neuroscientists might announce a “significant improvement” in memory after people drink blueberry juice for several weeks. Sounds impressive until you realise it means a tiny but reliable effect, not that blueberries turn you into Einstein. So what exactly makes a result significant? In science, “significant” doesn’t mean “big” or “important.” It means the difference is unlikely to be due to chance. The real question is: when we see a difference, how confident can we be that it’s genuine, and not just random variation? We begin with the assumption that there is no real difference — any variation is just noise. This is called the “null hypothesis.” In practice, it’s just healthy scepticism: it takes convincing evidence to show there’s something genuinely different going on. That evidence comes from statistical tests. If the differences in our data can’t easily be explained by random variation, we call them “significant.” Here’s a classic neuroscience example. Imagine I measure the brains of 1,000 people and find that, on average, men’s brains weigh 1,400 grams and women’s brains 1,300 grams. Have I uncovered a genuine sex difference? My starting assumption must be: probably not — this could just be random variation. So I run a statistical test (a t-test) and discover the difference is unlikely to be explained by noise. Statistically speaking, men’s brains are significantly heavier than women’s. But what does that actually tell us? 👉 Not that all men’s brains are bigger. In fact, it doesn’t say anything at all about individuals because there’s a lot of variability in brain size. Any given man may have a smaller brain than a given woman. 👉 It says nothing about function. For instance, there’s no evidence of differences in intelligence between men and women. 👉 It’s most likely down to body size: bigger bodies tend to have bigger brains, on average, and men tend to have larger bodies than women. So yes, the difference is statistically significant. But that doesn’t make it meaningful or important. All it really says is that the difference is unlikely to be due to chance. And that’s why “significance” matters: it’s a way of asking whether we’re seeing a real effect, or just the random bumps in the data. So the next time you hear a scientist declare something “significant,” remember — it doesn’t mean important. It means “we’re reasonably confident this isn’t just chance.” Whether it’s meaningful is another question altogether. Have you ever seen a headline about a “significant finding” that turned out to be less exciting than it first sounded? Please share in the comments.
Biostatistics In Research
Explore top LinkedIn content from expert professionals.
-
-
NEW EDITORIAL ON EMPIRICAL EXECUTION: STATISTICAL AND SUBSTANTIVE SIGNIFICANCE JM EXPECTS Too much empirical research still treats statistical significance as the finish line. It is not. In our new Journal of Marketing editorial, we make a direct case for a higher standard: authors should report exact p-values, abandon threshold-based “star” thinking, and place far greater weight on effect sizes and substantive significance. A low p-value may indicate statistical evidence, but it does not by itself establish that a finding is important. What makes this editorial especially practical is that it goes beyond exhortation and offers a roadmap. Table 2 provides a structured overview of common effect size metrics across empirical settings, organized around association strength, impact, and global fit together with benchmarks. It gives authors and reviewers a clearer framework for deciding which effect size metric fits the research design, instead of relying on habit or convention. Table 3 shows what better empirical reporting actually looks like. It illustrates how to present findings without hiding behind “significant” versus “not significant,” how to report exact p-values and effect sizes together, and how to interpret results in terms of their substantive impact on stakeholders, firms, brands, and society. Rigor in marketing research means clearer evidence, better judgment, and stronger insight into whether an effect is actually consequential in the real world. The editorial is open access. You find the link in the first comment below. If you enjoyed this, share it with others and follow me, Jan-Benedict Steenkamp, for more writing. #JournalOfMarketing #MarketingScience #MarketingResearch #EmpiricalResearch #EffectSizes #ResearchTransparency #OpenScience #AcademicResearch
-
Machine learning beats traditional forecasting methods in multi series forecasting. In one of the latest M forecasting competitions, the aim was to advance what we know about time series forecasting methods and strategies. Competitors had to forecast 40k+ time series representing sales for the largest retail company in the world by revenue: Walmart. These are the main findings: ▶️ Performance of ML Methods: Machine learning (ML) models demonstrate superior accuracy compared to simple statistical methods. Hybrid approaches that combine ML techniques with statistical functionalities often yield effective results. Advanced ML methods, such as LightGBM and deep learning techniques, have shown significant forecasting potential. ▶️ Value of Combining Forecasts: Combining forecasts from various methods enhances accuracy. Even simple, equal-weighted combinations of models can outperform more complex approaches, reaffirming the effectiveness of ensemble strategies. ▶️ Cross-Learning Benefits: Utilizing cross-learning from correlated, hierarchical data improves forecasting accuracy. In short, one model to forecast thousands of time series. This approach allows for more efficient training and reduces computational costs, making it a valuable strategy. ▶️ Differences in Performance: Winning methods often outperform traditional benchmarks significantly. However, many teams may not surpass the performance of simpler methods, indicating that straightforward approaches can still be effective. Impact of External Adjustments: Incorporating external adjustments (ie, data based insight) can enhance forecast accuracy. ▶️ Importance of Cross-Validation Strategies: Effective cross-validation (CV) strategies are crucial for accurately assessing forecasting methods. Many teams fail to select the best forecasts due to inadequate CV methods. Utilizing extensive validation techniques can ensure robustness. ▶️ Role of Exogenous Variables: Including exogenous/explanatory variables significantly improves forecasting accuracy. Additional data such as promotions and price changes can lead to substantial improvements over models that rely solely on historical data. Overall, these findings emphasize the effectiveness of ML methods, the value of combining forecasts, and the importance of incorporating external factors and robust validation strategies in forecasting. If you haven’t already, try using machine learning models to forecast your future challenge 🙂 Read the article 👉 https://buff.ly/3O95gQp
-
How can we decrease pharmacy spend on high-cost drugs by double digits without worse outcomes? --- Uplift modeling is a common tactic in marketing to target the specific people for a promotion that otherwise wouldn’t buy the product. While marketing in general can lead to overconsumption, in healthcare/#pharmacy, the same mathematical techniques used for uplift modeling could be repurposed to support #PrecisionMedicine or personalized medicine, where the goal is to identify which patients are most likely to benefit from a specific treatment while avoiding unnecessary treatments for patients who might not respond well. Identifying the cohort that is getting most of the outcomes from a drug varies by drug, but some drugs have only a fraction of the total population driving a larger share of clinical results. --- Here's the basic process for using #UpliftModeling (you can find more details from my Milliman white paper in the comments): 1. Treatment: Identify the treatment for which you want to predict response (e.g., a high-cost brand/specialty drug like GLP-1s). This could also be done for a medical device or any intervention. 2. Data collection: Gather comprehensive data and studies about patients, including their medical history, genetic information, and any other relevant attributes. This is often the limiter of building a good model. 3. Control group: Assemble a control group of patients who are similar to those receiving the treatment but are not receiving the treatment themselves. This helps establish a baseline for comparison. 4. Outcome measurement: Measure the effectiveness of the treatment for both the treatment group and the control group. This could involve monitoring health improvements, cardiac events, or other relevant medical outcomes. For FDA-approved drugs, this could come from published research on the “absolute risk reduction” or “number needed to treat.” 5. Model building: Develop predictive models using machine learning algorithms that estimate the likelihood of a positive response to the treatment for each individual. 6. Uplift calculation: Calculate the difference in response rates between the treatment group and the control group to determine the net impact of the treatment. 7. Segment: Divide patients into different segments based on their predicted response probabilities. 8. Action: Use the insights from uplift modeling to guide treatment, coverage, or other decisions. --- A payer or employer can use this information how they’d like, but I imagine it will be used to adjust formularies or utilization management strategies. It could also be used when setting up contracts for how a drug should be used while carving out certain drugs or disease states (e.g. oncology drugs at a center of excellence). There are more potential use cases in the white paper in the comments. --- Would you use this strategy for #PharmacyBenefits or #ValueBasedCare models that take on risk for cost of care?
-
🤔 Ever wondered you get hard core scientific proof that your correlations and model results aren't just spurious ❓ 🥇 The example here is the gold standard. Let's take the #Tech sector #XLK We have produced a factor model, where XLK returns are a function of macro factor returns like real GDP Nowcasting, inflation, real/nominal rates, credit spreads, the US Dollar. 12 factors in total (with the data all normalized and "de-correlated" using a Partial Least Squares Regression PLSR) 👉 We ran a Null Hypothesis test: A statistical method for determining if a REAL relationship exists in a population, or if an observed relationship in a sample is just due to chance It involves assuming that a “null hypothesis” of “no effect” is true and then using sample data to decide if there is enough evidence to reject it in favour of an alternative hypothesis The test’s outcome helps researchers make inferences about the larger population based on sample data, ensuring statistical rigour and managing the risk of false conclusions For a particular day, given that one has the (historical) factors mean return and their CoVar matrix (125 trading days, 90 half-life) ..and assuming the factor return jointly follows a multivariate Gaussian distribution (or any other distribution like an alpha-stable) ..it is possible to generate (simulate) multivariate random draws of our factor returns that follow that distribution (correlations included). We generate 125 of these simulated random draws in each step (the same as the historical window) Then we take these random generated factor returns and we regress them against the target (e.g. XLK) For this operation (PLSR), we also get the value for the macro exposures. ❗ These exposures were obtained from a random sample, therefore they are the result of chance. 🤖 We repeat the above process 10,000 times and we record those 10,000 exposures (and the R^2) and we do a histogram (in blue) with them This histogram give us the "range" of exposure values one can get from a pure chance process Then we do one extra PLSR this time with the REAL factor return data We plot the real exposures over the previous histograms (red line) ❓ And the question is: Are the red lines (real exposures) well inside the histograms of the random samples or not ? If they are, then those exposures (or R^2) are NOT significant because they could have been obtained just by chance However, when we look at these plots we see that the R^2 are every time very far from the histograms , and many of the model exposures are on the distribution tails (> 95% tail) or much further away One can only conclude that: 1️⃣ Macro is driving XLK: Significant R^2 (outside of the histograms by over 40 std deviations) 2️⃣ Many macro exposures also are significant (outside the histograms), because they couldn't have been the result of chance 👉 A null hypothesis test on your model is a very rigorous way to test for spuriousness #equities #factorinvesting
-
Recent population health data suggests that residents in the northern region of Singapore, particularly towns such as Woodlands, Yishun, and Sembawang, have a higher prevalence of diabetes and hypertension compared with national averages. While the numbers are clear, the underlying causes are less certain. Several plausible factors may contribute: • Demographic structure, including an older population profile • Socioeconomic gradients that influence diet, stress, and health behaviours • Differences in physical activity and lifestyle patterns • Ethnic distribution and associated metabolic risk profiles • More active screening and detection through primary care networks However, these remain hypotheses. More rigorous research is needed to understand the drivers behind this geographic clustering of chronic disease. One promising approach is the use of artificial intelligence for population health monitoring. AI can integrate multiple data sources such as electronic medical records, screening programmes, pharmacy data, wearable devices, and socioeconomic indicators to detect emerging patterns of disease. With machine learning and geospatial analytics, health systems could identify high-risk neighbourhoods earlier, monitor behavioural risk factors such as physical activity, and predict which communities are most vulnerable to chronic disease. This would allow health systems to move beyond reactive care toward proactive population health management. Instead of waiting for complications to appear, we can anticipate risk, target prevention programmes, and evaluate whether interventions are working. Understanding why disease burden concentrates in specific communities is essential for designing effective public health strategies. Combining epidemiological research with AI-driven monitoring may help us better understand these patterns and ultimately improve the health of our population. #PopulationHealth #Diabetes #Hypertension #AIinHealthcare #PublicHealth #Singapore
-
‘Getting over ANOVA’ is the title of our paper on multi-group data, out today in Nature Methods. The break-up is overdue. ANOVA asks whether all the groups are the same—a question nobody wants answered—and then sends you off to a pile of post-hoc tests nobody wants either. So what replaces it? We need methods that answer what we actually want to know: which groups differ, in which direction, and by how much. And they should not just report it, but also show it. The new paper describes a software package, DABEST 2.0, that brings estimation graphics to multi-group data: repeated measures, two-factor interactions via delta-delta effects, binary outcomes, and internal replicates via mini-meta. Some graphics can directly replace an ANOVA method. Each graphic shows the raw data, the effect size, and the uncertainty. Building software to visualize multi-group effect sizes has been a collaborative effort by the DABEST team, and I'm proud of what we built. DABEST is open source and available in Python, R, and through a web app. Data analysis should be easy to practice, and give you direct answers to the questions your experiments were designed to ask. Shout out to the team: Zinan Lu, Jonathan Anns, Yishan Mai, ROU ZHANG, CFA, Kahseng Lian, Nicole L., Shan Hashir, Zhuoyu Wang, Yixuan Li, A. Rosa Castillo, Joses Ho, Hyungwon Choi, Sangyu Xu #DABEST #EstimationStatistics #DataVisualization #OpenScience
-
Statistics nearly ended a friend’s PhD. He wept. Not because it was difficult. Because he kept asking the wrong questions. One poor statistical decision can wreck solid research. He chased methods and skipped thinking. Everything shifted when I told him to pause and ask these five compelling questions. 1. How many variables are you dealing with? One variable → describe it. Means, charts, frequencies. Many variables → reduce or group. PCA, factor analysis, clustering. 2. What are you actually trying to do? Describe → summaries and visuals. Compare → t-tests, ANOVA. Classify → logistic models, decision trees. Predict → regression, time series. Explain → multiple regression, path analysis. 3. What type of data do you have? Nominal → chi-square, logistic models. Ordinal → non-parametric tests, ordinal models. Continuous → correlation, regression, ANOVA. 4. Do variables have clear roles? Dependent and independent → model the relationship. No clear roles → explore patterns and structure. 5. Is time, space, or sequence involved? Time → trends, ARIMA. Space → spatial analysis. Sequence → check drift, bias, instability. Statistics is not magic. It is disciplined decision-making. Ask better questions. The method follows. ♻️If this helped, — Like + comment + repost to one person stuck with data. 🔔Follow Edidiong Ukpong(PhD Architecture) for clear, grounded research thinking.
-
Me, watching someone misdescribe p-values at a conference ……Do you think you can pass the P-value explanation test❓. First A p-value is → Not a badge of truth or a certificate of real-world impact ➊ 𝗔 𝗽-𝘃𝗮𝗹𝘂𝗲 𝗶𝘀 → The probability of observing results as extreme (or more extreme) as yours → G𝗶𝘃𝗲𝗻 𝘁𝗵𝗮𝘁 𝘁𝗵𝗲 𝗻𝘂𝗹𝗹 𝗵𝘆𝗽𝗼𝘁𝗵𝗲𝘀𝗶𝘀 𝗶𝘀 𝘁𝗿𝘂𝗲 ————————— For example: ➋ A 𝗽-𝘃𝗮𝗹𝘂𝗲 𝗼𝗳 𝟬.𝟬𝟯 𝗱𝗼𝗲𝘀 𝗻𝗼𝘁 𝗺𝗲𝗮𝗻: → “My intervention worked” → “There’s a 97% chance the null is false” → “We’ve found definitive proof” Instead… → It means there’s a 3% chance that you would observe results this strong (or stronger) if there were truly no effect. ➌ 𝗪𝗵𝘆 𝗱𝗼𝗲𝘀 𝟬.𝟬𝟱 𝗺𝗮𝘁𝘁𝗲𝗿? → In research, we often use 0.05 as a conventional cutoff for statistical significance → If your p-value is less than 0.05, we say the result is “statistically significant” → This means: it’s unlikely the observed results happened by chance under the null BUT → Statistical significance ≠ practical relevance → p < 0.05 doesn’t mean “definitely effective” → And p > 0.05 doesn’t mean “no effect at all” ➍ 𝗧𝗵𝗶𝘀 𝗶𝘀 𝘄𝗵𝗲𝗿𝗲 𝗺𝗼𝘀𝘁 𝗴𝗲𝘁 𝗶𝘁 𝘄𝗿𝗼𝗻𝗴 → They treat p-values as a truth switch: “Yes” or “No” → But statistics is nuance. Interpretation matters. → And misrepresenting the basics undermines public trust in science. ————————- 💬 Have you seen p-values misused or misunderstood in public discussions? ♻️ Repost to help raise the bar for statistical literacy in public health. #StatisticalThinking #PValueMyths
-
What statistical test would you use in this UX study? You are evaluating three new interface designs in a UX experiment. Each participant interacts with all three interfaces, and you collect two key outcomes: task satisfaction and task completion time. Your goal is to determine whether the design meaningfully affects the user experience. At this point, many researchers divide the data into multiple comparisons and run several t-tests. They compare satisfaction scores between each pair of designs and then do the same for completion time. While this approach might feel intuitive and convenient, especially when using familiar tools, it introduces serious issues. Running multiple t-tests increases the likelihood of false positives and treats each outcome independently, ignoring the fact that satisfaction and time are often related. This fragmented approach weakens statistical validity and risks overlooking meaningful patterns in how interface design influences overall experience. A more appropriate method, particularly when dealing with continuous dependent variables such as satisfaction and completion time, is MANOVA, which stands for Multivariate Analysis of Variance. This technique evaluates whether design has a combined effect on both outcomes while accounting for their potential correlation. It offers a more comprehensive and accurate understanding of how design affects the user experience. Not all UX study designs are this straightforward. Often, participants complete multiple tasks, interact with various designs across sessions, or respond to stimuli of varying complexity. These scenarios create repeated or nested structures that traditional ANOVA or MANOVA cannot handle well. In such cases, mixed-effects models are more appropriate. They allow researchers to model both fixed effects, like interface design, and random effects, such as variation across users or tasks. These models are particularly useful with unbalanced data, hierarchical structures, or irregular repeated measures. While powerful, both MANOVA and mixed-effects models require assumptions like multivariate normality, linear relationships, and sphericity to be checked. When applied correctly, they offer the flexibility needed to analyze complex UX data without losing valuable variability. Selecting the right test can be challenging, especially with so many possible designs such as between-subjects, within-subjects, repeated measures, or studies with multiple outcomes. That is why I created the table below. It summarizes common parametric tests based on study structure to help researchers choose more confidently. Although it focuses on standard comparisons, it also highlights when advanced methods like mixed-effects models are more appropriate for complex designs.