Error vs Effect Size

Understanding When Error Obscures Biological Effects

The Central Problem: An experiment can only reliably detect a biological effect if the effect size is substantially larger than the measurement error. When error approaches or exceeds effect size, you cannot confidently conclude whether observed differences are real or just noise.

Key Concept: Effect size ÷ Error size = Signal-to-Noise Ratio. A ratio of 2:1 or greater is generally needed for confident conclusions. Below this, biological effects become "lost in the noise."

Accuracy & Precision — the Dart Board Analogy

Precision — how consistent are repeated measurements? Darts landing close together = high precision = small error bars.
Accuracy — how close are measurements to the true value? Darts centred on the bullseye = high accuracy.
These are independent. A mis-calibrated pH probe can give precise but inaccurate results. Error bars on a graph measure precision, not accuracy.

High Precision · High Accuracy

Tight cluster on the bullseye. Small error bars. Result is reliable and correct.

High Precision · Low Accuracy

Consistent but off-target — e.g. a sensor with a fixed calibration offset. Small error bars but a wrong conclusion.

Low Precision · High Accuracy

Wide scatter centred on target. Large error bars. More replicates improve confidence.

Low Precision · Low Accuracy

Wide scatter AND off-target. Large error bars and a wrong conclusion. A method problem, not just a replicate problem.

Interactive Scenario Builder

0.40 pH units
±0.050 pH units
±0.08 pH units
3 replicates
Total Error (±)
0.094
Signal-to-Noise
7.3:1
Effect / Error
7.3×
Conclusion
Detectable
Your scenario — live
Adjust the sliders above to see your scenario's precision mapped to the dart board.
What is an error bar? Each column shows the average (mean) pH measured across all trials. The red ⊢ symbol extending above and below each column is called an error bar — it shows how much the measurement could vary. The top cap is the highest the true average could reasonably be; the bottom cap is the lowest. A shorter error bar means more precise, consistent measurements. If the error bars of two groups overlap, you cannot confidently claim a real difference between them.
Current Scenario Interpretation
✓ Reliable Detection Possible

With a signal-to-noise ratio of 7.3:1, this experiment can reliably detect the biological effect.

Why this works:

  • True effect (0.40 pH) is 7.3× larger than measurement uncertainty
  • With 3 replicates, effective error reduced to ±0.054 pH
  • Observed differences are clearly larger than "noise"
  • Confident conclusion about biological effect is justified

Realistic VCE Biology Scenarios

Click to load these scenarios into the visualisation above

✓ Ideal: Blue Light vs Dark (Method 3)

Effect: 0.6 pH units (strong photosynthesis)
Method Error: ±0.026 pH (spectrophotometry)
Execution: ±0.03 pH (careful technique)
Result: Effect is 15× larger than error → Highly reliable conclusion

⚠ Marginal: Small Effect with pH Meter

Effect: 0.3 pH units (weak response)
Method Error: ±0.1 pH (pH meter drift)
Execution: ±0.15 pH (variable technique)
Result: Effect only 1.7× larger than error → Questionable conclusion

✗ Undetectable: Tiny Effect, High Error

Effect: 0.15 pH units (minimal response)
Method Error: ±0.1 pH (visual comparison)
Execution: ±0.2 pH (poor standardisation)
Result: Error exceeds effect → Cannot draw conclusion

⚠ Few Replicates Problem

Effect: 0.4 pH units (moderate)
Method Error: ±0.05 pH (good method)
Execution: ±0.12 pH (student variation)
Replicates: Only 2
Result: Need more data → Statistical power too low

✓ Solved by Replicates

Effect: 0.25 pH units (small but real)
Method Error: ±0.026 pH (spectrophotometry)
Execution: ±0.05 pH (some variation)
Replicates: 6 (good design)
Result: Averaging improves confidence → Reliable detection

✓ Method Dominates Error

Effect: 0.5 pH units (strong)
Method Error: ±0.1 pH (moderate precision)
Execution: ±0.05 pH (excellent technique)
Result: Good technique can't fix poor method, but effect still detectable → Conclusion possible

Key Teaching Points

1. Effect Size Matters Most

A large biological effect (e.g., blue light producing ΔpH = 0.6) can be detected even with moderate error. A tiny effect (ΔpH = 0.1) may be undetectable even with excellent precision. Choose experimental conditions that produce large effects when possible.

2. Total Error = √(Method² + Execution²)

Errors combine via root-sum-of-squares, not simple addition. This means the largest error source dominates. If method error is ±0.1 and execution is ±0.05, total error is ±0.11 (not ±0.15). Improving execution can't fix a fundamentally imprecise method.

3. The 2:1 Rule of Thumb

When effect size is at least twice the error magnitude, you can usually draw conclusions. Between 1:1 and 2:1 is questionable. Below 1:1, the effect is "lost in the noise" and no reliable conclusion is possible.

4. Replicates Reduce Effective Error

With n replicates, effective error reduces by √n. Three replicates reduce error by 1.73×, six replicates by 2.45×. This is why the 96-well plate format (enabling duplicates/triplicates) is so powerful - it directly improves signal-to-noise ratio.

5. When Students Blame Themselves

If students are using Method 2 (pH meters, ±0.1 pH error) and observe a small effect (ΔpH = 0.15), they may blame "student error" when they can't detect it reliably. The truth: the method itself cannot resolve such small effects. Even perfect execution won't help.

6. Method Selection Is Hypothesis-Dependent

If your hypothesis predicts large effects (ΔpH > 0.5), Method 2 (pH meters) may suffice. If predicting subtle effects (ΔpH < 0.2), you must use Method 3 (spectrophotometry). The method must match the expected effect size.

7. Statistical Significance ≠ Biological Importance

With enough replicates, even tiny effects become "statistically significant." But ask: is ΔpH = 0.05 biologically meaningful for algal photosynthesis? The converse is also true: a large, important effect might not be statistically significant with too few replicates or too much error.

8. Experimental Design Trade-offs

Better methods reduce error but may increase complexity/cost. More replicates improve confidence but require more time/resources. Good experimental design balances precision needed (based on expected effect size) with practical constraints.