A fishing expedition runs test after test, across every tool, step, shift and parameter, until something comes up with p below 0.05. With 20 independent tests at alpha 0.05, the chance of at least one false hit is about 64 percent, so whatever surfaces is often noise. The defenses are stating hypotheses before looking, correcting for multiple comparisons and confirming any catch with fresh data.
He ran 300 commonality tests and found one tool at p of 0.03. That's a fishing expedition, not a root cause.
Also heard as
- data dredging
- p-hacking
- torturing the data
Related terms
p-value
Probability of seeing a result at least this extreme if there were truly no effect, used to judge statistical significance.
Significant
Describes a result unlikely to have come from chance alone at a stated risk level, which is not the same as important.
Confounding
Mixing of two effects so that the data cannot separate them, whether by design or through an uncontrolled variable.
Power
Probability that a test will detect a real effect of a given size, set by sample size, noise and alpha.