Publications
Books on Corpus Statistics
Two titles available now and a third in preparation — written for language researchers rather than statisticians, by Muhammad Shoaib Tahir.
Publications
Featured Books
Practical knowledge for modern language research.

Stats in Corpus Linguistics
A practical guide to statistical analysis for corpus-based research — frequency, significance testing and effect size.

Measure by Measure
A practical handbook of association statistics — MI, t-score, log-Dice and log-ratio, and how to report them.

Introduction to Computational Linguistics
An accessible introduction to computational approaches to language, for readers from a linguistics background.
01 — Available now
Stats in Corpus Linguistics
A practical guide to statistical analysis for corpus-based research.
The statistics corpus linguists actually report, built up from first principles: what a frequency count is, why it must be normalised, when a difference is significant, and why significance alone is never enough.
What it covers
Written for postgraduate students and researchers who read statistics in papers and want to understand, check and report them properly.
02 — Available now
Measure by Measure
A Practical Handbook of Association Statistics for Corpus Linguistics.
Collocation depends entirely on which association measure you choose, and the measures disagree. This handbook sets out what each one rewards, how they rank the same data differently, and how to choose and report one honestly.
MI
Rewards rarity; needs a frequency threshold.
T-score
Rewards frequency; surfaces grammatical pairings.
log-Dice
Size-independent, so scores compare across corpora.
Log-ratio
An effect size in doublings, readable at a glance.
03 — Second edition in preparation
Introduction to Computational Linguistics
An accessible introduction to computational approaches to language.
Written for readers arriving from a language background rather than computer science. A revised and expanded second edition is on the way, updated from teaching the course to BS and MS students.
- Text processing foundations
- Tagging, morphology and parsing
- Language models explained in plain terms
- Practical work in Python, assuming no prior programming
- New in the second edition: expanded exercises and worked examples
Practise what the books teach
Free MCQ papers with answer keys, a browser practice system, and thirteen analysis tools for frequency, collocation and keyness.