Pooled CRISPR screens: from read counts to hits you can trust
A genome-wide screen costs months of work and a good deal of money, and the analysis decides whether any of it counts. This guide is about the decisions that separate a hit list from a wish list.
Who it is for
You are running, or about to run, a pooled loss-of-function screen, in cells or in animals, and you want to know before you start what the data will and will not be able to tell you, and afterwards how to rank what you found.
What is in it
- What a screen measures. Fitness, not function, and why that distinction decides how you interpret every hit.
- Design choices that become statistics later. Library coverage, the number of guides per gene, replicates, and the bottleneck that an animal or a selection step imposes on a population.
- Reading the standard tools. What the common analysis packages actually compute, where their assumptions hold, and how their outputs differ on the same data.
- Controls. Non-targeting guides, essential-gene guides, and a positive control that you expect to find, run through the same pipeline as everything else.
- Ranking and choosing. How to combine effect size, consistency across guides and prior knowledge into a shortlist worth the cost of validation.
- Screens in unusual settings. In vivo screens, screens under immune pressure, and screens in organisms other than human cell lines, with worked examples from my own lab’s work in a parasite.
- Checklists for design, analysis and reporting.
Why I wrote it
My lab has run genome-wide screens in Toxoplasma for a decade, in culture, in mice, and under interferon pressure, and has published a step-by-step protocol for screens in animals. The hardest lessons were never about the wet work. They were about what the counts meant.