CRISPR Positive Selection via Hypergeometric Probability Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current CRISPR/Cas9 genome engineering techniques face challenges in addressing inter-experiment variability and gene identification due to differences in gRNA virus titers, infection rates, and target efficiency, leading to inconsistent data analysis in positive and negative selection screens.
Innovation Solution
A method involving infecting cas9-positive cells with a library of viral vectors, sequencing to determine read counts, and calculating probabilities to identify genes associated with a phenotype by categorizing cells with a designated phenotype and applying statistical formulas to account for variability, thereby enhancing the accuracy of gene identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If existing techniques rely on read counts to identify gRNAs associated with a phenotype, then the analysis is simple, but the inter-experiment variability issues are not addressed and the reliability is poor
Solution Approach 1:
The patent transforms the analysis from using raw read counts to using probability values derived from hypergeometric distributions. This parameter change accounts for variability in gRNA virus titers, infection rates, and target efficiency by modeling the selection process statistically, thereby improving reliability while maintaining computational feasibility
Solution Approach 2:
The patent introduces probability calculations as an intermediary between raw sequencing data and gene identification. By computing the probability of observing certain gRNA frequencies by chance, the method mediates between the complex variability sources and the final phenotype association, filtering out false positives while preserving true signals
2Reliability
If multiple replicates are used to reduce variability, then the reliability improves, but the complexity of data analysis increases
Solution Approach 1:
The patent merges multiple replicate datasets into a unified probabilistic framework. Instead of analyzing each replicate separately and then combining results, the method combines all gRNA observations across replicates into a single hypergeometric probability calculation, simplifying the analytical process while maintaining statistical power
Solution Approach 2:
The patent changes the analytical parameter from individual replicate read counts to aggregate probability values that incorporate all replicates. This transformation reduces computational complexity by working with summary statistics rather than full replicate datasets, while still capturing the variability information
3Productivity
If gRNA virus titers and infection rates vary among experiments, then the productivity increases, but the measurement precision of gRNA abundance decreases
Solution Approach 1:
The patent changes the measurement parameter from absolute gRNA abundance (read counts) to relative probability values that are normalized for variability. The hypergeometric probability calculation inherently accounts for differences in virus titers and infection rates by modeling the selection process, thereby maintaining measurement precision across experiments with varying productivity
Data Source
AI summary
Methods and systems for CRISPR positive selection are described. CRISPR positive selection uses DNA sequencing to identify genes that their perturbation by CRISPR guide RNAs is correlated to the phenotype. In some aspects, disclosed are genome-wide CRISPR/Cas9 screening methods to identify genetic modifiers. Also disclosed are the apparatuses used for performing the methods.


