CRISPR Guide Library Design for Targeted Gene Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing genome-scale CRISPR-Cas9 knockout libraries require large numbers of cells to maintain genome-scale representation and lack the ability to design custom libraries targeting specific gene sets with higher coverage for genes like kinases, transcription factors, and chromatin modifiers.
Innovation Solution
A computer-implemented method for designing CRISPR-Cas system guide sequences by identifying target regions, ranking them based on off-target avoidance and on-target efficiency scores, and generating guide sequences to optimize targeting specific genes, using tissue-specific expression data and protein domain presence, with optional exclusion criteria for homopolymer repeats and transcriptional terminators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If genome-scale CRISPR-Cas9 knockout libraries are designed to cover all genes in the genome, then comprehensive gene coverage is achieved, but large numbers of cells (>10^8) are required to maintain representation
Solution Approach 1:
The patent segments the genome into specific gene sets of interest (e.g., kinases, transcription factors, chromatin modifiers, druggable genome) rather than attempting to cover all genes. This segmentation allows custom libraries to focus resources on biologically relevant targets, reducing the total number of guides needed while maintaining high coverage for priority genes.
Solution Approach 2:
The patent applies local quality by assigning different coverage depths to different genes based on their importance. Highly prioritized genes receive multiple guide sequences (higher coverage), while less critical genes receive fewer or no guides. This non-uniform distribution optimizes library size while ensuring adequate representation of key targets.
2Reliability
If guide sequences are selected without optimization, then library construction is simpler, but on-target efficiency and off-target avoidance are not optimized
Solution Approach 1:
The patent performs preliminary action by pre-computing on-target efficiency scores and off-target avoidance scores for all potential guide sequences before library construction. This advance calculation allows the selection of optimized guides without adding complexity to the actual library building process, as the computational work is completed beforehand.
Solution Approach 2:
The patent implements feedback by using predicted on-target and off-target scores to iteratively select and refine guide sequence candidates. Guides are ranked based on these scores, and selection criteria can be adjusted based on performance metrics, creating a feedback loop that optimizes guide quality while managing library complexity.
3Object-affected harmful factors
If all potential target regions are included in the library, then comprehensive target coverage is achieved, but off-target effects increase
Solution Approach 1:
The patent converts the potential harm of off-target effects into a benefit by using off-target prediction algorithms to identify and exclude problematic guide sequences. Rather than randomly including guides, the system actively screens out those likely to cause off-target effects, transforming the challenge of off-target prediction into a quality control mechanism.
Solution Approach 2:
The patent applies parameter changes by adjusting selection thresholds for on-target efficiency and off-target avoidance scores. By setting minimum cutoff values for these parameters, the system dynamically controls which guides are included, balancing off-target reduction with adequate target coverage based on specific experimental requirements.
Data Source
AI summary
Embodiments disclosed herein provide methods, including computer-implemented methods, for designing guide sequence which may be incorporated into custom, large scale guide sequence libraries. The methods require only a list of target genes as input and utilize on target and off target scores to generate an optimal set of guide sequences for a set of target genes. In certain embodiments, the methods may also utilize multi-tissue RNA-sequencing data and/or protein annotation to design targets to genes that are highly expressed and/or contain a functional protein domain. The invention further comprises guide libraries, cells comprising said guide libraries. Computer-implemented embodiments further improve computer system function by reducing excessive user wait time through the use of data structures that reduce search from linear to logarithmic time.


