CRISPR Guide RNA Efficiency Prediction via Segmented Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current CRISPR technology design tools fail to accurately predict guide RNA targeting efficiency due to their inability to account for gene-specific factors such as transcriptional activity and length, leading to variable and incomplete silencing efficiency.

Innovation Solution

An automated machine learning algorithm is used to predict guide RNA targeting efficiency by separating and combining models trained on gene intrinsic and guide features, allowing for improved prediction of guide depletion and selection of optimal guide RNAs for CRISPRi and other CRISPR technologies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current CRISPR technology design tools are used to predict guide RNA targeting efficiency, then the prediction process is simple, but the prediction accuracy is low due to inability to account for gene-specific factors

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The prediction model is segmented into two separate components: a gene-level model that processes gene-specific features (transcriptional activity, length, GC content) and a guide-level model that processes guide RNA features. These two models are trained separately and then combined, allowing each to specialize in specific aspects of targeting efficiency without overwhelming complexity in a single monolithic model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary decomposition step where the overall targeting efficiency is separated into gene-level effects and guide-level effects. This intermediary structure allows the system to account for confounded factors by processing them through separate computational pathways before combining the results, thereby improving accuracy without requiring a single overly complex model.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If gene-specific factors such as transcriptional activity and length are not accounted for, then the model complexity is low, but the silencing efficiency is variable and incomplete

Engineering Contradiction:
Improvesilencing efficiencyVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The model segments the analysis into gene-specific features (transcriptional activity, length, GC content) processed by a gene-level model and guide-specific features processed by a guide-level model. This segmentation allows reliable incorporation of gene-specific factors without creating an unwieldy single model, as each segment can be optimized independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent incorporates multiple gene-specific parameters (transcriptional activity, gene length, GC content) as input features to the gene-level model. By changing the parameters considered from a simple single-factor model to a multi-parameter gene-level model, the silencing efficiency reliability improves while the complexity is managed through the segmented architecture.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If data from screens providing levels of targeting characteristics efficiency are used without separating gene-level effects, then the data utilization is straightforward, but the guide efficiency prediction is confounded with gene-level effects

Engineering Contradiction:
Improveguide efficiency predictionVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts gene-level effects from the screen data by processing them through a separate gene-level model. This extraction removes the confounding influence of gene-specific factors from the guide efficiency prediction, allowing for more accurate guide-level predictions. The gene-level effects are taken out and handled separately rather than being mixed with guide-level analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The data processing is segmented into two distinct computational pathways: one for gene-level features and one for guide-level features. This segmentation allows the system to utilize screen data effectively while preventing gene-level effects from confounding the guide efficiency predictions, as each level is processed independently and then combined.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4179538B1Method for prediction of the guide efficiency when targeting a gene of interest
Publication Date: 2024.08.28 HELMHOLTZ ZENTRUM FUER INFEKTIONSFORSCHUNG GMBH
  • EP4179538B1 patent drawingFigure 1
  • EP4179538B1 patent drawingFigure 2
  • EP4179538B1 patent drawingFigure 3

AI summary

The present invention relates to a Method for prediction of the targeting efficiency of guides comprising guide RNA (gRNA) targeting a gene of interest, said guides locating a respective gene of interest as a target in a gene sequence, by evaluating data provided by screens, the screens providing levels of targeting characteristics efficiency of guides comprising the level of targeting efficiency confounded with gene-specific effects. The method comprising the steps of a) Selecting a set of guides along with a first set of gene intrinsic features related to their respective gene of interest, and a second set of guide features; b) Inputting the selected first set of gene intrinsic features and the selected second set of guide features into an automated machine learning algorithm and c) Calculating an estimate of targeting efficiency of said selected guides by use of the automated machine learning algorithm, wherein the automated machine learning algorithm comprising two models being separately trained from each other, wherein the first of the two models using the selected first set of gene intrinsic features and the second of the two models using the selected second set of guide features, and wherein the automated learning algorithm combines the two models with the result of a prediction of the depletion of the guides targeting the gene of interest.