Deep Learning sgRNA Prediction Model for CRISPR/dCas Epigenetic Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing prediction models for CRISPR/Cas gene editing systems are not accurate for predicting the epigenetic editing activity of CRISPR/dCas systems, as they rely only on sequence features and lack consideration of epigenetic features.

Innovation Solution

A deep learning model that incorporates specific epigenetic features, such as distance between the transcription start site and the sgRNA target site, DNA methylation level, RNA expression level, and chromosome accessibility, to improve the accuracy of sgRNA editing efficiency prediction in CRISPR/dCas epigenetic editing systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing prediction models rely only on sequence features, then the model complexity is low, but the prediction accuracy for epigenetic editing activity is insufficient

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines sequence features with multiple epigenetic features (DNA methylation, histone modifications, chromatin accessibility, etc.) into a unified prediction model. This merging of different feature types enables accurate prediction of epigenetic editing activity while maintaining reasonable model complexity through efficient feature integration strategies.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The prediction model uses a composite feature structure that integrates multiple types of biological data (sequence information, epigenetic marks, chromatin states) similar to how composite materials combine different substances to achieve superior properties. This composite approach allows the model to capture complex biological interactions that single feature types cannot represent.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If multiple epigenetic features are added to improve prediction accuracy, then the prediction performance improves, but the training time increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of epigenetic features including normalization, filtering, and pre-computation of relevant epigenetic marks before model training. This preliminary action reduces the computational burden during training while preserving the predictive power of multiple epigenetic features, thus shortening training time without sacrificing accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The model extracts only the most relevant epigenetic features from the complete set of available epigenetic data through feature selection and importance ranking. By taking out only the essential features needed for accurate prediction, the model reduces training complexity and time while maintaining high prediction accuracy for epigenetic editing activity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250037801A1Method for predicting gene editing activity by deep learning and use thereof
Publication Date: 2025.01.30 CENT FOR EXCELLENCE IN BRAIN SCI & INTELLIGENCE TECH CHINESE ACAD OF SCI
  • US20250037801A1 patent drawing
  • US20250037801A1 patent drawing
  • US20250037801A1 patent drawing

AI summary

The present invention provides a model or tool for predicting the editing efficiency of sgRNA in a CRISPR/Cas gene editing system, in particular a CRISPR/dCas epigenetic editing system, a training and prediction method thereof, and a related computer system, computer storage medium, and application. In particular, one or more sgRNA and target gene related epigenetic features are added to an input of the model.