Deep Learning Prime Editing Efficiency Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting prime editing efficiency are limited by the lack of computational models that effectively identify factors affecting prime editing and predict activity at a given target sequence, hindering the optimization of prime editing processes.

Innovation Solution

A deep learning-based system is developed to predict prime editing efficiency by generating predictive models using high-throughput data sets of pegRNA-encoding sequences and target sequences, incorporating features such as PBS length, RT template length, editing type, and editing position to accurately forecast editing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning models are used to predict prime editing efficiency, then prediction accuracy is improved, but computational complexity and data requirements increase

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the complex biological sequence data into numerical feature representations (parameters) including sequence length, GC content, secondary structure properties, and positional features. This parameter transformation enables the deep learning model to process biological data efficiently while maintaining high prediction accuracy, resolving the contradiction between accuracy and computational complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The model architecture is segmented into distinct functional components: embedding layers for sequence representation, convolutional layers for local pattern recognition, recurrent layers for sequential dependencies, and fully connected layers for final prediction. This segmentation allows each component to specialize in specific computational tasks, improving overall efficiency while maintaining high accuracy.

Inventive Principle:
Principle #1Segmentation

2Reliability

If high-throughput experimental data is collected to train prediction models, then model reliability is improved, but experimental time and resource consumption increase

Engineering Contradiction:
Improvemodel reliabilityVSAvoidexperimental time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary data processing and feature extraction during the experimental data collection phase. By pre-processing the high-throughput sequencing data and extracting relevant features before model training, the system reduces the time required for subsequent model development and validation, while maintaining the reliability gained from comprehensive data collection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic training data by generating in silico pegRNA sequences with known features and simulated editing outcomes. These synthetic copies augment the experimental data, allowing the model to be trained on larger datasets without proportionally increasing experimental time and resources, thus maintaining reliability while reducing time loss.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230274792A1System and method for prime editing efficiency prediction using deep learning
Publication Date: 2023.08.31 IND ACADEMIC COOP FOUND YONSEI UNIV
  • US20230274792A1 patent drawing
  • US20230274792A1 patent drawing
  • US20230274792A1 patent drawing

AI summary

A system for predicting prime editing efficiency by using deep learning, including: an information input unit that receives an input of data on prime editing efficiency of a prime editor, a predictive model generator for generating prime editing efficiency predictive models by performing deep learning to learn a relationship between features affecting prime editing efficiency and prime editing efficiency, by using the data received from the information input unit, a candidate sequence input unit that receives an input of a candidate target sequence for prime editing; and an efficiency predictor for predicting prime editing efficiency by applying the candidate target sequence input into the candidate sequence input unit to an efficiency predictive model generated in the predictive model generator.