BE-Hive Machine Learning Model for Base Editing Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current base editing technologies face challenges in reliably predicting outcomes and selecting optimal base editors and guide RNAs for precise genome editing, due to the complex interplay between base editors and target sequences, leading to inefficient and often empirical optimization processes.

Innovation Solution

Development of machine learning models, referred to as BE-Hive, that predict genome editing outcomes and facilitate the selection of appropriate base editors and guide RNAs for specific genomic loci and desired genotype outcomes, considering various determinants such as target sequence context, cell type, and DNA repair proteins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If base editing is performed without predictive tools, then empirical optimization can be used to achieve editing outcomes, but the process becomes time-consuming and inefficient

Engineering Contradiction:
Improvebase editing efficiencyVSAvoidoptimization time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by developing and applying predictive machine learning models (BE-Hive) before conducting base editing experiments. These models analyze target sequence context, cell type, and DNA repair protein characteristics to predict optimal base editor and guide RNA combinations in advance, eliminating the need for time-consuming empirical optimization during the experimental phase

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If simple guidelines are used for target selection, then the process is easy to follow, but viable targets that do not fit canonical guidelines are overlooked

Engineering Contradiction:
Improvetarget selection simplicityVSAvoidtarget selection scope
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by using machine learning models that evaluate multiple parameters simultaneously (sequence context, cell type, DNA repair proteins) rather than relying on simple canonical guidelines. This allows the system to identify viable targets across a broader range of conditions while maintaining ease of use through automated prediction algorithms

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive analysis of base editor and target sequence determinants is performed, then prediction accuracy improves, but the complexity of the system increases

Engineering Contradiction:
Improveoutcome prediction accuracyVSAvoidpredictive system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies the intermediary principle by introducing machine learning models as intermediaries between the complex biological system (base editors, target sequences, cell types) and the user. These models handle the complexity of analyzing multiple determinants (sequence context, DNA repair proteins, cell type characteristics) internally, providing accurate predictions through a user-friendly interface without exposing the user to the underlying complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230123669A1Base editor predictive algorithm and method of use
Publication Date: 2023.04.20 THE BROAD INST INC
  • US20230123669A1 patent drawing
  • US20230123669A1 patent drawing
  • US20230123669A1 patent drawing

AI summary

The present disclosure provides a novel machine learning model capable of assisting those of ordinary skill in the art to conduct base editing by, inter alia, facilitating the selection of an appropriate guide RNA and base editor combination which are capable of conducting base editing at a certain level of efficiency and specificity on a given input target DNA sequence desired to be edited to produce an outcome genotype of interest. The disclosure also provides base editors (e.g., ABEs and CBEs), napDNAbps, cytidine deaminases, adenosine deaminases, nucleic acid sequences encoding base editors and components thereof, vectors, and cells. In addition, the disclosure provides methods of making biological or experimental training and/or validation data for training and/or validating the machine learning computational models, as well as, vectors, libraries, and nucleic acid sequences for use in obtaining said experimental training and/or validation data.