Nucleic Acid Library Production via Machine Learning Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for producing nucleic acid libraries encoding desired proteins are limited by the need for clear positive mutants, which are often not obtained through biopanning operations, leading to small-scale second libraries due to cost and accuracy constraints.

Innovation Solution

The method involves calculating the estimated binding strength using sequence data from sublibraries at various stages, particularly at the target-binding sequence elution stage, and combining this with machine learning to predict mutant sequences. Additionally, degenerate codon design is used to construct a secondary library that includes sequences similar to those predicted by machine learning, thereby expanding the library size while maintaining cost-effectiveness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning is performed using data from after E. coli infection or phage amplification in the phage display method, then sequences with high enrichment rate can be obtained, but bias selection occurs depending on infection and amplification processes, so sequences with higher enrichment rate do not necessarily have improved desirable function

Engineering Contradiction:
Improveaccuracy of function predictionVSAvoidreliability of enrichment rate as function indicator
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts only the essential selection pressure signal by using data from the target-binding sequence elution stage, removing the bias introduced by subsequent E. coli infection and phage amplification processes. This extraction of pure selection signal allows accurate correlation between enrichment rate and actual binding function without contamination from procedural biases.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary measurement step - actual binding function measurement of selected mutants - to validate and calibrate the machine learning model. This intermediary data serves as a bridge between the enrichment rate data and the desired function prediction, ensuring the model learns accurate structure-function relationships rather than procedural artifacts.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If a second library is constructed based on machine learning prediction results, then the number of sequences to be evaluated can be limited in terms of cost, but the scale of the second library remains small due to accuracy constraints of training data

Engineering Contradiction:
Improvelibrary sizeVSAvoidaccuracy of training data
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by conducting actual binding function measurements on a representative subset of mutants from the first library before constructing the second library. These preliminary measurements generate high-quality training data that captures the true structure-function relationship, enabling the machine learning model to accurately predict which sequences should be included in the expanded second library.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses machine learning to copy the essential functional characteristics of measured mutants into predictions for unmeasured sequences. The model learns from the measured data and generates accurate predictions for a much larger set of sequences, effectively copying the functional information from a small measured set to a large predicted set, thereby expanding library size while maintaining accuracy.

Inventive Principle:
Principle #26Copying

3Measurement precision

If direct association data set is used for machine learning, then high-quality data with direct measurement of function and physical property values can be obtained, but the data set size is limited to several tens to several hundreds and searchable sequence is limited

Engineering Contradiction:
Improvequality of function dataVSAvoiddata set size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the data collection process into two stages: first, obtain high-quality direct measurement data from a small number of mutants; second, use this segmented high-quality data to train a machine learning model that can then evaluate a much larger segmented set of predicted sequences. This segmentation allows the high precision of direct measurement to inform a large-scale screening effort.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal machine learning model trained on high-quality direct measurement data that can universally predict the function of any sequence in the searchable database. The model serves multiple functions: it evaluates predicted sequences, guides library construction, and enables accurate function prediction across the entire sequence space, not just the measured subset.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250182853A1Method for producing library by machine learning
Publication Date: 2025.06.05 TOHOKU UNIV
  • US20250182853A1 patent drawing
  • US20250182853A1 patent drawing
  • US20250182853A1 patent drawing

AI summary

A method for producing a nucleic acid library. The method includes: preparing, by a phage display method, a first library composed of mutants obtained by randomly introducing a mutation into a nucleic acid sequence encoding a protein bound to or configured to be bound to a target; performing biopanning on the first library and obtaining data to be used for machine learning from an obtained sublibrary; and performing machine learning using the data and obtaining a second library from the first library based on machine learning prediction. The data to be used for machine learning includes a sequence of a mutant population included in a sublibrary at a target-binding sequence elution stage, an estimated binding strength to the target, and an actual measurement value of binding of some mutants included in the mutant population to the target.