RBM Aptamer Selection Using Sparse Sequence Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning methods face challenges in predicting and generating aptamers with high binding affinity to target biomolecules due to limited training datasets, sequencing errors, and interpretability issues, particularly in the context of aptamer selection.

Innovation Solution

Utilizing a Restricted Boltzmann Machine (RBM) model trained on sequence information of aptamers with minimum threshold binding affinity, incorporating a maximum likelihood algorithm to curate datasets and enforce sparse connections, enabling the generation of candidate aptamers and classifiers that can synthesize aptamers with high binding affinity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep neural networks are used to predict aptamer binding interactions, then predictive capability is improved, but interpretability and training feasibility deteriorate due to the large number of free parameters and difficulty in identifying important features

Engineering Contradiction:
Improvebinding prediction accuracyVSAvoidinterpretability of molecular features
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts only the essential pairwise residue coupling information from sequence data using Direct Coupling Analysis, rather than using deep neural networks that process all possible features. This extraction approach identifies the most important contacts between residues while discarding redundant information, achieving both accuracy and interpretability by focusing on the critical few features that drive binding interactions

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of starting with a deep neural network and trying to interpret its complex parameters, the patent inverts the approach by directly analyzing sequence correlations to identify important features first, then building a simpler predictive model based on those identified features. This reversal prioritizes interpretability from the outset while maintaining predictive power

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If deep neural networks with many parameters are trained on biological sequence datasets, then predictive power is improved, but training difficulty and computational complexity worsen due to limited training examples and presence of errors

Engineering Contradiction:
Improvebinding interaction predictionVSAvoidtraining complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses simple statistical models (Potts model, direct coupling analysis) that are computationally inexpensive and easy to train, rather than complex deep neural networks. These simpler models can be trained quickly on limited biological sequence data without requiring extensive computational resources or large datasets, making them practical for applications with limited training examples

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent changes the parameters of the model from thousands of weights in deep neural networks to a small number of coupling parameters between residue pairs. This parameter reduction makes the model tractable for training on limited biological data while maintaining the ability to capture essential binding interactions through the statistical analysis of sequence covariation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250279163A1Methods, systems, and computer readable media for aptamer selection
Publication Date: 2025.09.04 THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
  • US20250279163A1 patent drawing
  • US20250279163A1 patent drawing
  • US20250279163A1 patent drawing

AI summary

Provided herein are methods of generating a trained classifier at least partially using a computer. The methods include training a Restricted Boltzmann Machine (RBM) using at least a first training dataset that comprises sequence information corresponding to a population of aptamers, and/or one or more descriptors thereof, which aptamers comprise a minimum threshold binding affinity to a target biomolecule to produce a trained RBM model. Additional methods as well as related systems and computer readable media are also provided.