RBM Aptamer Selection Using Sparse Sequence Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning methods face challenges in predicting and generating aptamers with high binding affinity to target biomolecules due to limited training datasets, sequencing errors, and interpretability issues, particularly in the context of aptamer selection.
Innovation Solution
Utilizing a Restricted Boltzmann Machine (RBM) model trained on sequence information of aptamers with minimum threshold binding affinity, incorporating a maximum likelihood algorithm to curate datasets and enforce sparse connections, enabling the generation of candidate aptamers and classifiers that can synthesize aptamers with high binding affinity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep neural networks are used to predict aptamer binding interactions, then predictive capability is improved, but interpretability and training feasibility deteriorate due to the large number of free parameters and difficulty in identifying important features
Solution Approach 1:
The patent extracts only the essential pairwise residue coupling information from sequence data using Direct Coupling Analysis, rather than using deep neural networks that process all possible features. This extraction approach identifies the most important contacts between residues while discarding redundant information, achieving both accuracy and interpretability by focusing on the critical few features that drive binding interactions
Solution Approach 2:
Instead of starting with a deep neural network and trying to interpret its complex parameters, the patent inverts the approach by directly analyzing sequence correlations to identify important features first, then building a simpler predictive model based on those identified features. This reversal prioritizes interpretability from the outset while maintaining predictive power
2Measurement precision
If deep neural networks with many parameters are trained on biological sequence datasets, then predictive power is improved, but training difficulty and computational complexity worsen due to limited training examples and presence of errors
Solution Approach 1:
The patent uses simple statistical models (Potts model, direct coupling analysis) that are computationally inexpensive and easy to train, rather than complex deep neural networks. These simpler models can be trained quickly on limited biological sequence data without requiring extensive computational resources or large datasets, making them practical for applications with limited training examples
Solution Approach 2:
The patent changes the parameters of the model from thousands of weights in deep neural networks to a small number of coupling parameters between residue pairs. This parameter reduction makes the model tractable for training on limited biological data while maintaining the ability to capture essential binding interactions through the statistical analysis of sequence covariation
Data Source
AI summary
Provided herein are methods of generating a trained classifier at least partially using a computer. The methods include training a Restricted Boltzmann Machine (RBM) using at least a first training dataset that comprises sequence information corresponding to a population of aptamers, and/or one or more descriptors thereof, which aptamers comprise a minimum threshold binding affinity to a target biomolecule to produce a trained RBM model. Additional methods as well as related systems and computer readable media are also provided.


