Aptamer Design via Reinforcement Learning Fine-Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying aptamers with high binding affinity and specificity to molecular targets are resource-intensive, time-consuming, and often result in suboptimal selections due to the impracticality of testing a septillion potential sequences and the limitations of experimental validation techniques.
Innovation Solution
A computer-implemented method using reinforcement learning from experimental feedback (RLEF) to fine-tune a decoder model for generating novel aptamer sequences, leveraging a generative language model pre-trained on non-coding RNA sequences, and iteratively optimizing the model parameters to achieve desired binding characteristics through experimental validation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a full experimental evaluation of all possible aptamer sequences is conducted, then the binding affinity and specificity can be optimized, but the resources and time required become prohibitively enormous
Solution Approach 1:
The patent applies preliminary action by using reinforcement learning to pre-screen and prioritize aptamer sequences before experimental evaluation. The RL agent learns from simulated binding data to identify high-priority candidates, so that only the most promising sequences undergo costly wet-lab validation. This preliminary computational filtering dramatically reduces the number of sequences requiring experimental testing while maintaining high selection accuracy.
Solution Approach 2:
The patent uses copying by creating computational models and simulations that replicate the binding behavior of aptamers. Instead of testing all possible sequences experimentally, the system uses in-silico models to copy and predict binding outcomes, allowing virtual screening of the full sequence space. Only the top predicted candidates are then validated with physical experiments, significantly reducing experimental burden.
2Reliability
If the number of aptamer sequences tested experimentally is increased, then the likelihood of finding optimal binders improves, but the scalability and cost of the process deteriorates
Solution Approach 1:
The patent implements feedback through a closed-loop system where experimental binding data from high-priority sequences is fed back into the reinforcement learning model to refine its predictions. The RL agent continuously learns from actual experimental outcomes, adjusting its sequence prioritization strategy. This feedback mechanism ensures that the system becomes increasingly accurate over time, maintaining high binding quality while improving discovery efficiency with each iteration.
Solution Approach 2:
The patent applies dynamics by making the sequence prioritization strategy adaptive and evolving. Rather than using a static filtering approach, the reinforcement learning model dynamically adjusts its predictions based on accumulated learning from experimental data. The system evolves its understanding of which sequence features lead to high-affinity binders, allowing the discovery process to become progressively more efficient while maintaining rigorous quality standards.
3Manufacturing precision
If traditional SELEX processes are used to identify high-affinity aptamers, then selective binding can be achieved, but the process becomes resource-intensive and time-consuming
Solution Approach 1:
The patent replaces the mechanical and manual aspects of traditional SELEX with an automated reinforcement learning system. Instead of manual library design and iterative selection rounds, the RL agent automatically prioritizes sequences and guides the selection process. This substitution of computational automation for manual experimental procedures reduces both the time required and the operational complexity while maintaining the ability to identify high-specificity binders.
Solution Approach 2:
The patent introduces an intermediary computational layer between the aptamer library and experimental validation. The reinforcement learning model acts as an intermediary that translates the vast sequence space into a manageable set of high-priority candidates. This intermediary filtering step simplifies the overall process by reducing the burden on experimental workflows while preserving the ability to discover highly specific binders through targeted validation.
Data Source
AI summary
The present disclosure relates to a closed loop aptamer development system that leverages in vitro experiments and in silico computation and artificial intelligence-based techniques to iteratively improve a process for identifying binders that can bind a molecular target. Particularly, aspects of the present disclosure are directed to obtaining, using an experimental assay, experimental data for a set of aptamers. The experimental data includes multiple pairs of data, each pair of data having: (i) an aptamer sequence for an aptamer from a set of aptamers, and (ii) a measurement for the characteristic of the aptamer with respect to a given target. A reward model is fine-tuned, using the experimental data, to predict a function-approximation metric for the characteristic of each aptamer in the set of aptamers. A decoder model is fine-tuned for generating novel aptamer sequences based on the function-approximation metric generated by the reward model for the novel aptamer sequences.


