Generative Thermodynamics Neural Networks for Binding Affinity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for predicting binding affinities between query compounds and target proteins are inefficient, require excessive computational resources, and lack operational flexibility due to reliance on extensive training data and physics-based algorithms.

Innovation Solution

A generative thermodynamics neural network is trained using a base-to-energy distribution transformation process to predict binding conformations and metrics without the need for ground truth data, allowing for efficient and flexible prediction across various domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning systems use large volumes of training data to improve prediction accuracy, then accuracy is improved, but computational expense and data requirements increase

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent transforms the training objective from predicting binding affinities directly to learning the energy distribution parameters (mean and standard deviation) that characterize binding energy landscapes. This parameter transformation allows the model to generalize from fewer examples by learning the statistical properties of energy distributions rather than memorizing specific binding outcomes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional supervised learning mechanics with a physics-informed approach that uses thermodynamic principles (Boltzmann distribution) to generate training targets. Instead of requiring extensive labeled binding affinity data, the system uses energy distributions derived from molecular dynamics simulations to train the neural network, significantly reducing data requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If physics-based algorithms are used to predict binding affinities, then theoretical accuracy is improved, but computational resources and time increase

Engineering Contradiction:
Improvebinding affinity prediction accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent segments the binding prediction problem into two distinct components: (1) learning the energy distribution parameters through a neural network, and (2) using these parameters to generate binding affinity predictions via Boltzmann sampling. This segmentation allows the computationally intensive energy calculation to be performed once during training, while predictions can be generated efficiently by sampling from the learned distribution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary computation of energy distributions during the training phase, storing the learned mean and standard deviation parameters. When making predictions on new compounds, the system only needs to perform lightweight sampling operations rather than full physics-based simulations, dramatically reducing computational resources required for prediction.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If conventional systems rely on extensive training data to achieve accurate predictions, then prediction reliability is improved, but operational flexibility and adaptability to new domains decrease

Engineering Contradiction:
Improveprediction reliabilityVSAvoidoperational flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal energy distribution model that can be applied across different protein targets and compound types. By learning to predict energy distribution parameters rather than target-specific binding affinities, the model gains the ability to generalize to untested domains and new protein targets without requiring extensive retraining on domain-specific data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If machine learning models are trained with conventional approaches, then model performance is improved, but training time and operational efficiency worsen

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent changes the training target from continuous binding affinity values to discrete energy distribution parameters (mean and standard deviation). This parameter transformation simplifies the learning task, allowing the model to achieve good performance with fewer training iterations and less computational time, as it only needs to learn two parameters per example rather than predicting complex binding outcomes.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system reduces computational expense and enhances operational flexibility by accurately predicting binding affinities using energy distributions, enabling accurate predictions in untested data domains without requiring extensive training data.

Implementation Method 1

map the query compound from a known distribution (e.g., a Gaussian distribution) to an energy distribution (e.g., a Boltzmann distribution) and predict a binding conformation for the query compound and a target protein

Methodology Applied
Scientific EffectBoltzmann distribution:

Data Source

PatentUS20260066039A1Training and using generative thermodynamics neural networks to determine binding affinities from energy distributions
Publication Date: 2026.03.05 RECURSION PHARMACEUTICALS INC
  • US20260066039A1 patent drawing
  • US20260066039A1 patent drawing
  • US20260066039A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for training and utilizing generative thermodynamics neural networks to utilize an energy-to-base distribution transformation process to determine a binding conformation for a query compound and a target protein. For example, the disclosed systems can sample a conformation of the query compound from a known distribution and utilize the base-to-energy distribution transformation process to map the compound from the known distribution to a binding conformation. Moreover, the disclosed systems can determine an energy value associated with the binding conformation. In some bases, the disclosed systems can utilize an energy-to-base distribution transformation process to determine a binding metric for the binding conformation.