3D Spatial Modeling for DNA-Encoded Library Binding Affinity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning models for DNA-encoded libraries (DELs) face challenges in accurately predicting binding affinities due to noise and biases in data generation, with existing approaches limited to molecule-level representations, failing to effectively utilize 3-D spatial information from docked compound-target complexes.

Innovation Solution

Incorporating 3-D spatial information from docked compound-target complexes into machine learning models to improve the prediction of binding affinities, enabling the denoising of DEL count data and reducing reliance on external supervision from protein crystal structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models use only molecule-level representations, then the model complexity remains manageable, but the prediction accuracy of binding affinities deteriorates due to lack of 3-D spatial information

Engineering Contradiction:
Improveprediction accuracy of binding affinitiesVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from 2D molecule-level representations to 3D spatial representations by incorporating docked compound-target complex structures. This dimensional enhancement provides critical spatial information about binding modes, interaction geometries, and conformational arrangements that 2D representations cannot capture, thereby improving prediction accuracy despite increased model complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent segments the binding affinity prediction task into multiple components: (1) generating multiple docked poses representing different binding configurations, (2) extracting spatial features from each pose including intermolecular distances, angles, and contact points, (3) aggregating features across multiple poses, and (4) feeding processed features to the machine learning model. This segmentation manages complexity by breaking down the 3D analysis into systematic, computable steps

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If DEL experiments are conducted to generate training data, then a quantitative readout for billions of compounds is obtained, but the signal is obfuscated by noise introduced in the complicated data-generation process

Engineering Contradiction:
Improvenumber of compounds screenedVSAvoidsignal quality of DEL data
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces 3D spatial structural information from molecular docking as an intermediary that mediates between the noisy DEL count data and the underlying binding affinity signal. This intermediary provides physically meaningful constraints and contextual information that help disentangle true binding signals from experimental noise, enabling more accurate model training despite the noisy input data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary molecular docking calculations to generate predicted compound-target poses and extract spatial features before training the machine learning model. This preliminary action pre-processes the data by incorporating physical chemistry principles and structural biology knowledge, creating enriched training features that improve signal-to-noise ratio before the actual binding affinity prediction task

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240177012A1Molecular Docking-Enabled Modeling of DNA-Encoded Libraries
Publication Date: 2024.05.30 INSITRO INC
  • US20240177012A1 patent drawing
  • US20240177012A1 patent drawing
  • US20240177012A1 patent drawing

AI summary

Embodiments of the disclosure involve training machine learned models using DNA-encoded library experimental data outputs and for deploying the trained machine learned models for conducting a virtual compound screen, for performing a hit selection and analysis, or for predicting binding affinities between compounds and targets. Machine learned models are trained using one or more augmentations that selectively expand molecular representations of a training dataset. Furthermore, machine learned models are trained to account for confounding covariates, thereby improving the machine learned models' abilities to conduct a virtual screen, perform a hit selection, and to predict binding affinities.