Machine Learning Models for DNA-Encoded Library Signal Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional DNA encoded library (DEL) experiments suffer from low signal-to-noise ratio due to experimental noise and biases, which negatively impact the performance of machine learning models used for drug discovery by complicating the identification of compounds that bind to protein targets.

Innovation Solution

The development of machine learning models, including classification and regression models, that selectively expand molecular representations and account for confounding sources of noise and biases, enabling improved accuracy in virtual compound screens, hit selection, and binding affinity predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional DEL experiments are used to generate training data, then large-scale compound screening is enabled, but the signal-to-noise ratio decreases due to experimental noise and biases

Engineering Contradiction:
Improvenumber of compounds screenedVSAvoidsignal-to-noise ratio
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts and removes confounding factors (noise and biases) from the DEL experimental data through computational methods. Specifically, it separates the true binding signal from experimental artifacts by identifying and eliminating sources of noise such as non-specific binding signals and sequence composition biases, thereby improving the signal-to-noise ratio while preserving the large-scale screening capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces computational processing as an intermediary step between raw DEL experimental data and machine learning model training. This intermediary processing layer applies algorithms to correct biases and filter noise, transforming the raw low-quality data into cleaned, high-quality training data that maintains the original large-scale compound information while removing detrimental experimental artifacts

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If DEL experimental data with noise and biases is used for training, then model training can proceed with available data, but model performance deteriorates

Engineering Contradiction:
Improvemodel training feasibilityVSAvoidmodel performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent converts the harmful effect of noise and biases in DEL data into a benefit by using these confounding factors as explicit training targets. The machine learning model is trained to recognize and account for these biases, learning to distinguish true binding signals from experimental artifacts. This approach transforms previously harmful noise into useful information that improves model robustness and performance

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent changes the parameters and features used in model training by incorporating corrected and normalized DEL data parameters. It transforms raw sequencing counts into bias-corrected binding affinity estimates by adjusting for factors such as library composition, experimental conditions, and detection biases, thereby improving model reliability while maintaining training feasibility

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230130619A1Machine learning pipeline using DNA-encoded library selections
Publication Date: 2023.04.27 INSITRO INC
  • US20230130619A1 patent drawing
  • US20230130619A1 patent drawing
  • US20230130619A1 patent drawing

AI summary

Embodiments of the disclosure involve training machine learned models using DNA-encoded library experimental data outputs and for deploying the trained machine learned models for conducting a virtual compound screen, for performing a hit selection and analysis, or for predicting binding affinities between compounds and targets. Machine learned models are trained using one or more augmentations that selectively expand molecular representations of a training dataset. Furthermore, machine learned models are trained to account for confounding covariates, thereby improving the machine learned models' abilities to conduct a virtual screen, perform a hit selection, and to predict binding affinities.