Molecular Property Enhancement Using Structure-Informed ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for enhancing molecular properties, particularly in small and large molecule therapeutics, face challenges in efficiently improving properties such as binding affinity, specificity, and developability, especially in the context of drug design, where traditional approaches are limited in their ability to generate molecules with superior properties compared to input molecules.

Innovation Solution

A machine learning-based technique utilizing a molecule design computation model that encodes and decodes molecular representations to generate output molecules with enhanced properties, trained on datasets of molecule pairs with different property values, allowing for the generation of molecules with compositional and conformational modifications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional molecular design methods are used, then the process is simple and easy to understand, but the ability to generate molecules with superior properties is limited

Engineering Contradiction:
Improveability to generate molecules with superior propertiesVSAvoidcomplexity of molecule design system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a molecule design computation model as an intermediary between traditional design methods and molecular property enhancement. This model, trained on matched datasets of molecule pairs, serves as a mediator that learns complex structure-property relationships and generates molecules with improved properties without requiring direct complex experimental trial-and-error

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the molecule design problem into a parameter optimization problem by training the computation model on matched datasets where molecules are represented by structural parameters and properties. The model learns to map structural parameters to property values, enabling systematic generation of molecules with superior properties through parameter optimization rather than traditional trial-and-error

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If machine learning models are trained on large datasets to improve property prediction accuracy, then the prediction precision improves, but the training time and computational resources increase

Engineering Contradiction:
Improveproperty prediction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training the molecule design computation model on comprehensive matched datasets before actual molecule design tasks. The model is pre-trained to learn general structure-property relationships from diverse molecule pairs, so that when deployed for specific property enhancement tasks, it can quickly generate accurate predictions without requiring extensive re-training

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses partial action by training the model on matched datasets that focus specifically on the property of interest rather than attempting to model all molecular properties simultaneously. This selective training approach achieves high prediction accuracy for the target property while reducing overall training complexity and time requirements

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250364073A1Structure-informed machine learning enabled enhancement of molecular properties
Publication Date: 2025.11.27 GENENTECH INC
  • US20250364073A1 patent drawing
  • US20250364073A1 patent drawing
  • US20250364073A1 patent drawing

AI summary

An input molecule exhibiting a value for one or more properties may be identified. A molecule design computation model may be applied to generate one or more output molecule exhibiting a different value for the one or more properties than the input molecule. The molecule design computation model may generate the one or more output molecules by at least encoding the input molecule to generate an embedding of the input molecule, and decoding the embedding of the input molecule to generate the one or more output molecules. In some cases, the molecule design computation model may generate the one or more output molecules by denoising an input molecule while conditioned on the input molecule. In some cases, the molecule design computation model may operate on a joint representation of the input molecule that combines a linear and a three-dimensional representation of the input molecule.