Molecule Design Model Using Pseudo-Matched Pairs for Property Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for enhancing molecular properties, particularly in small and large molecule therapeutics, face challenges in efficiently improving properties such as binding affinity, specificity, and developability, due to limitations in chemical synthesis and delivery methods.

Innovation Solution

A machine learning enabled system for enhancing molecular properties through iterative training with pseudo-matched molecule pairs, utilizing a molecule design computation model that encodes and decodes molecular embeddings to generate output molecules with improved properties, incorporating compositional and conformational modifications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional chemical synthesis methods are used to enhance molecular properties, then molecules can be produced, but the efficiency and effectiveness of improving properties such as binding affinity, specificity, and developability are limited

Engineering Contradiction:
Improvemolecular property enhancement precisionVSAvoidmolecule generation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces traditional chemical synthesis methods with a machine learning-based computational system. The molecule design computation model uses iterative training with pseudo-matched molecule pairs to generate optimized molecular structures, substituting wet-lab chemical processes with in-silico computational design. This enables precise control over molecular properties through algorithmic optimization rather than trial-and-error synthesis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system employs iterative training that progressively optimizes molecular parameters by generating pseudo-matched molecule pairs with improved binding affinity, specificity, and developability. Each iteration refines the molecular structure by making targeted parameter changes based on learned patterns from training data, enabling systematic enhancement of multiple properties simultaneously.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If machine learning models are trained on matched datasets to generate molecules with improved properties, then binding affinity and specificity can be enhanced, but the complexity of the training process and model architecture increases

Engineering Contradiction:
Improvebinding affinityVSAvoidmodel training complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The training process is segmented into distinct phases: initial model training on matched datasets, generation of pseudo-matched molecule pairs, and iterative refinement cycles. The molecule design computation model is divided into separate components including the encoder for processing input molecules and the decoder for generating optimized structures. This segmentation allows each component to be optimized independently while contributing to overall reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs self-service mechanisms where the model generates its own training data through pseudo-matched molecule pair creation. The iterative process automatically identifies areas for improvement and generates corresponding training examples without external intervention, enabling the model to self-optimize its performance for binding affinity and specificity while managing training complexity through autonomous learning.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250364091A1Iterative training with pseudo-matched molecule pairs for machine learning enabled enhancement of molecular properties
Publication Date: 2025.11.27 GENENTECH INC
  • US20250364091A1 patent drawing
  • US20250364091A1 patent drawing
  • US20250364091A1 patent drawing

AI summary

An input molecule exhibiting a value for one or more properties may be identified. A molecule design computation model may be applied to generate one or more output molecule exhibiting a different value for the one or more properties than the input molecule. The molecule design computation model may generate the one or more output molecules by at least encoding the input molecule to generate an embedding of the input molecule, and decoding the embedding of the input molecule to generate the one or more output molecules. In some cases, the molecule design computation model may generate the one or more output molecules by denoising an input molecule while conditioned on the input molecule. In some cases, the molecule design computation model may operate on a joint representation of the input molecule that combines a linear and a three-dimensional representation of the input molecule.