Molecule Generation Model Using Difference Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying molecules with specific properties in drug design are inefficient due to reliance on virtual screening techniques, which struggle to effectively navigate the large chemical space.

Innovation Solution

A method and apparatus for training a molecule generation model by obtaining molecular samples, determining molecular difference information, and training encoding and generation modules to generate optimized molecules with specific properties, utilizing techniques such as multilayer perceptrons and transformer models for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If virtual screening technique is used to identify molecules with specific properties, then the process can be automated, but the precision of molecule generation is insufficient due to the large chemical space

Engineering Contradiction:
Improveprecision of molecule generationVSAvoidcomplexity of navigating chemical space
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the molecule generation process into distinct modules: an encoding module that processes input molecules into latent representations, and a generation module that synthesizes new molecules. This segmentation allows each module to specialize in specific tasks, improving overall precision while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a latent space dimension between the input chemical space and output molecule space. By projecting molecules into this intermediate latent representation space and performing operations there, the system can efficiently navigate chemical space and generate precise molecular structures that would be difficult to obtain through direct screening.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If traditional virtual screening is used, then existing molecules can be evaluated, but the productivity of discovering new molecules with desired properties is low

Engineering Contradiction:
Improveefficiency of drug design processVSAvoidability to generate molecules with specific properties
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a self-service mechanism where the model generates its own training data through iterative optimization. The system uses the trained modules to automatically generate new molecules with desired properties, which are then used to further refine the model, creating a self-improving cycle that increases both productivity and reliability without requiring extensive manual intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback loops where the generation module's output is evaluated against target properties, and this feedback is used to adjust the latent representations and subsequent generation steps. This feedback mechanism ensures that the generated molecules progressively converge toward the desired properties, improving reliability while maintaining high productivity through automated iterative optimization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230115984A1Method and apparatus for training model, method and apparatus for generating molecules
Publication Date: 2023.04.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20230115984A1 patent drawing
  • US20230115984A1 patent drawing
  • US20230115984A1 patent drawing

AI summary

The present disclosure provides a method for training a model, a method and an apparatus for generating molecules, and relates to the technical field of computer technology, particularly the technical field of artificial intelligence. The particular implementation may include: obtaining first molecular samples and second molecular samples; determining molecular difference information based on the first molecular samples and the second molecular samples; training an initial encoding module and an initial generation module based on the molecular difference information to obtain a target encoding module and a target generation module; and determining a molecule generation model based on the target encoding module and the target generation module.