Entangled Conditional Adversarial Autoencoder for Molecular Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Supervised Adversarial Autoencoders (SAAE) face challenges in generating complex molecular structures with thousands of variations, requiring numerous conditions and struggling to achieve high performance in producing novel chemical structures with specific properties.

Innovation Solution

The development of an Entangled Conditional Adversarial Autoencoder (ECAAE) that processes latent codes through reparameterization and disentanglement techniques to generate molecular structures with defined properties, such as activity against specific proteins, solubility, or ease of synthesis, by optimizing neural networks to match prior distributions and reduce mutual information between variables.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If Supervised Adversarial Autoencoders (SAAE) are used to generate molecular structures with complex conditions, then the model can generate novel chemical structures with desired properties, but the model requires a large number of complex conditions with thousands of variations and achieves poor generation performance

Engineering Contradiction:
Improvegeneration capability under complex conditionsVSAvoidgeneration performance
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies segmentation by decomposing the complex conditional generation task into independent latent factors. The latent space is divided into multiple independent dimensions, each representing a specific molecular property or condition. This allows the model to generate molecules with complex conditions by combining independent latent factors rather than processing all conditions simultaneously, thereby improving generation performance while maintaining versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism through the use of a discriminator network that mediates between the generator and the complex conditions. The discriminator evaluates generated molecules against the conditions and provides feedback, enabling the generator to learn how to satisfy complex conditions without requiring explicit processing of all condition variations, thus improving generation performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the number of conditions and variations is increased to generate diverse molecular structures, then the model can explore more chemical space, but the computational complexity and training requirements increase significantly

Engineering Contradiction:
Improvechemical space explorationVSAvoidmodel architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex condition space into independent latent factors, allowing the model to represent diverse molecular structures through combinations of a limited number of independent factors. This segmentation reduces the effective complexity of the model by avoiding the need to explicitly model all possible condition variations, while still enabling exploration of extensive chemical space through factor combination.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal latent representation that can generate molecules across multiple conditions and chemical spaces using the same model architecture. The latent factors are designed to be multi-functional, where each factor can contribute to different molecular properties depending on its activation state, allowing a single model to handle diverse generation tasks without requiring condition-specific models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If traditional machine learning models are used to estimate compound properties, then the models can guide the drug optimization process, but the hit rate of new drug candidates remains limited

Engineering Contradiction:
Improvehit rate of new drug candidatesVSAvoidaccuracy of property estimation
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent incorporates feedback mechanisms through the adversarial training process, where the discriminator provides continuous feedback to the generator about how well generated molecules satisfy the desired conditions. This feedback loop enables the model to iteratively improve its property estimation accuracy and generate higher-quality candidates, thereby improving both the hit rate and reliability of drug discovery predictions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces traditional mechanical property estimation methods with a neural-based generative model that learns implicit representations of molecular properties. Instead of using explicit rules or simple regression models, the system uses deep neural networks to capture complex relationships between molecular structures and properties, significantly improving both accuracy and productivity in drug candidate identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20230331723A1Entangled conditional adversarial autoencoder for drug discovery
Publication Date: 2023.10.19 INSILICO MEDICINE IP LTD
  • US20230331723A1 patent drawing
  • US20230331723A1 patent drawing
  • US20230331723A1 patent drawing

AI summary

A method is provided for generating new objects having given properties, such as a specific bioactivity (e.g., binding with a specific protein). In some aspects, the method can include: (a) receiving objects (e.g., physical structures) and their properties (e.g., chemical properties, bioactivity properties, etc.) from a dataset; (b) providing the objects and their properties to a machine learning platform, wherein the machine learning platform outputs a trained model; and (c) the machine learning platform takes the trained model and a set of properties and outputs new objects with desired properties. The new objects are different from the received objects. In some aspects, the objects are molecular structures, such as potential active agents, such as small molecule drugs, biological agents, nucleic acids, proteins, antibodies, or other active agents with a desired or defined bioactivity (e.g., binding a specific protein, preferentially over other proteins).