Graph-Based Random Access for Combinatorial Synthesis Libraries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual high throughput screening (vHTS) methods struggle with navigating ultra-large combinatorial synthesis libraries due to challenges in query-based random access, chemical validity, and synthetic accessibility, particularly in non-enumerative settings, where existing generative models face issues with long autoregressive chains and computational inefficiencies.

Innovation Solution

A graph-based generative model that utilizes a molecular encoder and reaction/synthon query generators to efficiently navigate and retrieve molecules from ultra-large combinatorial synthesis libraries, ensuring chemical validity and synthetic accessibility, by learning a hierarchy of keys and using minimal autoregression for parallelization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit enumeration of compounds is used for virtual screening, then screening accuracy is improved, but computational complexity and time consumption increase linearly with library size

Engineering Contradiction:
Improvescreening accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the compound library into modular building blocks and reaction templates, allowing the system to navigate the chemical space through combinatorial assembly rather than exhaustive enumeration. This segmentation enables efficient query-based access to specific molecular structures without processing the entire library.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary indexing system that maps queries to building blocks and reaction templates. This intermediary layer enables indirect access to compound structures through modular components, avoiding the need to explicitly enumerate and store all possible compounds in the library.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If query-based random access is implemented in non-enumerative libraries, then access efficiency is improved, but chemical validity and synthetic accessibility become harder to ensure

Engineering Contradiction:
Improveaccess efficiencyVSAvoidchemical validity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary validation of building blocks and reaction templates before they are used in library construction. By pre-verifying the chemical validity and synthetic accessibility of modular components and their combination rules, the system ensures that all generated structures are chemically sound without requiring post-hoc validation during query processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service mechanisms where the modular building block system inherently ensures chemical validity through its design. The predefined building blocks and reaction templates are constructed to automatically satisfy chemical validity constraints, eliminating the need for external validation systems.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If generative models use long autoregressive chains to navigate large molecular graphs, then coverage of chemical space is improved, but computational time and resource requirements increase significantly

Engineering Contradiction:
Improvecoverage of chemical spaceVSAvoidcomputational time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments molecular graphs into smaller building blocks that can be independently processed and recombined. This segmentation allows the system to navigate large chemical spaces by assembling validated modular components rather than generating entire molecular graphs through long autoregressive chains, significantly reducing computational time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by generating only the necessary portions of molecular structures through combinatorial assembly of building blocks, rather than generating complete molecular graphs from scratch. This approach achieves sufficient coverage of chemical space with minimal computational effort by focusing on assembling pre-validated fragments.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250316344A1Systems and method for query-based random access into virtual chemical combinatorial synthesis libraries
Publication Date: 2025.10.09 ATOMWISE INC
  • US20250316344A1 patent drawing
  • US20250316344A1 patent drawing
  • US20250316344A1 patent drawing

AI summary

Systems and methods for querying a combinatorial synthesis library comprising a plurality of compounds and representing a plurality of reaction types, where each reaction type maps to a plurality of reactants, and each reactant maps to a plurality of synthons, accepts a query in the form of a single graph into a molecular encoder model, thereby obtaining a query vector. The query vector is inputted into a reaction query generator model thereby obtaining a first reaction type and a first plurality of reactants. A synthon is determined for each reactant by inputting the reactant into a synthon query generator model. A set of synthons is therefore determined, each corresponding to a reactant in the first plurality of reactants. A molecular structure in the combinatorial synthesis library is identified that includes the set of synthons arranged in accordance with a synthesis rule associated with the first reaction type.