Hierarchical Reinforcement Learning for Targeted Compound Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Pharmaceutical companies face challenges in efficiently identifying compounds that strongly bind to disease targets while minimizing off-target interactions and optimizing pharmacological properties, leading to high costs and failure rates in drug development due to toxicity and poor ADME profiles.

Innovation Solution

A hierarchical reinforcement learning approach is employed to generate and optimize compounds by using a parent and child model to iteratively modify initial compounds based on their interaction with a target macromolecule, updating model parameters to achieve desired properties, and testing derived compounds in a wet lab assay.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large libraries of compounds are screened to find compounds with strong binding to target macromolecules, then the probability of finding potent compounds increases, but storage constraints, shelf stability issues, and chemical costs increase significantly

Engineering Contradiction:
Improvebinding potencyVSAvoidlibrary size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system performs preliminary computational screening and hierarchical reinforcement learning optimization before physical experimentation. The parent model generates candidate compounds with predicted high binding affinity, and the child model optimizes their properties, allowing researchers to focus experimental resources on a small subset of promising candidates rather than screening entire libraries

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses computational models (parent and child neural networks) to create virtual representations and predictions of compound behavior. These digital twins allow researchers to evaluate binding affinity, selectivity, and ADME properties in silico before synthesizing or ordering physical compounds, reducing the need to handle and store large quantities of actual chemical substances

Inventive Principle:
Principle #26Copying

2Reliability

If more compounds are tested to increase the odds of finding compounds with desirable ADME profiles, then the likelihood of finding drug candidates with good pharmacological properties increases, but the cost and time needed for physical assay increases prohibitively

Engineering Contradiction:
ImproveADME profile qualityVSAvoidassay time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary computational evaluation of ADME properties using the hierarchical reinforcement learning framework before physical assays. The parent model generates compounds while the child model predicts and optimizes ADME characteristics, enabling virtual filtering of candidates based on pharmacological properties before committing to time-consuming wet lab experiments

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces mechanical/physical assay processes with computational modeling and prediction. Neural network models predict ADME properties from molecular structures, substituting physical experimentation with in silico evaluation for initial screening rounds, thereby dramatically reducing assay time and enabling evaluation of thousands of compounds computationally before selecting a small subset for physical testing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If compounds are optimized for strong binding to target macromolecules, then potency increases, but off-target side effects and toxicity increase due to non-selective binding

Engineering Contradiction:
Improvebinding potencyVSAvoidtoxicity
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system applies different optimization criteria to different aspects of compound properties. The parent model focuses on generating compounds with high binding affinity to the target macromolecule, while the child model simultaneously optimizes for selectivity against off-targets and favorable ADME profiles. This multi-objective hierarchical approach ensures that potency and selectivity are optimized independently and simultaneously

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system incorporates feedback loops where the child model evaluates predicted off-target binding and ADME properties, then feeds this information back to the parent model to refine compound generation. The hierarchical reinforcement learning framework uses reward functions that penalize off-target binding and toxicity predictions, guiding the optimization process toward compounds that are both potent and safe

Inventive Principle:
Principle #23Feedback

4Productivity

If hierarchical reinforcement learning with parent and child models is used to optimize compound generation, then compound identification efficiency increases, but model complexity and computational requirements increase

Engineering Contradiction:
Improvecompound identification efficiencyVSAvoidmodel complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the complex task of compound optimization into separate hierarchical levels: the parent model handles high-level compound generation and structural optimization, while the child model handles lower-level property prediction and fine-tuning. This segmentation allows each model to specialize in specific aspects of the optimization problem, improving overall efficiency while managing computational complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260080972A1Systems and methods for discovering compounds using hierarchical reinforcement learning
Publication Date: 2026.03.19 DEEPCURE INC
  • US20260080972A1 patent drawing
  • US20260080972A1 patent drawing
  • US20260080972A1 patent drawing

AI summary

A method for identifying derived compounds exhibiting activity for a target macromolecule generates experiences. Each experience uses an initial compound in plurality of initial compounds to construct a derived compound through a hierarchical proximal policy. The policy has a parent molecular reaction model and a child reactant model that uses an environment of the target macromolecule. The parent model evaluates a plurality of molecular reactions. The child model evaluates a corresponding plurality of reactants for a selected molecular reaction. Using the plurality of experiences, the parameters of the parent model are updated in accordance with a first surrogate objective while the parameters of the child model are updated in accordance with a second surrogate objective. The generation of derived compounds and hierarchical proximal policy updating continues until convergence. Then, a subset of the derived compounds from the experiences is tested for activity against the target macromolecule.