AI Compound Library Construction for Virtual Screening Bias Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual screening methods face prediction biases due to hidden biases in compound data sets, such as domain bias and noncausal bias, which hinder the identification of valuable compounds.

Innovation Solution

An artificial intelligence-based compound processing method that generates a first candidate compound with specific attributes and performs molecular docking to identify a second candidate compound, constructing a compound library that alleviates these biases by combining both, thereby enhancing the accuracy of virtual screening.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a compound data set is used for virtual screening, then screening efficiency is improved, but hidden biases (domain bias or causal bias) are introduced causing prediction errors

Engineering Contradiction:
Improvescreening efficiencyVSAvoidprediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts and removes biased compounds from the training data set by identifying and excluding compounds with specific structural characteristics (such as those containing certain functional groups or molecular patterns) that cause domain bias or causal bias, thereby creating a cleaner data set for more accurate virtual screening

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements a feedback mechanism where the virtual screening results are used to validate and refine the compound data set. By evaluating the performance of screened compounds and comparing them with actual biological activity data, the system identifies and corrects biases in the training data set iteratively, improving both efficiency and accuracy

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If compound generation processing is performed to increase structural diversity, then domain bias is reduced, but computational complexity increases

Engineering Contradiction:
Improvestructural diversityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary compound generation and diversity analysis before the main virtual screening process. By pre-generating a diverse set of candidate compounds and pre-identifying structural patterns that cause bias, the system prepares optimized training data that reduces domain bias while minimizing computational complexity during the actual screening

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes key parameters of compound generation such as molecular weight ranges, functional group restrictions, and structural constraints to optimize structural diversity. By adjusting these parameters strategically, the system generates diverse compounds without excessive computational overhead, balancing versatility with complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240055071A1Artificial intelligence-based compound processing method and apparatus, device, storage medium, and computer program product
Publication Date: 2024.02.15 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20240055071A1 patent drawing
  • US20240055071A1 patent drawing
  • US20240055071A1 patent drawing

AI summary

An artificial intelligence-based compound processing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product relates to an artificial intelligence technology. The method includes obtaining an active compound for a target protein; performing compound generation processing on an attribute property of the active compound to obtain a first candidate compound; performing molecular docking processing on the active compound and the target protein to obtain molecular docking information respectively corresponding to a plurality of molecular conformations of the active compound; screening the plurality of molecular conformations based on the molecular docking information respectively to identify a second candidate compound corresponding to the active compound; and constructing a compound library for the target protein based on the first candidate compound and the second candidate compound.