AI Compound Library Construction for Virtual Screening Bias Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Virtual screening methods face prediction biases due to hidden biases in compound data sets, such as domain bias and noncausal bias, which hinder the identification of valuable compounds.
Innovation Solution
An artificial intelligence-based compound processing method that generates a first candidate compound with specific attributes and performs molecular docking to identify a second candidate compound, constructing a compound library that alleviates these biases by combining both, thereby enhancing the accuracy of virtual screening.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a compound data set is used for virtual screening, then screening efficiency is improved, but hidden biases (domain bias or causal bias) are introduced causing prediction errors
Solution Approach 1:
The patent extracts and removes biased compounds from the training data set by identifying and excluding compounds with specific structural characteristics (such as those containing certain functional groups or molecular patterns) that cause domain bias or causal bias, thereby creating a cleaner data set for more accurate virtual screening
Solution Approach 2:
The patent implements a feedback mechanism where the virtual screening results are used to validate and refine the compound data set. By evaluating the performance of screened compounds and comparing them with actual biological activity data, the system identifies and corrects biases in the training data set iteratively, improving both efficiency and accuracy
2Adaptability or versatility
If compound generation processing is performed to increase structural diversity, then domain bias is reduced, but computational complexity increases
Solution Approach 1:
The patent performs preliminary compound generation and diversity analysis before the main virtual screening process. By pre-generating a diverse set of candidate compounds and pre-identifying structural patterns that cause bias, the system prepares optimized training data that reduces domain bias while minimizing computational complexity during the actual screening
Solution Approach 2:
The patent changes key parameters of compound generation such as molecular weight ranges, functional group restrictions, and structural constraints to optimize structural diversity. By adjusting these parameters strategically, the system generates diverse compounds without excessive computational overhead, balancing versatility with complexity
Data Source
AI summary
An artificial intelligence-based compound processing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product relates to an artificial intelligence technology. The method includes obtaining an active compound for a target protein; performing compound generation processing on an attribute property of the active compound to obtain a first candidate compound; performing molecular docking processing on the active compound and the target protein to obtain molecular docking information respectively corresponding to a plurality of molecular conformations of the active compound; screening the plurality of molecular conformations based on the molecular docking information respectively to identify a second candidate compound corresponding to the active compound; and constructing a compound library for the target protein based on the first candidate compound and the second candidate compound.


