Iterative Molecular Subset Selection for QSAR Prediction Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting the properties of chemical molecules, particularly toxicity, are inefficient and prone to errors due to the reliance on databases with molecules that are too different from the substance being predicted, leading to incorrect predictions.

Innovation Solution

An iterative process for selecting a subset of reference molecules based on overall similarity measures, using a molecule descriptor to evaluate and update the similarity between molecules, allowing for a more accurate and adaptive prediction of molecular properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If QSAR is applied directly to the entire database, then the prediction process is simple, but the prediction accuracy deteriorates because the database contains molecules too different from the target molecular substance

Engineering Contradiction:
Improvesimplicity of prediction processVSAvoidprediction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the entire database into multiple subsets based on structural similarity to the target molecule. Instead of applying QSAR to the whole database at once, it divides molecules into groups using structural keys (MACCS fingerprints) and selects only those subsets with sufficient structural similarity, thereby improving prediction accuracy while maintaining operational simplicity through automated segmentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by treating different regions of the database differently based on their structural similarity to the target molecule. Molecules with high structural similarity (above a threshold) are selected for QSAR analysis, while dissimilar molecules are excluded. This localized selection ensures that QSAR is applied only to relevant molecular subsets, improving prediction reliability.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If the database is filtered to include only structurally similar molecules, then prediction accuracy improves, but the complexity of the prediction process increases due to additional similarity search steps

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomplexity of prediction process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-calculating structural keys (MACCS fingerprints) for all molecules in the database before the actual QSAR prediction. This preprocessing step creates ready-to-use structural representations that enable rapid similarity comparison during the prediction phase, reducing the computational complexity burden at the time of prediction while maintaining high accuracy through pre-filtered molecular subsets.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces structural keys (MACCS fingerprints) as an intermediary between the target molecule and the database molecules. These binary vectors serve as mediators that enable efficient similarity comparison without requiring direct complex structural analysis. The intermediary representation simplifies the similarity search process while preserving the essential structural information needed for accurate prediction.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a strict similarity threshold is used to select reference molecules, then prediction reliability improves, but the number of available reference molecules decreases

Engineering Contradiction:
Improveprediction reliabilityVSAvoidnumber of reference molecules
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies dynamics by making the similarity threshold adaptable rather than fixed. The threshold can be adjusted based on the specific prediction context, target molecule characteristics, and database composition. This dynamic approach allows the system to optimize between reliability and the number of reference molecules available, selecting a threshold that ensures sufficient similarity while maintaining an adequate sample size for robust QSAR analysis.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230154571A1Method and device for selecting a subassembly of molecules for use in predicting at least one property of a molecular structure
Publication Date: 2023.05.18 ARIANEGRP SAS
  • US20230154571A1 patent drawing
  • US20230154571A1 patent drawing

AI summary

A selection process is iterative and includes an initialization associating with a so-called current molecule a value of a predetermined molecule descriptor associated with the target molecular structure, and during each iteration of the selection process, the process includes evaluating, for each molecule of a database including a plurality of molecules each associated with a value of the descriptor, a so-called overall similarity measure between the value of the descriptor associated with the molecule and the value of the descriptor associated with the current molecule; selecting molecules from the database having an overall similarity measure greater than a predetermined threshold, the selected molecules being added to the reference subset; and updating the value of the descriptor associated with the current molecule from the values of the descriptors associated with at least some of the molecules belonging to the reference subset.