Active Learning Selection Model for Compound Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The generation of accurate and reliable machine learning (ML) models for predicting compound properties is hindered by the shortage of labeled training data, leading to costly, time-consuming, and error-prone processes, especially when dealing with multiple properties.
Innovation Solution
A selection model is developed using reinforcement learning (RL) to iteratively select and validate a shortlist of compounds from prediction results, enhancing the training dataset and improving the predictive performance of ML models by determining the best validation methods, such as laboratory experimentation or computer analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If more compounds are selected for validation to improve model accuracy, then the quality of the property model improves, but the cost and time required for validation increases
Solution Approach 1:
The patent extracts only the most valuable compounds for validation from the full prediction result list. The selection model identifies and extracts a shortlist of compounds that will provide the maximum information gain for model improvement, rather than validating all compounds. This resolves the contradiction by taking out only the essential elements needed for model enhancement.
Solution Approach 2:
The patent applies partial action by validating only a subset of compounds rather than all compounds. The selection model determines the optimal partial set of compounds that will yield the best model improvement per unit of validation cost, achieving sufficient model accuracy without the excessive time and resource investment of validating all compounds.
2Reliability
If manual generation of labelled training datasets is performed to improve model quality, then the accuracy of property models improves, but the process becomes costly, time-consuming and error-prone
Solution Approach 1:
The system performs self-service by automatically generating and curating labelled training datasets through the selection model. The model autonomously identifies compounds for validation, determines appropriate validation methods, and integrates results back into the training dataset without requiring manual intervention at each step, thereby improving productivity while maintaining model quality.
Solution Approach 2:
The patent implements feedback loops where validation results of selected compounds are automatically fed back into the training dataset, which then retrains the property model. This automated feedback mechanism eliminates manual data generation steps, reduces errors, and improves both the quality and efficiency of model development by continuously iterating with validated data.
3Adaptability or versatility
If a larger number of properties are predicted to improve comprehensiveness, then the coverage of compound characterization improves, but the complexity of generating labelled training datasets increases exponentially
Solution Approach 1:
The patent segments the complex task of multi-property prediction into manageable components. The selection model handles multiple properties simultaneously by identifying compounds that are most valuable for validating multiple properties at once, breaking down the exponential complexity into linearly scalable selections based on information gain across all properties.
Solution Approach 2:
The selection model serves multiple functions: it selects compounds for validation, determines appropriate validation methods, and prioritizes compounds based on their value across multiple properties simultaneously. This multi-functional approach allows the system to handle comprehensive property prediction without the exponential complexity increase that would result from separate processing of each property.
Data Source
AI summary
Method(s) and apparatus are provided for generating a selection model based on a machine learning (ML) technique, the selection model for selecting a shortlist of compounds requiring validation with a particular property. An iterative procedure or feedback loop for generating the selection model may include: receiving a prediction result list output from a property model for predicting whether a plurality of compounds are associated with a particular property and an property model score; retraining the selection model based on the property model score and/or the prediction result list; selecting a shortlist of compounds using the retrained selection model from the plurality of compounds associated with the prediction result list; sending the selected shortlist of compounds for validation with the particular property, where another ML technique is used to update the property model based on the validation; repeating the receiving and retraining of the selection model until determining the selection model has been validly trained.


