Machine Learning Model for Chemical Space Exploration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing drug discovery processes are time-consuming and expensive due to the computational infeasibility of evaluating vast chemical libraries for identifying top hit molecules with high binding affinity against drug targets.
Innovation Solution
A processor-implemented method using a machine learning model to explore the chemical space by vectorizing molecules, clustering, sampling, determining docking scores, training a Gaussian process model, and computing acquisition functions to identify top hit molecules efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If computational techniques are used to evaluate each molecule in chemical libraries to identify hit molecules with high binding affinity, then the accuracy of identifying top hit molecules is improved, but the computation time and cost increase significantly making the process computationally infeasible
Solution Approach 1:
The patent divides the vast chemical library into smaller manageable subsets or batches. Instead of evaluating all molecules simultaneously, the system processes molecules in segments, allowing computational resources to be efficiently utilized and reducing the overall computation time while maintaining evaluation accuracy for each segment.
Solution Approach 2:
The patent employs preliminary filtering techniques to pre-assess molecules based on simpler criteria before applying full computational docking simulations. This preliminary action eliminates obviously unsuitable candidates early in the process, reducing the number of molecules that require intensive computational evaluation and thus reducing total computation time.
2Reliability
If computational techniques are used to evaluate each molecule in chemical libraries to identify top hit molecules, then the quality of drug discovery results is improved, but the cost of the process increases significantly
Solution Approach 1:
The patent applies partial evaluation by assessing only a subset of molecules with full computational docking simulations, while using faster screening methods for the remainder. This partial action approach maintains high reliability for the evaluated subset while reducing overall computational cost, achieving a balance between quality and expense.
Solution Approach 2:
The patent uses simplified molecular representations or surrogate models as copies of the full molecular structures for initial screening. These copies allow rapid assessment of potential candidates without the computational expense of evaluating the complete molecular structures, reducing costs while maintaining the ability to identify promising candidates for full evaluation.
3Loss of information
If the entire chemical library is evaluated to explore comprehensive chemical space, then the completeness of molecular exploration is improved, but the computational feasibility deteriorates
Solution Approach 1:
The patent segments the chemical space exploration into multiple phases or stages, each evaluating a different subset of molecules with varying levels of computational depth. This segmentation allows comprehensive coverage of chemical space over time while maintaining computational efficiency at each stage, preventing resource exhaustion.
Solution Approach 2:
The patent implements periodic re-evaluation of chemical space using updated computational models and newly synthesized data. Instead of attempting to evaluate everything at once, the system periodically returns to previously unexplored or partially explored regions of chemical space, maintaining completeness while managing computational load through time-distributed processing.
Data Source
AI summary
A system and method for exploring a chemical space during molecular design for at least one top hit molecule using a machine learning (ML) model are provided. The method includes (i) representing the at least one molecule stored in a drug library into at least one vector; (ii) clustering the at least one vector to obtain at least one cluster of molecules into one or more clusters; (iii) uniformly sampling a first subset of molecules from each cluster of molecules; (vi) determining a docking score for sampled subset of molecules; (iv) training the ML model by correlating sampled subset of molecules with docking score; (viii) computing acquisition function values for a second subset of molecules from each cluster; and (ix) determining at least one top hit molecule based on the computed acquisition function values, thereby exploring the chemical space for the at least one top hit molecule.


