Molecule Design Selection Using Multi-Objective Bayesian Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of designing molecules, particularly proteins, involves a vast and sparsely populated combinatorial search space, making brute force approaches computationally expensive and inefficient, with limited wet lab resources further constraining in vitro and in vivo assessments, leading to suboptimal candidate selection.
Innovation Solution
A multi-objective active learning technique using property computational models and a selection engine for multi-objective Bayesian optimization to identify joint positive molecule designs that satisfy multiple criteria, prioritizing certain properties over others, ensuring better candidates are selected for in vitro and in vivo assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If brute force approaches are used to search the combinatorial search space, then all possible molecule designs can be evaluated, but the computational cost becomes excessively high and the process is inefficient
Solution Approach 1:
The patent applies partial action by using active learning to evaluate only a subset of the most promising molecule designs rather than exhaustively evaluating all possible designs. The selection engine identifies and evaluates only those candidates with highest expected improvement, achieving reliable candidate selection without the prohibitive computational cost of brute force enumeration of the entire combinatorial search space
Solution Approach 2:
The system uses multi-objective Bayesian optimization where the selection engine learns from previous evaluation results to automatically identify promising candidates for further assessment. The computational model improves its predictions based on accumulated data from wet lab resources, enabling the system to self-optimize the candidate selection process without requiring exhaustive brute force search
2Measurement precision
If wet lab resources are used for in vitro and in vivo assessments, then accurate molecular property data can be obtained, but the limited resources constrain the number of candidates that can be evaluated
Solution Approach 1:
The patent applies partial action by using computational property predictions to pre-screen and prioritize molecule designs before wet lab assessment. The selection engine evaluates many candidates computationally and selects only the most promising subset for expensive wet lab testing, thereby obtaining accurate molecular property data for critical candidates while maintaining high productivity by limiting wet lab resources to a manageable number of high-priority candidates
Solution Approach 2:
The patent introduces computational property prediction models as an intermediary between candidate generation and wet lab assessment. These models serve as a filter that identifies promising candidates from the combinatorial search space, enabling the system to allocate limited wet lab resources efficiently to candidates most likely to succeed while maintaining measurement precision for the selected subset
3Reliability
If multi-objective optimization is applied to prioritize certain properties, then the selection of better candidates is improved, but the complexity of the selection process increases
Solution Approach 1:
The patent applies segmentation by decomposing the multi-objective optimization problem into separate objective functions for different molecular properties (e.g., binding affinity, solubility, stability). The selection engine evaluates candidates against multiple independent objectives and uses Pareto optimality concepts to identify candidates that balance multiple properties, improving selection quality while managing complexity through structured decomposition of the optimization task
Data Source
AI summary
One or more property computational models may be applied to determine a first probability of a molecule design exhibiting a first property and a second probability of the molecule design exhibiting a second property. A plurality of samples in which each sample includes a first value of the first property and a second value of the second property exhibited by the molecule design may be determined based on the output of the property computational models. A set of samples in which the first value of the first property satisfies a criterion may be identified. A utility metric corresponding to an expected improvement of the first property and the second property of the molecule design over that of baseline molecule designs may be determined based on the set of samples. One or more molecule designs may be identified as candidates for synthesis based on the corresponding utility metrics.


