Multi-objective optimized molecular design with partially ordered, hybrid variable molecular properties
Through multi-objective active learning technology, multi-objective optimization of molecular design is used using characteristic calculation models, which solves the problems of resource bottlenecks and poor molecular design quality in the existing technology, and achieves more efficient molecular design and evaluation.
Patent Information
- Application Number
- CN202380070729.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-10-03
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively multi-objective optimization in molecular design, resulting in resource bottlenecks and poor molecular design quality during in vitro and in vivo evaluation.
Multi-objective active learning technology is used to use trained characteristic calculation models through computer program products to determine the characteristic probability of molecular design, and determine utility metrics based on the output to identify synthetic candidates.
Improves the multi-objective optimization capability of molecular design, ensuring that the selected molecular design has better molecular properties when evaluated in vitro and in vivo, reducing resource waste and the possibility of failure.
Smart Images

Figure CN120113003A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 378,186, filed on October 3, 2022, entitled “Multi-Objective Active Learning for Molecular Design,” and U.S. Provisional Application No. 63 / 385,609, filed on November 30, 2022, entitled “Multi-Objective Active Learning for Molecular Design,” the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] The subject matter described herein relates generally to molecular design and, more specifically, to multi-objective active learning techniques for molecular design. Background Art
[0004] A molecule is a group of two or more atoms linked together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of the substance. An example of a molecule is a protein molecule, while examples of non-protein molecules include small molecules, ions, nucleic acids, polysaccharides, glycolipids, etc. The functions and properties of a molecule may depend on its three-dimensional structure. For example, proteins are responsible for many important cellular functions, including, for example, enzymatic reactions, molecular transport, regulation and execution of many biological pathways, cell growth, proliferation, nutrient uptake, morphology, movement, cell-to-cell communication, etc.
[0005] Protein structure can include one or more polypeptides, which are amino acid residue chains connected together by peptide bonds. The amino acid residue sequence in the polypeptide chain forming the protein structure determines the three-dimensional structure of the protein (for example, the tertiary structure of the protein). In addition, the amino acid sequence in the polypeptide chain forming the protein determines the basic function of the protein. Therefore, a goal of protein design can include building one or more amino acid residue sequences that exhibit various desired properties. For example, in the case of macromolecular drug discovery, de novo protein design will generally seek to identify sequences of amino acid residues (for example, such as antibodies, etc.) that can be combined with target antigens (for example, viral antigens, tumor antigens, etc.), including by adopting a three-dimensional structure complementary to the three-dimensional structure of the target antigen. Summary of the invention
[0006] Systems, methods, and articles of manufacture, including computer program products, for molecular design with multi-objective active learning are provided. In one aspect, a system for molecular design with multi-objective active learning is provided. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that, when executed by the at least one data processor, cause an operation. The operations may include: applying one or more trained property computation models to a first molecular design to determine a first probability that the first molecular design exhibits a first property and a second probability that the first molecular design exhibits a second property; determining, based at least on outputs of the one or more property computation models, a first plurality of samples associated with the first molecular design, each sample in the first plurality of samples including a first value of the first property exhibited by the first molecular design and a second value of the second property exhibited by the first molecular design having the first value for the first property; identifying, among the first plurality of samples, a first group of samples in which the first value of the first property meets a first criterion; determining, based at least on the first group of samples, a first utility metric corresponding to a first expected improvement of the first property and the second property of the first molecular design relative to the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.
[0007] In another aspect, a method for molecular design with multi-objective optimization is provided. The method may include: applying one or more trained property computation models to a first molecular design to determine a first probability that the first molecular design exhibits a first property and a second probability that the first molecular design exhibits a second property; determining a first plurality of samples associated with the first molecular design based at least on the output of the one or more property computation models, each sample in the first plurality of samples including a first value of the first property exhibited by the first molecular design and a second value of the second property exhibited by the first molecular design having the first value for the first property; identifying a first group of samples in which the first value of the first property meets a first criterion; determining a first utility metric based at least on the first group of samples, the first utility metric corresponding to a first expected improvement of the first property and the second property of the first molecular design relative to the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.
[0008] In another aspect, a non-transitory computer program product for molecular design with multi-objective active learning is provided. The non-transitory computer program product may store instructions that, when executed by at least one data processor, cause operations. The operations may include: applying one or more trained property computation models to a first molecular design to determine a first probability that the first molecular design exhibits a first property and a second probability that the first molecular design exhibits a second property; determining a first plurality of samples associated with the first molecular design based at least on the output of the one or more property computation models, each sample in the first plurality of samples including a first value of the first property exhibited by the first molecular design and a second value of the second property exhibited by the first molecular design having the first value for the first property; identifying a first group of samples in which the first value of the first property meets a first criterion; determining a first utility metric based at least on the first group of samples, the first utility metric corresponding to a first expected improvement of the first property and the second property of the first molecular design relative to the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.
[0009] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may optionally be included in any feasible combination.
[0010] In some variations, the first utility metric may be determined by applying expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), or joint entropy search (JES).
[0011] In some variations, the one or more property computation models may be retrained based at least on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecular designs. The one or more retrained property computation models may be applied to determine the first property and the second property of the one or more baseline molecules.
[0012] In some variations, a second set of samples may be identified within the first plurality of samples in which the first value of the first characteristic fails to meet the first criterion.The first utility metric may be determined to include a first contribution from the first set of samples and exclude a second contribution from the second set of samples.
[0013] In some variations, one or more trained property computation models may be applied to the first molecular design to determine a third probability that the first molecular design exhibits a third property. A first plurality of samples may be determined based at least on the output of the one or more property computation models to further include as part of each sample a third value of the third property exhibited by the first molecular design. A first set of samples may be further identified based on a first value of the first property that satisfies a first criterion and a second value of the second property that satisfies a second criterion. A first utility metric may be determined based at least on the first set of samples to further correspond to a first expected improvement of the first property, the second property, and the third property of the first molecular design relative to the first property, the second property, and the third property of one or more baseline molecular designs.
[0014] In some variations, the first characteristic and the second characteristic may occupy the same level in the hierarchy above the third characteristic, such that the first molecular design is required to satisfy a first criterion associated with the first characteristic and a second criterion associated with the second characteristic before the first molecular design is evaluated for the third characteristic.
[0015] In some variations, the first characteristic and the second characteristic may occupy different levels in the hierarchy above the third characteristic, such that the first molecular design is required to satisfy a first criterion associated with the first characteristic before the first molecular design is evaluated for the second characteristic. The first molecular design may further be required to satisfy a second criterion associated with the second characteristic before the first molecular design is evaluated for the third characteristic.
[0016] In some variations, the one or more property computation models may include a first property computation model trained to determine a first probability that a first molecular design exhibits a first property.
[0017] In some variations, the first property computation model may include a first probabilistic binary classifier trained to output a first value when the first probability satisfies a second threshold, and to output a second value when the first probability fails to satisfy the second threshold. The first property computation model may further include a first probabilistic regressor trained to determine a first value of a first property exhibited by the first molecular design.
[0018] In some variations, the one or more property computation models may further include a second property computation model trained to determine a second probability that the first molecule exhibits a second property.
[0019] In some variations, the second property calculation model may include a second binary classifier trained to output a first value when the second probability meets a second threshold and to output a second value when the second probability fails to meet the second threshold. The second property calculation model may further include a second regressor trained to determine a second value of the second property exhibited by the first molecular design.
[0020] In some variations, the one or more property computation models may include an integration of property computation models. A first probability that the first molecular design exhibits a first property and / or a second probability that the first molecular design exhibits a second property may be determined based at least on an output of the integration of property computation models.
[0021] In some variations, one or more property calculations may be applied to the second molecular design to determine a third probability that the second molecular design exhibits a first property and a fourth probability that the second molecular design exhibits a second property. A second plurality of samples associated with the second molecular design may be determined based at least on the output of the one or more property calculation models. Each sample of the second plurality of samples may include a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design. A second group of samples in which the third value of the first property meets the first standard may be identified within the second plurality of samples. A second utility metric may be determined based at least on the second group of samples, the second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design relative to the first property and the second property of one or more baseline molecular designs. Based at least on the second utility metric of the second molecular design, the second molecular design may be identified as another candidate for synthesis.
[0022] In some variations, one or more baseline molecular designs may be updated to include a first molecular design such that the second expected improvement includes an expected improvement in the first and second properties of the second molecular design relative to the first and second properties of the first molecular design.
[0023] In some variations, one or more baseline molecular designs may be updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or the second property exhibited by the first molecular design.
[0024] In some variations, one or more baseline molecular designs may be updated to include an average of the first plurality of samples associated with the first molecular design.
[0025] In some variations, each of the first probability that a first molecular design exhibits a first property and / or the second probability that a first molecular design exhibits a second property may include (i) a first probability distribution over a first value indicating that the corresponding property is present in the first molecular design and a second value indicating that the corresponding property is not present in the first molecular design, and (ii) a second probability distribution over a range of possible values indicating the magnitude of the corresponding property exhibited by the first molecular design.
[0026] In some variations, the first molecular design may be identified as a candidate for synthesis based at least on the first utility metric of the first molecular design satisfying one or more thresholds.
[0027] In some variations, the N number of molecular designs with the highest utility metric may be selected as candidates for synthesis.The first molecular design may be identified as a candidate for synthesis based at least on the first molecular design being one of the N number of molecular designs with the highest utility metric.
[0028] In some variations, a first molecular design may be identified as a candidate for synthesis based at least on the presence or absence of one or more particular amino acid residues in the first molecular design.
[0029] In some variations, the first plurality of samples may include a distribution of a second value of a second property exhibited by the first molecular design over a first value of a first property exhibited by the first molecular design.
[0030] Specific implementations of the current subject matter may include, but are not limited to, methods consistent with the description provided herein and articles including tangibly embodied machine-readable media that are operable to cause one or more machines (e.g., computers, etc.) to cause operations that implement one or more of the features described. Similarly, a computer system that may include one or more processors and one or more memories coupled to the one or more processors is also described. A memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method consistent with one or more implementations of the current subject matter may be implemented by one or more data processors present in a single computing system or multiple computing systems. Such multiple computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.) via a direct connection between one or more computing systems in the multiple computing systems, etc.
[0031] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will become apparent with reference to the description and drawings, and to the claims. Although certain features of the subject matter disclosed herein are described for illustrative purposes related to the design of biological sequences of, for example, protein molecules, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the subject matter protected. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, help explain some principles associated with the disclosed embodiments.
[0033] Figure 1 depicts a system diagram illustrating an example of a molecular design system according to some exemplary embodiments;
[0034] Figure 2A depicts a flow chart illustrating an example of a molecular design process for multi-objective optimization with partially ordered, mixed variable properties, according to some exemplary embodiments;
[0035] Figure 2B depicts a flow chart illustrating another example of a molecular design process for multi-objective optimization with partially ordered, mixed variable properties according to some exemplary embodiments;
[0036] Figure 3 depicts a block diagram illustrating a molecular design pipeline with multi-objective optimization of a partially ordered, mixed-variable nature, according to some exemplary embodiments;
[0037] Figure 4 depicts a schematic diagram showing an example of a hierarchy associated with properties of a molecular design according to some exemplary embodiments;
[0038] Figure 5 depicts a diagram illustrating the effect of resampling a proxy posterior on an example of an acquisition function, according to some exemplary embodiments;
[0039] Fig. 6A depicts a flow chart illustrating an example of a process for determining a probability that a molecular design exhibits a property, according to some exemplary embodiments;
[0040] Figure 6B depicts a flow chart illustrating another example of a process for determining a probability that a molecular design exhibits a property according to some exemplary embodiments;
[0041] Figure 7 depicts a graph showing the change in the number of joint positive molecule designs over multiple active learning iterations according to some exemplary embodiments;
[0042] Figure 8 depicts a diagram illustrating a pairwise Pareto front view of an example of a penicillin production task, according to some exemplary embodiments;
[0043] Fig. 9depicts a graph showing a distribution of examples of molecular designs selected as candidates for synthesis and testing, according to some exemplary embodiments;
[0044] Fig.10 depicts a graph showing the number of joint positive molecule designs and the log posterior density on binding affinity for an example of an antibody design task according to some exemplary embodiments; and
[0045] Fig.11 Depicted is a block diagram illustrating an example of a computing system in accordance with some example embodiments.
[0046] When applicable, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION
[0047] Designing a molecule, including a biological sequence such as a protein or a non-biological small molecule, requires searching in a vast combinatorial design space. For example, de novo protein design aims to identify protein sequences (e.g., amino acid residue sequences) that exhibit a range of desirable properties, such as expression, binding affinity to another molecule (e.g., viral antigen, tumor antigen, etc.), lack of nonspecificity, stability, lack of immunogenicity, human nature, lack of self-association, etc. De novo protein design is a particularly challenging and resource-intensive task, at least because the combinatorial search space for every possible arrangement of amino acid residues that can form a protein structure is huge, but the number of amino acid residue sequences corresponding to actual functional proteins is scarce. That is, the vast majority of protein sequences in the combinatorial search space will not exhibit any function at all, let alone a combination of the aforementioned desired properties. When considering candidate protein sequences of variable length (e.g., candidate protein sequences formed by different numbers of amino acid residues), a search of this huge combinatorial search space becomes even more computationally difficult. Therefore, the brute force approach of indiscriminately examining every possible amino acid residue sequence to identify sequences that exhibit the desired properties is too computationally expensive to be a viable solution, even if performed in a computer.
[0048] In some exemplary embodiments, instead of exploring a vast combinatorial space sparsely populated by functional molecules, a molecular design engine can generate one or more molecular designs by sampling data distributions associated with various known molecules, including, for example, protein molecules, small molecules, ions, nucleic acids, polysaccharides, glycolipids, etc. For example, a molecular design engine can include a molecular design computational model that is trained using known molecules, including, for example, molecules known to exhibit certain functions and molecules without any known functions. In doing so, the molecular design computational model can learn data distributions corresponding to reduced-dimensional representations of the composition and / or structure of various known molecules. For example, in the case of de novo protein design, the data distribution may correspond to a reduced-dimensional representation of the amino acid residue sequences that form various known protein sequences. In some cases, the data distribution may occupy a topological space (e.g., a manifold) occupied by known molecules, which describes the relationships between known molecules. Although the high dimensionality of the data associated with known molecules tends to blur the relationships between groups of molecules with compositional and / or structural similarities, the data distribution learned by the molecular design computational model can occupy a lower-dimensional space, where one or more groups of molecules with compositional and / or structural similarities can form identifiable clusters.
[0049] Nevertheless, although computational molecular design models (including the above-mentioned molecular design computational models) can accelerate the initial molecular design process, limited wet laboratory resources still constitute a bottleneck to the rate at which candidate molecular designs are evaluated in vitro and in vivo. In a typical drug development pipeline, the molecular design must be verified in vitro and optimized for multiple rounds before the molecular design can proceed to preclinical development and clinical trials to test the performance of the molecule in vivo. Although computational molecular design models, such as the above-mentioned molecular design computational models, may be able to generate a large number of molecular designs (e.g., on the order of millions of molecular designs), limited wet laboratory resources preclude the synthesis and in vitro evaluation of each molecular design. Instead, a subset of molecular designs generated by the molecular design computational model can be selected for in vitro and / or in vivo evaluation.
[0050] Indiscriminate selection of molecular designs for in vitro and / or in vivo evaluation may increase the possibility of selecting molecular designs with poor molecular properties (e.g., suboptimal pharmacology and physiochemical properties, which increase the possibility of failure in subsequent preclinical development and clinical trials), while ignoring better candidates. In particular, in some cases, molecular designs can be generated and evaluated in continuous design iterations, wherein each design iteration (or design round) includes generating one or more new molecular designs to make improvements on the molecular designs from one or more previous design iterations. Therefore, the molecular designs selected for in vitro and / or in vivo evaluation during the current design iteration should exhibit better molecular properties than the molecular properties of the molecular designs of the previous design iterations. Therefore, as described in more detail below, the selection engine can perform multi-objective optimization on a set of partially ordered, mixed variable molecular properties when selecting the computationally generated molecular designs for further in vitro and / or in vivo evaluation. Doing so may increase the possibility of selecting better molecular designs for in vitro and / or in vivo evaluation, such as molecular designs that exhibit better molecular properties than the molecular designs of the previous design iterations.
[0051] In some exemplary embodiments, the selection engine can perform multi-objective Bayesian optimization (BO) that utilizes one or more probabilistic proxy models and a utility function to principledly weigh exploration (evaluating highly uncertain molecular designs) against exploitation (evaluating molecular designs that are believed to increase or maximize a goal). For example, the selection engine can apply one or more trained property computational models to determine a first probability that a first molecular design generated by a molecular design computational model exhibits a first property. In addition, the selection engine can apply one or more property computational models to determine a second probability that the first molecular design exhibits a second property. In this context, one or more property computational models can be used as in silico proxies for in vitro and / or in vivo evaluations that are too resource intensive to apply to every molecular design generated by a molecular design computational model.
[0052] In some exemplary embodiments, one or more characteristic calculation models may be implemented as one or more zero-inflated probability proxy models, each of which includes a probability binary classifier and a probability regressor model. Therefore, in some cases, the output of one or more characteristic calculation models may include a first plurality of prediction samples associated with a first molecular design or a prediction sample associated with a first molecular design. Each prediction sample associated with a first molecular design may include a first value of a first characteristic exhibited by the first molecular design and a second value of a second characteristic exhibited by the first molecular design having a first value of the first characteristic. In some cases, multiple characteristic calculation models may be applied to generate a first plurality of prediction samples in order to explain the uncertainty of each characteristic calculation model output. In this case, uncertainty may refer to the confidence level of the accuracy of the output of the characteristic calculation model, such as the accuracy of the characteristic value predicted by the characteristic calculation model for the molecular design. For example, in some cases, multiple characteristic calculation models may be applied to determine the first value of the first characteristic exhibited by the first molecular design, while multiple characteristic calculation models may also be applied to determine the second value of the second characteristic exhibited by the first molecular design. Therefore, the first prediction sample may include the output of a characteristic calculation model different from the characteristic calculation model used to generate the second prediction sample. In some cases, a first plurality of samples associated with a first molecular design may form a first distribution of a second value of a second property exhibited by the first molecular design over a first value of a first property exhibited by the first molecular design. In the case of an antibody design, the output of the one or more property computational models may include, for example, a distribution of binding affinity levels exhibited by the first molecular design over expression levels exhibited by the first molecular design. That is, in the case of an antibody design, the output of the one or more property computational models may include a prediction sample or a plurality of prediction samples, each of which includes an expression level of the first molecular design and a corresponding binding affinity of the first molecular design.
[0053] In some exemplary embodiments, the selection engine can determine a first utility metric, which indicates that the first molecular design is improved on the first magnitude of one or more baseline molecular designs with respect to the first characteristic and the second characteristic. For example, in the case of antibody design, the first utility metric of the first molecular design can indicate the magnitude of the improvement of the first molecular design on the expression level and binding affinity of the baseline molecular design. In addition, in some cases, the selection engine can select the first molecular design as a candidate for synthesis and testing based at least on the first utility metric of the first molecular design. In the process of doing so, the selection engine can ensure that the selection of the first molecular design as a candidate for synthesis and testing is a so-called joint positive molecular design, which refers to a molecular design that meets the specific standards of the first characteristic and the second characteristic. As described in more detail below, the selection engine can perform multi-objective optimization to select the aforementioned joint positive molecular design. It should be understood that the joint positive molecular design is not necessarily a molecular design with the best value in each molecular characteristic (e.g., the highest expression and the highest binding affinity), at least because such a molecular design may not exist at all. On the contrary, in this case, multi-objective optimization (MOO) can include a joint positive molecular design that identifies a first value showing a first characteristic, which cannot be improved without deteriorating the second value of the second characteristic.
[0054] In some exemplary embodiments, the selection engine may apply an active learning approach, wherein the first molecular design becomes one of the baseline molecular designs during subsequent design iterations. For example, the selection engine may determine to select a second molecular design generated by the molecular design computational model as the next candidate for synthesis based at least on a second utility metric indicating a second magnitude of improvement in the first characteristic and the second characteristic of the second molecular design relative to the baseline molecular design including the first molecular design. The relative values of the first characteristic and the second characteristic exhibited by the first molecular design may be determined based on the output of one or more characteristic computational models. Alternatively and / or additionally, the relative values of the first characteristic and the second characteristic exhibited by the first molecule may be determined based on one or more in vitro measurements or in vivo characterizations associated with the first molecular design. In some cases, in order to account for noise that may exist in one or more in vitro measurements or in vivo characterizations associated with the first molecular design (e.g., measurement errors associated with the laboratory device 130), the relative values of the first characteristic and the second characteristic exhibited by the first molecular design may be determined based on the output of one or more characteristic computational models after the one or more characteristic computational models have been updated, for example, by retraining based on one or more in vitro measurements or in vivo characterizations associated with the first molecular design.
[0055] In some exemplary embodiments, when determining a first utility metric associated with a first molecular design, the selection engine may apply a partial ranking to, for example, prioritize a first characteristic over a second characteristic. As described above, the first utility metric associated with the first molecular design may indicate a first magnitude of improvement in a first characteristic and a second characteristic of the first molecular design relative to a first characteristic and a second characteristic of one or more baseline molecules. In some cases, the partial ranking of the first characteristic and the second characteristic may require that the first characteristic of the first molecular design meets the first criterion before evaluating the first molecular design to determine whether the second characteristic of the first molecular design meets the second criterion. For example, in the context of antibody design, before evaluating the first molecular design for its binding affinity to a target antigen, the selection engine may require that the expression level of the first molecular design meet one or more criteria to reflect experimental and / or biological dependencies, wherein the first molecular design may be required to reach a certain expression level before a sufficient number of the first molecular design can be synthesized and assayed for other characteristics, such as binding affinity to a target antigen. Thus, to determine a first utility metric associated with a first molecular design while applying a partial ranking to prioritize a first characteristic over a second characteristic, the selection engine may include contributions from first samples in a first distribution where a first value of the first characteristic satisfies one or more criteria while excluding contributions from second samples in the first distribution where the first value of the first characteristic fails to satisfy the one or more criteria. In doing so, the first utility metric associated with the first molecular design may indicate a first magnitude of improvement in a first characteristic and a second characteristic of the first molecular design relative to a first characteristic and a second characteristic of one or more baseline molecular designs, where the first characteristic of the first molecular design satisfies the one or more criteria.
[0056] In some exemplary embodiments, when the first molecular design is selected as a candidate for synthesis and in vitro measurement and / or in vivo characterization, the selection engine may further evaluate the third characteristic of the first molecular design. For example, the selection engine may apply one or more characteristic calculation models to determine the third probability that the first molecular design exhibits the third characteristic. In this case, each of the first plurality of prediction samples output by the one or more characteristic calculation models may include a third value of the third characteristic exhibited by the first molecular design, which exhibits a first value of the first characteristic and a second value of the second characteristic. In addition, the first utility metric associated with the first molecular design may indicate how much the first characteristic, the second characteristic, and the third characteristic of the first molecular design are improved relative to the first characteristic, the second characteristic, and the third characteristic of one or more baseline molecules.
[0057] In some exemplary embodiments, the partial sorting that prioritizes the first characteristic over the second characteristic may further include a third characteristic. In some cases, the selection engine may apply a partial sorting so that the combination of the first characteristic and the third characteristic takes precedence over the second characteristic. In this particular scenario, before evaluating the first molecule design to determine whether the second characteristic of the first molecule design meets the second standard, the first characteristic and the third characteristic of the first molecule design may be required to meet the first standard. Alternatively and / or additionally, the selection engine may apply a partial sorting so that the first characteristic takes precedence over the second characteristic, and the second characteristic further takes precedence over the third characteristic. In this case, before evaluating the first molecule design for the second characteristic, the first characteristic of the first molecule design may be required to meet the first standard, and before further evaluating the first molecule design to determine whether the third characteristic of the first molecule design meets the third standard, the second characteristic of the first molecule design may be required to meet the second standard. Referring again to the antibody design example, the selection engine may require the expression level of the first molecule design to meet the first standard before evaluating the first molecule design for its binding affinity to the target antigen. In addition, before the selection engine evaluates various developability features (e.g., specificity, thermal stability, etc.) of the first molecule design, the binding affinity of the first molecule design may be further required to meet the second standard. By including the third property, the selection engine can ensure that the first molecular design selected as a candidate for synthesis and testing is a joint positive molecular design that meets the specified criteria regarding the first property, the second property, and the third property.
[0058] As described above, in some exemplary embodiments, the selection engine can perform multi-objective optimization, such as multi-objective Bayesian optimization, on the characteristics (or objectives) of multiple partially ordered, mixed variables. This framework may reflect some scenarios in drug design, where the molecular design may need to satisfy a first characteristic (e.g., expression) before evaluating the molecular design for a second characteristic (e.g., affinity) and / or a third characteristic (e.g., specificity). As described in more detail below, multi-objective optimization (e.g., multi-objective Bayesian optimization) may include applying a partial sort during drug design that prioritizes satisfying the first characteristic (e.g., expression) over satisfying the second characteristic (e.g., affinity) and / or satisfying the third characteristic (e.g., specificity). For example, for each molecular design, the selection engine can modify the posterior probability distribution of each objective (e.g., determined by one or more probabilistic proxy models) so that the characteristics exhibited by the molecular design are modeled as a zero-inflated distribution (a mixture of continuous distributions of zero and non-zero values), and certain characteristics are prioritized over other characteristics. In doing so, the selection engine can identify significantly more joint positive molecular designs, such as in molecular designs that meet the criteria on each property being improved, compared to conventional techniques such as standard Bayesian optimization. Therefore, the selection engine increases or maximizes the possibility of selecting better molecular designs for in vitro and / or in vivo evaluation. In particular, selecting molecular designs from the current design iteration can utilize experimental knowledge from previous design iterations, so that candidates with progressively better properties are selected in successive design iterations.
[0059] Figure 1 A system diagram is depicted that illustrates an example of a protein design system 100 according to some exemplary embodiments. Figure 1 , the molecular design system 110 may include a molecular design engine 110, a selection engine 120, one or more wet lab devices 130, and a client device 140. Figure 1 As shown, the molecular design engine 110, the selection engine 120, one or more laboratory devices 130 and the client device 140 can be communicatively coupled via a network 150. The one or more laboratory devices 130 may include any wet laboratory and dry laboratory equipment capable of performing in vitro measurements and / or in vivo characterizations. Examples of the one or more laboratory devices 130 may include sequencers, mass spectrometers, centrifuges, and the like. The client device 140 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smart phone, a tablet computer, a wearable device, and the like. The network 150 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and the like.
[0060] Reference again Figure 1, the molecular design engine 110 can apply the molecular design computational model 115 to generate a plurality of molecular designs, including, for example, a first molecular design 160a, a second molecular design 160b, and the like. For example, in some cases, the molecular design computational model 115 can be a trained machine learning model to learn data distributions corresponding to reduced-dimensional representations of the composition and / or structure of various known molecules (e.g., protein sequences). In some cases, the data distribution can be a topological space (e.g., a manifold) occupied by known molecules, which describes the relationships existing between known molecules. The molecular design computational model 115 can generate each of the first molecular design 160a and the second molecular design 160b by sampling the data distribution (e.g., topological space). For example, in some cases, the first molecular design 160a and the second molecular design 160b can each be a protein sequence that occupies a data distribution (e.g., a topological space) corresponding to its reduced-dimensional representation.
[0061] In some exemplary embodiments, the molecular design engine 110 may be able to generate a large number of molecular designs, but not every molecular design generated by the molecular design engine 110 can be evaluated in vitro and in vivo. Instead, the selection engine 120 may perform one or more active learning iterations to identify one or more joint positive molecular designs that meet specific criteria about multiple properties for synthesis and testing by one or more laboratory devices 130. In some cases, one or more criteria associated with a property may include that a property value meets one or more thresholds, falls within one or more intervals of values, or is a member of a set. For example, in the case of antibody design, the selection engine 120 may perform one or more active learning iterations to identify one or more molecular designs that exhibit sufficient expression levels and sufficient binding affinity for a target antigen. In some cases, the selection engine 120 may further perform one or more iterations of active learning to identify one or more molecular designs that exhibit certain developability features, such as specificity, thermal stability, etc., in addition to having sufficient expression levels and sufficient binding affinity.
[0062] Figure 2A A flow chart is depicted that illustrates an example of a process 200 for molecular design with multi-objective optimization of partially ordered, mixed variable properties, according to some exemplary embodiments. Figure 1 -2, process 200 can be performed by the molecular design engine 110 and the selection engine 120 to identify, for example, a subset of molecular designs generated by the molecular design engine 110 as candidates for in vitro and / or in vivo evaluation.
[0063] At 202, the molecular design engine 110 may generate a plurality of molecular designs. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 115 to generate a plurality of molecular designs, including, for example, a first molecular design 160a, a second molecular design 160b, and the like.
[0064] At 204, for each of the plurality of molecular designs, the selection engine 120 may determine a utility metric corresponding to a magnitude of improvement in a combination of properties exhibited by each molecular design relative to a combination of properties exhibited by one or more baseline molecular designs. In some exemplary embodiments, for each of the plurality of molecular designs generated by the molecular design engine 110, the selection engine 120 may determine a corresponding utility metric indicating how much the combination of properties exhibited by each molecular design (e.g., expression, binding affinity, specificity, and thermal stability, etc.) is improved relative to a combination of the same properties exhibited by one or more baseline molecules. For example, in some cases, for the first molecular design 160a, the selection engine 120 may determine a first utility metric indicating a first magnitude of improvement in a combination of properties exhibited by the first molecular design 160 relative to a combination of properties exhibited by one or more baseline molecular designs. Additionally, for second molecular design 160b, selection engine 120 can determine a second utility metric indicating a second magnitude of improvement in a property of second molecular design 160b relative to a property of a baseline molecular design, which in some cases can include first molecular design 160a.
[0065] At 206, the selection engine 120 may select one or more molecular designs as candidates for synthesis and testing based at least on the utility metrics associated with each of the plurality of molecular designs. For example, in some cases, the selection engine 120 may identify the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing if the corresponding utility metrics meet one or more thresholds. Alternatively and / or additionally, the selection engine 120 may select N number of molecular designs with the highest utility metrics from the plurality of molecular designs generated by the molecular design engine 110 as candidates for synthesis and testing. In this case, if the first molecular design 160a and / or the second molecular design 160b are part of the N number of molecular designs with the highest utility metrics among the plurality of molecular designs generated by the molecular design engine 110, the first molecular design 160a and / or the second molecular design 160b may be selected as candidates for synthesis and testing. In some cases, in addition to the utility metric associated with each of the first molecular design 160a and the second molecular design 160b, the selection engine 120 may impose additional conditions when selecting the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing. For example, in the case of antibody design, when selecting the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing, the selection engine 120 may further require the presence (or absence) of certain amino acid residues (or amino acid residue sequences).
[0066] Figure 2B A flow chart is depicted that illustrates another example of a process 250 for molecular design with multi-objective optimization of partially ordered, mixed variable properties, according to some exemplary embodiments. Figure 1 and Figure 2A-2B , process 250 may be performed, for example, by selection engine 120 to determine a utility metric for each molecular design generated by molecular design engine 110. In some cases, process 250 may implement Figure 2A At least a portion of operation 204 of process 200 is shown.
[0067] At 252 , the selection engine 120 can receive a molecular design. For example, in some cases, the selection engine 120 can receive a first molecular design 160 a generated by the molecular design computational model 115 from the molecular design engine 110 .
[0068] At 254, the selection engine 120 may apply one or more property computation models to determine a first probability that the molecular design exhibits a first property and a second probability that the molecular design exhibits a second property. For example, in some cases, the first property computation model may be applied to determine the first probability that the first molecular design 160a exhibits the first property by at least listing the probability of occurrence of each possible value of the first property of the first molecular design 160a. In some cases, the first property computation model or the second property computation model may be applied to determine the second probability that the first molecular design 160a exhibits the second property by at least listing the probability of occurrence of each possible value of the second property of the first molecular design 160a. In some cases, multiple property computation models (e.g., an integration of property computation models) may be applied to determine the value of each property of the first molecular design 160a.
[0069] At 256, based at least on the output of the one or more property computation models, the selection engine 120 may determine a plurality of prediction samples associated with the molecular design, wherein each prediction sample includes a first value of a first property exhibited by the molecular design and a second value of a second property exhibited by the molecular design having the first value for the first property. For example, in some cases, the output of the one or more property computation models may include a plurality of intermediate a posteriori samples, each of which includes a first value of the first property exhibited by the first molecular design 160a and a second value of the second property exhibited by the first molecular design 160a. The intermediate a posteriori samples associated with the first molecular design 160a may correspond to a distribution of the second value of the second property exhibited by the first molecular design 160a over the first value of the first property exhibited by the first molecular design 160a.
[0070] At 258, the selection engine 120 may identify a set of prediction samples within the plurality of prediction samples where the first value of the first characteristic satisfies the criterion. In some exemplary embodiments, the selection engine 120 may determine a utility metric that indicates a magnitude of improvement in the first characteristic and the second characteristic of the first molecular design 160a relative to the first characteristic and the second characteristic of one or more baseline molecular designs. Furthermore, when determining the utility metric associated with the first molecular design 160a, the selection engine 120 may apply a partial ranking, for example, to prioritize the first characteristic over the second characteristic. To further illustrate, Figure 4 A schematic diagram is depicted showing an example of a hierarchical structure 400 in which a first characteristic y of a first molecular design 160a is 0.0 Prioritizes the second characteristic y of the first molecular design 160a 1.0 And the third characteristic y 1.1 . Figure 4 The example of the hierarchy 400 shown in FIG. 4 includes three properties y distributed over two levels indexed by l. l,mThe features occupying the same level of the hierarchy 400 are further indexed by m. 0.0 To the second characteristic y 1.0 and the third characteristic y 1.1 The arrows for each of the may symbolize dependencies between them, such as experimental and / or biological dependencies. Alternatively and / or additionally, from the first characteristic y 0.0 To the second characteristic y 1.0 and the third characteristic y 1.1 The arrows in each of the 0.0 Priority over the second characteristic y 1.0 and the third characteristic y 1.1 As explained in more detail below, from the output r of the probabilistic regressor model 315, l,m and the predicted sample b output by the probabilistic binary classifier 313 {l,m} To feature y l,m The arrows indicate each y l,m are all modeled as zero-inflated, where b l,m Controlling zero events and r l,m Controls consecutive non-zero events.
[0071] At 260, the selection engine 120 may determine a utility metric based at least on the set of predicted samples, the utility metric corresponding to the expected improvement of the first characteristic and the second characteristic of the molecular design relative to the first characteristic and the second characteristic of one or more baseline molecular designs. For example, in some cases, the selection engine 120 may determine a utility metric indicating the magnitude of the improvement of the first characteristic and the second characteristic of the first molecular design 160a relative to the first characteristic and the second characteristic of the one or more baseline molecular designs. In addition, the selection engine 120 may apply a partial sorting by resampling at least the intermediate a posteriori samples associated with the first molecular design 160a to generate a plurality of corresponding a posteriori samples, the partial sorting giving priority to the first characteristic over the second characteristic. These a posteriori samples are then subjected to a multi-objective acquisition function to determine the utility metric of the first molecular design 160a. Examples of multi-objective acquisition functions applied to determine the utility metric of the first molecular design 160a may include expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc.
[0072] Figure 3 A system diagram is depicted that illustrates an example of a protein design system 300 according to some exemplary embodiments. Figure 3As shown in the depicted molecular design pipeline 300, the selection engine 120 can apply one or more property calculation models 310 to determine, for example, a first probability that the first molecular design 160a exhibits a first property and a second probability that the first molecular design 160b exhibits a second property. In some cases, the one or more property calculation models 310 can output a first probability distribution indicating the first probability that the first molecular design 160a exhibits the first property by at least listing the probability of occurrence of each possible value of the first property exhibited by the first molecular design 160a. In addition, in some cases, the one or more property calculation models 310 can output a second probability distribution indicating the second probability that the second molecular design 160a exhibits the second property by at least listing the probability of occurrence of each possible value of the second property exhibited by the first molecular design 160a.
[0073] In some exemplary embodiments, the one or more characteristic calculation models 310 may be implemented as a zero-inflated probability proxy model including a probabilistic binary classifier 313 and a probabilistic regressor model 315. In addition, in some cases, the one or more characteristic calculation models 310 may include an independent trained characteristic calculation model to determine the probability that the first molecular design 160a exhibits each individual characteristic. For example, the one or more characteristic calculation models 310 may include a first characteristic calculation model trained to determine the first probability that the first molecular design 160a exhibits the first characteristic and a second characteristic calculation model trained to determine the second probability that the first molecular design 160a exhibits the second characteristic. In some cases, the one or more characteristic calculation models 310 may include at least one trained characteristic calculation model to determine the probability that the first molecular design 160a exhibits multiple characteristics including, for example, the first characteristic, the second characteristic, etc. In addition, in some cases, the one or more characteristic calculation models 310 may include multiple trained characteristic calculation models (e.g., an integration of characteristic calculation models) to determine the probability that the first molecular design 160a exhibits the same characteristic. For example, the one or more property calculation models 310 may include a first property calculation model and a third property calculation model, each of the one or more property calculation models being trained to determine a first probability that the first molecular design 160 exhibits a first property.
[0074] Including multiple property calculation models (or property calculation model integration) for a single property can compensate for at least some uncertainties that may exist in the output of a separate property calculation model. For example, in some cases, for some molecular designs encountered by the property calculation model, the output of a separate property calculation model may be less uncertain (or have a higher accuracy confidence), and for other molecular designs encountered by the property calculation model, the output of a separate property calculation model may be less certain (or have a lower accuracy confidence). Alternatively and / or additionally, for a specific molecular design, the output of a separate property calculation model may not be as uncertain (or have a higher accuracy confidence) as the output of another separate property calculation model for the same molecular design. Therefore, when multiple property calculation models (or an integration of property calculation models) are applied to determine the properties of a molecular design, the lower uncertainty in the output of some property calculation models can compensate for the higher uncertainty in the output of other property calculation models.
[0075] Reference again Figure 3 , the output of one or more characteristic calculation models 310 may include multiple intermediate a posteriori samples 320. Each sample in the multiple intermediate a posteriori samples 320 may include a first value of the first characteristic exhibited by the first molecular design 160a and a second value of the second characteristic exhibited by the first molecular design 160b. The multiple intermediate a posteriori samples 320 may correspond to the distribution of the second value of the second characteristic exhibited by the first molecular design 160a on the first value of the first characteristic exhibited by the first molecular design 160a. For example, if the first characteristic is the expression level and the second characteristic is the binding affinity, each intermediate a posteriori sample may include the first value of the expression level exhibited by the molecular design and the second value of the binding affinity exhibited by the molecular design with the first value for the expression level. Therefore, the multiple intermediate a posteriori samples 320 may list the distribution of different expression levels on different binding affinities exhibited by the molecular design.
[0076] In some exemplary embodiments, the selection engine 120 may apply a partial ordering that prioritizes the first characteristic over the second characteristic by subjecting at least the intermediate a posteriori samples 320 to resampling 330 to generate a plurality of a posteriori samples 340. A multi-objective acquisition function 350 may be applied to the a posteriori samples 340 to determine a utility metric for the first molecular design 160a. Examples of the multi-objective acquisition function 350 may include expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), and the like. The resampling 330 may convert any intermediate a posteriori samples in which the first value of the first characteristic fails to meet one or more criteria to prevent these intermediate a posteriori samples from contributing to the utility metric for the first molecular design 160a.
[0077] To further illustrate, Figure 5 The effect of resampling 330 on intermediate a posteriori samples 320 output by one or more characteristic computation models 310 is depicted. Figure 5 The dashed line in represents a criterion (e.g., threshold) for goal 0 such that molecular designs selected as candidates for synthesis and testing should increase or maximize goal 1 while satisfying this criterion in goal 0. Figure 5 The points in constitute separate baseline molecular designs, which form the baseline Pareto front. The different shading in the grid represents the hypervolume improvement (HVI) computed from each posterior sample at a given location in the target space. Consider six intermediate posterior samples 320 (triangles) from the posterior (white boundary line), which are Figure 5 Their default state is shown in Figure 500 in (a). Figure 5 As shown in the diagram 550 in (b), the resampling 330 can transform the intermediate posterior samples 320 that do not meet the criteria in the target 0 (e.g., below the threshold) so that their hypervolume improvement (HVI) contribution is zero. That is, due to Figure 5 The resampling 330 shown in Figure 550 in (b) is for Figure 5 For those intermediate posterior samples 320 that fail to meet the target (0) (eg, dashed vertical lines) that are not zero in the default state shown in the graph 500 in (a), the value of the hyper volume improvement (HVI) is set to zero.
[0078] In some exemplary embodiments, the utility metric output by the multi-objective acquisition function 350 for the first molecular design 160a may indicate how much the first and second characteristics of the first molecular design 160a are improved over the first and second characteristics of one or more baseline molecular designs. In addition, as part of the active learning paradigm, the first molecular design 160a may become one of the baseline molecular designs for subsequent selection iterations. For example, the selection engine 120 may determine to select the second molecular design 160b generated by the molecular design engine 110 as another candidate for synthesis based at least on a second utility metric corresponding to an expected improvement in the first and second characteristics of the second molecular design 160b relative to the first and second characteristics of the baseline molecular design, in which selection iteration the baseline molecular design may be updated to include the first molecular design 160a. The values of the first and second characteristics associated with the first molecular design 160a may be determined empirically, for example based on in vitro measurements and / or in vivo characterizations. Alternatively and / or additionally, the values of the first and second properties of the first molecular design 160a may correspond to the intermediate a posteriori samples 320 associated with the first molecular design 160a (eg, an average of the values of the first and second properties included in the intermediate a posteriori samples 320).
[0079] As described above, the selection engine 120 can determine a utility metric that indicates the magnitude of improvement in the first and second properties of the first molecular design 160a relative to the first and second properties of one or more baseline molecular designs. Figure 3 As shown in the molecular design pipeline 300 in FIG. 1 , the selection engine 120 can apply a partial ordering that prioritizes a first characteristic over a second characteristic by resampling 330 at least the intermediate a posteriori samples 320 to generate a plurality of a posteriori samples 340, and then the a posteriori samples 340 are subjected to a multi-objective acquisition function 350 to determine a utility metric for the first molecular design 160a. Examples of the multi-objective acquisition function 350 can include expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), and the like.
[0080] In some exemplary embodiments, the selection engine 120 can compensate for noise (e.g., measurement errors associated with the lab equipment 130) that may be present in the observed characteristics of one or more baseline molecular designs. Figure 1 In the example shown in , the selection engine 120 determines a utility metric that indicates the magnitude of improvement in the first and second properties of the first molecular design 160a relative to the first and second properties of one or more baseline molecular designs, the values of the first and second properties of the one or more baseline molecular designs may be observed in a wet lab and may therefore include at least some noise caused by measurement errors associated with the laboratory equipment 130. The effects of this noise may be reduced or minimized by determining the utility metric based on the output of one or more property calculation models 310, which are retrained based on observed values of the first and second properties of the one or more baseline molecular designs. For example, the retrained property calculation model 310 may be applied to determine the values of the first and second properties for the first molecular design 160a and the values of the first and second properties for the one or more baseline molecular designs. The utility metric for the first molecular design 160a may be determined based on the output of the retrained property calculation model 310 rather than the observed values of the first and second properties of the one or more baseline molecular designs.
[0081] To further illustrate, Fig. 6A A flow chart is depicted that illustrates an example of a process 600 for determining the probability that a molecular design exhibits a property, according to some exemplary embodiments. Figure 1 , Figure 2A and Fig. 6A, process 600 may be performed by selection engine 120 and may implement, for example Figure 2A At least a portion of operation 204 of process 204 shown in .
[0082] At 602, for a baseline molecular design from a previous design iteration, the selection engine 120 may receive observations of properties of the baseline molecular design. In some exemplary embodiments, a molecular design from a previous design iteration may become a baseline molecular design for a subsequent design iteration. For example, one or more molecular designs from a previous design iteration may be selected for in vitro measurement and / or in vivo characterization of one or more desirable properties (e.g., expression, binding affinity to another molecule (e.g., viral antigen, tumor antigen, etc.), lack of non-specificity, stability, lack of immunogenicity, humanity, lack of self-association, etc.). As described in more detail below, the observations of properties exhibited by these baseline molecular designs may be used in subsequent design iterations to identify one or more molecular designs whose combination of properties exhibits the greatest improvement over the properties of the baseline molecular design.
[0083] At 604, based at least on the observed values of the characteristics of the baseline molecular design, the selection engine 120 can retrain one or more characteristic calculation models to determine the probability distribution over different possible values of the characteristics. The observed values of the characteristics of the baseline molecular design received in operation 602 may include at least some noise due to measurement errors present in, for example, the laboratory equipment 130 used to generate the observed values. Therefore, in some cases, the observed values of the characteristics of the baseline molecular design are not directly used to identify molecular designs from subsequent design iterations whose combinations of characteristics show improvement or maximum improvement relative to the characteristics of the baseline molecular design. Instead, in some exemplary embodiments, the observed values of the characteristics of the baseline molecular design are applied to retrain one or more characteristic calculation models 310. Doing so can, for example, generate a posterior probability distribution for the corresponding characteristics by updating the prior probabilities of these characteristics based on the observed values.
[0084] At 606, the selection engine 120 may apply one or more retrained characteristic calculation models to determine a first value for a characteristic of a baseline molecular design and a second value for a characteristic of a molecular design from a subsequent design iteration. For example, in some cases, the retrained characteristic calculation model 310 may be applied to determine characteristic values of a baseline molecular design from a previous design iteration and characteristic values of a molecular design generated during a subsequent design iteration. For example, in the case of expression levels, the selection engine 120 may apply a retrained characteristic calculation model 310 to determine a first expression level of a baseline molecular design from a previous design iteration and a second expression level of a molecular design from a subsequent design iteration. Although the expression level of the baseline molecular design has been observed through wet lab experiments, the utility metric of the molecular design from a subsequent design iteration is not directly determined based on the observed expression level of the baseline molecular design. On the contrary, as described in more detail below, the utility metric of the molecular design from a subsequent design iteration may be determined based on the first expression level of the baseline molecular design determined by the retrained characteristic calculation model 310.
[0085] At 608, the selection engine 120 may determine a utility metric based at least on the first value of the characteristic for the baseline molecular design and the second value of the characteristic for the molecular design from the subsequent design iteration, the utility metric corresponding to the magnitude of improvement of the combination of characteristics exhibited by the molecular design from the subsequent design iteration over the combination of characteristics exhibited by the baseline molecular design. In some exemplary embodiments, for the molecular design from the subsequent design iteration, the selection engine 120 may determine a utility metric corresponding to how much the characteristic of the molecular design is improved over the characteristic of the baseline molecular design from one or more previous design iterations. As described above, even if the observed values for the characteristic of the baseline molecular design are available, instead, the utility metric for the molecular design from the subsequent design iteration may be determined based on the output of the characteristic calculation model 310, which is retrained based on the observed values. Returning to the expression level example, the utility metric for the molecular design from the subsequent design iteration may be determined based on the first expression level of the baseline molecular design and the second expression level of the molecular design from the subsequent design iteration, wherein both the first expression level and the second expression level are determined by the retrained characteristic calculation model 310. In some cases, the utility metric can quantify the expected improvement (EI) for the combination of partially ordered, mixed-valued properties exhibited by the molecular design from the subsequent design iteration relative to the baseline molecular design. In addition, in some cases, the utility metric of the molecular design from the subsequent design iteration can be calculated by applying a utility function (e.g., a multi-objective acquisition function 350), the utility function including, for example, expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc.
[0086] In some exemplary embodiments, the selection engine 120 can compensate for the uncertainty that may exist in the output of each characteristic calculation model 310. For example, the selection engine 120 can apply one or more characteristic calculation models 310 to determine the first value of the first characteristic and the second value of the second characteristic exhibited by the molecular design. In some cases, the uncertainty in the output of each characteristic calculation model 310 can be reduced or minimized by applying at least a plurality of characteristic calculation models 310 (e.g., an integration of characteristic calculation models) to determine each of the first value of the first characteristic and the second value of the second characteristic for the molecular design. For example, the selection engine 120 can apply the first characteristic calculation model and the second characteristic calculation model to evaluate the first characteristic of the molecular design, and can determine the first value of the first characteristic based on the output of the first characteristic calculation model and the second characteristic calculation model. Similarly, the selection engine 120 can apply the third characteristic calculation model and the fourth characteristic calculation model to evaluate the second characteristic of the molecular design, and can determine the second value of the second characteristic based on the output of the third characteristic calculation model and the fourth characteristic calculation model.
[0087] To further illustrate, Figure 6B A flow chart is depicted that illustrates another example of a process 650 for determining the probability that a molecular design exhibits a property, according to some exemplary embodiments. Figure 1 , Figure 2A and Fig. 6A -B, process 650 may be performed by selection engine 120 and may implement, for example Figure 2B At least a portion of operation 254 of process 250 shown in Fig. 6A Operation 606 of process 600 is shown in FIG.
[0088] At 652, the selection engine 120 may apply a first characteristic computational model to determine a first value of a characteristic exhibited by the molecular design, and apply a second characteristic computational model to determine a second value of a characteristic exhibited by the molecular design. For example, in the case of expression level, the selection engine 120 may apply a set of characteristic computational model integration (e.g., a first characteristic computational model, a second characteristic computational model, etc.), which is trained to determine the expression level, so as to determine the expression level of each molecular design. Alternatively, in the case of binding affinity, the selection engine 120 may also apply a set of characteristic computational model integration (e.g., a third characteristic computational model, a fourth characteristic computational model, etc.), which is trained to determine the binding affinity, so as to determine the binding affinity of each molecular design. In some cases, the first value of the characteristic exhibited by the molecular design and the second value of the characteristic may be represented as a probability distribution. For example, the output of the first characteristic computational model may include a first probability distribution over a range of possible values for the characteristic (e.g., expression level, binding affinity, etc.), and the output of the second characteristic computational model may include a second probability distribution over a range of possible values for the characteristic (e.g., expression level, binding affinity, etc.). In some cases, the outputs of the first property computational model and the second property computational model can be zero-inflated, meaning that the output includes a first value (e.g., a binary value) indicating the presence (or absence) of a property (e.g., expression level, binding affinity, etc.) and a second value indicating the magnitude of the property exhibited by the molecular design.
[0089] At 654, the selection engine 120 may determine a third value of a characteristic exhibited by the molecular design for calculating a utility measure of the molecular design based at least on the first value and the second value. The output of each probability calculation model in the above-mentioned integration may be associated with at least some uncertainty. In particular, when applied to different molecular designs, differences in architecture and / or training may result in more (or less) certainty in the characteristic calculation model. For example, the first characteristic calculation model may generate a more certain output than the second characteristic calculation model for the first molecular design, but the second characteristic calculation model may generate a more certain output than the first characteristic calculation model for the second molecular design. Therefore, in some exemplary embodiments, the selection engine 120 may compensate for this uncertainty by at least determining the utility measure of the molecular design based on the output of multiple characteristic calculation models rather than the output of a single characteristic calculation model. In some cases, characteristic values (e.g., expression level, binding affinity, etc.) for determining the utility measure for the molecular design may be determined based on multiple values of the same characteristic determined by the integration of characteristic calculation models. For example, where a first property calculation model is applied to determine a first value of a property and a second property calculation model is applied to determine a second value of a property, the selection engine 120 may determine a third value of the property for determining a utility metric for a corresponding molecular design based at least on the first value and the second value. In some cases, the third value of the property may correspond to a mean, median, maximum, minimum, mode, and / or range of the first and second values.
[0090] As previously described, in some exemplary embodiments, the selection engine 120 can perform multi-objective Bayesian optimization (BO) by utilizing one or more probabilistic proxy models (e.g., one or more property computation models 310) and a utility function (e.g., multi-objective acquisition function 350) to weigh exploration (evaluating highly uncertain molecular designs) against exploitation (evaluating molecular designs that are believed to increase or maximize the objective). If the objective (or characteristic) is a black box function of the design space If evaluation is expensive (e.g., in a wet lab), then the goal of Bayesian optimization is to efficiently identify designs that increase or maximize f Thus, in this case, Bayesian optimization (BO) may include utilizing one or more probabilistic proxy models (e.g., one or more property computation models 310) and a utility function (e.g., multi-objective acquisition function 350) to weigh the tradeoffs of exploring the design space. To evaluate more uncertain molecular designs (e.g., molecular designs with unknown probability of increasing or maximizing f) versus developing more specific molecular designs that are believed to increase or maximize f.
[0091] Probabilistic proxy model (eg, characteristic calculation model 310) Beliefs about the probability distribution f can be constructed based on existing information. For example, where f is the expression level of a molecular design, a probabilistic proxy model (e.g., property computation model 310) can be trained based on network laboratory measurements of expression levels exhibited by molecular designs from previous design iterations. Given the presence of observation noise (e.g., measurement errors associated with the laboratory equipment 130), the characteristic calculation model 310 can be trained on a noisy data set available prior to a given design iteration t. In other words, each iteration t∈N can be trained with the data set associated, where each y (n) are all noisy observations of f. Therefore, a probabilistic proxy model (e.g., feature computation model 310) can be trained To infer the posterior distribution Posterior distribution quantification proxy target In the expression level example, the posterior distribution The probability distribution of possible expression levels exhibited by the next batch of molecular designs (eg, generated by the molecular design computational model 115) is quantified.
[0092] Utility function (eg, multi-objective acquisition function 350) α: The posterior distribution determined by the probabilistic proxy model (eg, the feature calculation model 310) may be ingested And determine the corresponding utility metric α for each molecular design. In some cases, the utility metric α can quantify the usefulness of each molecular design x. As described in more detail below, the usefulness (or utility) of each molecular design x may correspond to the possibility that the molecular design exhibits better characteristics than the molecular design from the previous design iteration. For example, in some cases, a molecular design that increases or maximizes the utility metric α can be selected for further in vitro measurements and / or in vivo characterization. That is, in some cases, the molecular design selected for in vitro measurement and / or in vivo characterization can be identified based on the difference between the expected value exhibited by each molecular design for a set of partially ordered, mixed variable characteristics (or targets) and the maximum value observed so far for the same characteristics (or targets) (for example, in the molecular design of the previous design iteration). The expected value for each characteristic (or target) can be calculated based on the aforementioned posterior distribution inferred by the corresponding characteristic calculation model 310.
[0093] In some exemplary embodiments, a utility function (e.g., multi-objective acquisition function 350) may determine an expected improvement (EI) in a property exhibited by a molecular design in a current design iteration relative to a property exhibited by a molecular design from a previous design iteration. Examples of utility functions (e.g., multi-objective acquisition function 350) a(x) may include expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), and the like. In the case of expected improvement (EI), the expected improvement (EI) acquisition function may be implemented by taking Obtain, where [·] + = max(·,0). In some cases, the integral can be approximated using a posteriori samples via Monte Carlo (MC) integration The maximized value of α can be selected as the molecular design for in vitro measurement and / or in vivo characterization. As described in more detail below, the actual value of the property exhibited by the molecular design can be measured (e.g., in a network laboratory) before the observations are appended to a data set for retraining a corresponding probabilistic proxy model (e.g., corresponding property computation model 310).
[0094] As described above, the utility metric α for the molecular design in the current design iteration can be calculated relative to the characteristic values of the molecular design from the previous design iteration. In some cases, the characteristic values of the molecular design from the previous design iteration can be determined based on the wet laboratory measurement results. Alternatively, the characteristic values of the molecular design from the previous design iteration can be determined by the corresponding probabilistic proxy model (e.g., the corresponding characteristic calculation model 310) after retraining the model based on the wet laboratory measurement results of these characteristics. For example, in some cases, after querying the probabilistic proxy model f for the molecular design (e.g., the characteristic calculation model 310) and obtaining the label pair (x) for the design, for example, from the wet laboratory measurement results. l ,y l ), you can append these values to the dataset The probabilistic proxy model (e.g., property calculation model 310) can be used to determine the posterior distribution of the corresponding property value for the molecular design from the next design iteration t+1 before being applied to the enhanced data set. Retrain on.
[0095] When there is a single target (property) of interest, the best molecular design can be identified based on a ranking of property values (e.g., highest expression level, highest binding affinity, etc.). When there are multiple targets (or properties) of interest, the best molecular design may not be the one with the best value for each target (or property), at least because a single molecular design that performs well in each target f may not exist. Therefore, in a scenario with K targets (or properties), there may be at least K number of probabilistic surrogate models (e.g., property calculation 310) for k=1,...,K In some cases, the goal of multi-objective optimization (MOO) may not be a single molecular design with optimal values for each property, but rather to identify a set of Pareto optimal trade-offs such that improving one objective (or property) within the set results in a deterioration of another objective (or property). For example, a Pareto optimal trade-off for optimization over expression level and binding affinity may be a set of molecular designs where an improvement in expression level is accompanied by a decrease in binding affinity.
[0096] To illustrate further, consider two molecular designs x (1) and x (2) , and represent the solution of each design x in the target space (of K targets) as f(x): = [f 1 (x),…,f k (x)]. In some cases, f k (x (1) ) can be said to dominate f k (x (2) ), or f k (x (1) )>f k (x (2) ), if for all K targets, f k (x (1) )≥f k (x (2) ), and for at least one of the K targets, f k (x (1) )>f k (x (2) ). …The Pareto front (PF) can be defined as the set of non-dominated solutions expressed by the following equation (1).
[0097]
[0098] The above expression of the Pareto front (PF) may produce a set of Pareto optimal designs This is why multi-objective optimization (MOO) determines the Pareto optimal design set within a reasonable number of iterations. One way to measure the quality of an approximate Pareto front (PF) is to compute the number of points dominated by P and given by a given reference point. The hypervolume HV(P|r ref ), the specified reference point is, for example, the value of the K target exhibited by the molecular design from the previous design iteration. The reference point is represented as The Pareto front P at design iteration t relative to the existing baseline can be t The expected improvement in expected hypervolume improvement (EHVI) is expressed as As described above, the expected hypervolume improvement (EHVI) acquisition function is an example of a multi-objective acquisition function 350 that can be applied by the selection engine 120. Other examples of the multi-objective acquisition function 350 may include the noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc., which are described in more detail below. According to the following formula (2), the expected hypervolume improvement (EHVI) acquisition function can be used in the posterior distribution output by the probabilistic proxy model (e.g., the characteristic calculation model 310) The expected hypervolume on improves the expectation of HVI. The integral in Eq. (2) can be obtained by Draw samples (e.g., Monte Carlo samples) from the dataset to evaluate.
[0099]
[0100] In a noise-free setting, the observed baseline Pareto front (PF) may be the true baseline Pareto front (e.g., in However, this does not hold true in many practical applications where wet lab measurements carry noise, such as in the form of measurement errors associated with the lab equipment 130. For example, given a zero-mean Gaussian measurement process with noise covariance ∑, the feedback for a molecular design x is y~N(f(x),∑) rather than f(x) itself. To account for this noise, the baseline Pareto front (PF) associated with the molecular designs from a previous design iteration can correspond to the property values determined by an updated probabilistic proxy model (e.g., property computation model 310) that has been retrained on the wet lab measurements for these molecular designs. That is, according to equation (3) below, the noise expected excess volume improvement (NEHVI) may be at a previously observed point It should be appreciated that while Noise Expected Improvement (EI) extends expected improvement to settings with observation noise (e.g., measurement errors associated with laboratory equipment 130), Noise Expected Hypervolume Improvement (NEHVI) extends noise expected improvement to optimization of multiple objectives (or properties).
[0101]
[0102] Due to the delay in feedback, sequential optimization or querying f for a single molecular design per iteration may be impractical for many applications. For example, in protein engineering, one may need to select a batch of molecular designs in a given iteration and wait several months to receive measurement results. Jointly selecting a batch of q molecular designs from a large number of q′>>q candidates may require combinatorial evaluation of utility functions (e.g., multi-objective acquisition function 350). In the context of optimizing molecular designs on gradients based on utility functions (e.g., multi-objective acquisition function 350), sequential greedy selection of q molecular designs in a given design iteration can achieve performance comparable to joint selection of q candidates for various utility functions (e.g., multi-objective acquisition function 350).
[0103] Many molecular design applications require some kind of hierarchy among multiple target properties of interest. In some cases, partial ordering may arise from experimental dependencies, where a molecular design must meet one or more criteria in one property (e.g., pass a certain threshold) before its other properties can be measured. In the context of antibody design, a design candidate is an amino acid sequence that represents an antibody that must first be expressed in cell culture. If the expression level does not exceed a certain threshold in terms of mass per volume, the lab cannot produce it in viable quantities and it cannot be assayed for other properties, such as binding affinity to the target antigen. A partial ordering that captures this experimental dependency might take the form: expression → affinity. Experimental dependencies like this create an asymmetry between the targets; it greatly reduces the information content of a molecular design that does not express compared to a design that does not bind, because a design that does not express cannot provide measurements of binding.
[0104] Alternatively, a partial ranking may encode a preference for a type of molecule design. For example, one or more properties may be prioritized such that a molecule design that performs poorly in these properties may be rejected regardless of how well it performs in other properties. In the context of antibody design, if a molecule design does not bind to the target antigen, its primary function has failed, and there is little interest in exploring its developability properties, such as specificity for the target antigen and thermal stability, although these developability properties are often still measurable, unlike non-expressors. A partial ranking that captures this preference may take the form: expression → affinity → {specificity, thermal stability}.
[0105] Reference again Figure 4 An example of a hierarchy 400 shown in FIG. 4 , where a partial ordering of some features takes precedence over other features can be represented as an ordered set of features: where y l,m represents the properties of level l∈{0,...,–-1} of the hierarchy 400, and m∈{0,...,M l– –1} is its M at the same level l l Indexes between sibling features. Various instances of multi-objective optimization (MOO) described herein, including the aforementioned Bayesian optimization, can include prioritizing features on level l over those on a subsequent level l+1, such that if a level l feature of a molecular design fails to meet a corresponding criterion, then a molecular design whose level l+1 meets its corresponding criterion is rejected.
[0106] Biological properties tend to carry excess zero or null values. That is, zero may be the most prevalent value for a biological property such as expression, binding affinity, lack of nonspecificity, stability, lack of immunogenicity, human nature, lack of self-association, etc. For example, in the case of antibody design, a large fraction of the molecule designs may not be expressed at all, which contributes to the high incidence of zero values (or null values) for that particular property. The zero-inflated nature of biological properties motivates the use of statistical models to explain the large incidence of zeros. For example, in some exemplary embodiments, one or more property computation models 310 may be implemented as a zero-inflated probabilistic proxy model. Thus, for each target (or property) y k , the probabilistic binary classifier 313 of the feature calculation model 310 can assign a binary random variable b k ∈{0,1} to indicate the characteristic y k The existence (or non-existence) of k The probability regressor model 115 can be used for the same feature y k ,generate It corresponds to the remaining discrete values of the continuous non-zero values of . In some cases, the probabilistic binary classifier 313 can be based on the feature yk Whether one or more thresholds are met to assign a binary random variable b k ∈{0,1}. Therefore, instead of just directly indicating the characteristic y k It should be understood that the output of the probability binary classifier 313 may indicate whether the feature y k Whether sufficient quantities or levels exist.
[0107] In one example where the property computation model 310 is trained to predict the expression level of a molecular design, the output of the property computation model 310 may include a first value determined by the probabilistic binary classifier 313 to indicate whether the expression level of the molecular design meets one or more thresholds. In addition, the output of the property computation model 310 may include a second value determined by the probabilistic regressor model 115 indicating the expression level of the molecular design if it is determined that the expression level of the molecular design meets one or more thresholds (e.g., a value assigned by the probabilistic binary classifier 313 of 1).
[0108] For further explanation, assume that f is non-negative (which is equivalent to it being bounded from below). Given a dataset D available at time t t , the probability binary classifier 313 can classify the target (or feature) y k The marginal predictive posterior p(b k |x) is modeled as a non-zero value.
[0109]
[0110] where Φ:R→(0,1) corresponds to the cumulative distribution function (CDF) of the standard normal distribution. The first term in the integral shown in equation (4) may be a Bernoulli distribution controlling for aleatoric uncertainty, while the second term may control for epistemic uncertainty.
[0111] At the same time, the probability regressor model 315 can be used to analyze the feature y k The marginal predictive posterior modeling of the non-zero mode (or non-zero value) of is shown in Equation (5).
[0112]
[0113] The probabilistic regressor model 315 can be used in the dataset Since Gaussian processes (GP) can be used as an alternative to Bayesian optimization (BO), and the common Gaussian process assumption fails for sparse multimodal data, separating the non-zero modes of the data can improve posterior inference.
[0114] Given the above, each target (or feature) y kThe marginal prediction posterior on can be a weighted mixture of the delta function and the above formula (5), where the relative weights on the latter are provided by the Bernoulli parameters in formula (4). Assuming p(r k =0|x)=0 is shown in the following formula (6). It should be understood that although for a single target y k Zero-inflated modeling is described, but zero-inflated modeling can be extended to multiple targets (or features), for example, using multi-task Gaussian processes (Gp) The joint posterior on .
[0115]
[0116] The above framework is expressed in terms of a zero-inflated continuous-valued target (a mixture of a delta function at zero and a continuous distribution), and is applied to binary targets and continuous-valued targets without zero inflation, which can be viewed as respectively converting p(r k |x,D t ,θ r )=p(r k )=N(0,σ 2 ) takes very large σ and p(b k =1|x,D t ,θ b )=1 is a special case.
[0117] In some exemplary embodiments, by Figure 3 As shown in the resampling 330 in FIG. 1 , the selection engine 120 may modify the plurality of intermediate posterior samples 320 output from one or more feature computation models 310 (e.g., zero-inflated probability proxy models) to further enforce hierarchical (e.g., parent-child) relationships between various features. Consider feature y k and its predecessor or parent node par(k). Consider an intermediate posterior sample from the plurality of intermediate posterior samples 320, which is calculated by one or more characteristic calculation models for a node with β k′ ~p(b k′ |x)∈{0,1} and The single molecule design output of , where for each k′∈{1,...,K}. Without any modification, Equation (7) will produce y k The following sample γ k :
[0118]
[0119] However, the sample γ k is agnostic to any hierarchy, e.g. Figure 4In some cases, the output of the probabilistic binary classifier from the parent node par(k) can be combined with the output of the probabilistic binary classifier and the output of the probabilistic binary classifier for the feature y. k The corresponding probabilistic regressor models of are combined (e.g., multiplicatively, etc.) to enforce the dependencies existing in the hierarchy 400. Therefore, if the output of the probabilistic binary classifier is "0" from one or more parent nodes par(k), indicating that the corresponding molecular design fails to exhibit certain characteristics that occupy the parent node par(k), then the intermediate posterior sample γ k are excluded from the contribution to the utility metric of the corresponding molecular design, such as expected hypervolume improvement (EHVI).
[0120] To further illustrate, for example, the selection engine 120 may alternatively start at the top level of the hierarchy 400 and proceed down the levels therein to select a k With its predecessor or parent attribute {b k′} k′∈par(k) If y k is a top-level feature, then and Otherwise, y k has the parent attribute, and The definition is as follows:
[0121]
[0122] The modified binary sample can then be used To obtain y k Valid samples
[0123]
[0124] Let γ: = [γ 1 ,…,γ K ]∈R K and And convert the sample levels described in formulas (8) and (9) into h:R K →R K , so that
[0125] The selection engine 120 may repeat the resampling 330 on other intermediate a posteriori samples 320 associated with the molecular design. Each intermediate a posteriori sample is represented as Then, each corresponding modified sample vector It can be used to evaluate multiple objective acquisition functions 350 (e.g., expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI) (Formula (4)), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc.) by, for example, Monte Carlo (MC) integration. More specifically, assume that S intermediate posterior samples are drawn simultaneously for designing candidate x * (reflecting chance and epistemic uncertainty) and previously observed designs (reflecting the accidental uncertainty), and for s = ..., S, each sampling is expressed as and Then, the Monte Carlo approximation of the multi-objective acquisition function (e.g., Noise Expected Hypervolume Improvement (NEHVI) in some cases) can be efficiently evaluated as:
[0126]
[0127] in and
[0128] Example tasks for multi-objective optimization (MOO)
[0129] The performance of the selection engine 120 is evaluated by simulated active experiments on two synthetic tasks and one real-world antibody design task. For these example use cases, the noise expected hypervolume improvement (NEHVI) (Formula (4)) is used as the multi-objective acquisition function 350. In addition, the multi-objective acquisition function 350 is evaluated by Monte Carlo integration (Formula (10)). Each experiment tests three types of acquisition: (1) batch multi-objective Bayesian optimization (BO) with partial ordering of features ("qNEHVI-DAG"), (2) batch multi-objective Bayesian optimization without any feature ordering ("qNEHVI") and (3) random. The main performance metric is the number of "jointly positive" molecular designs obtained, which, as mentioned above, refers to molecular designs that meet certain criteria (e.g., exceed a selected threshold) in each objective based on the specified partial ordering of features. Here, the size of each batch is denoted as q.
[0130] Example Task I: Penicillin Production
[0131] The first penicillin production task is based on a penicillin production simulator. The two objectives in this example can be defined as reducing or minimizing carbon dioxide (CO 2 ) by-product discharge, while ensuring that the fermentation time is below the set threshold and the output exceeds the set threshold (K = 3, X = R 7 ). The latter two objectives can be negated to define a maximization problem and assume a feature hierarchy {y0.0}→{y 1.0}→{y 2.0}, where y 0.0 = output ("target 0"), y 1.0 = negative fermentation time ("target 1") and y 2.0 = Negative CO 2 Byproduct ("Target 2"). Zero-mean Gaussian noise is added to the input.
[0132] The exact Gaussian process (GP) may be suitable for r k To model , an approximate Gaussian process may be suitable for each target k separately, using the variational evidence lower bound (ELBO) on b k In this example task, 512 posterior samples are drawn to evaluate qNEHVI.
[0133] Ten rounds of simulated active learning can be performed by initializing one or more feature calculation models 310 with 8 training points and selecting q=4 from a random sampling pool of 80 candidate points in each iteration. The three acquisition modes (qNEHVI-DAG, qNEHVI, and random) accept the same initial training points and candidate pool in each round. The entire experiment is repeated five times. Figure 7 (A) shows that qNEHVI-DAG identifies significantly more joint positives than qNEHVI and random overactivity learning iterations. Figure 8 Comparison of qNEHVI and qNEHVI-DAG selections for each pair of targets stacked in 5 replicates for final (after 10 iterations) selection. For each target, qNEHVI-DAG identifies more thresholds (black dashed lines) than qNEHVI and randomness.
[0134] Example Mission II: Branin-Currin
[0135] The next task is based on the analytical Branin-Currin test function, from which X = R 2 and K = 2. The Branin-Currin task is configured to simulate the antibody design task in a controlled environment. Here, the hierarchy of features can be defined as {y 0.0}→{y 1.0}, where y 0.0 = dimension 0 (“target 0”) and y 1.0 = Dimension 1 ("Target 1"). Target 0 is converted to binary value using a set threshold, while target 1 is zero-inflated and real-valued. A posteriori inference is performed following a similar procedure as for the penicillin production task.
[0136] Here, the selection engine 120 performs 20 rounds of simulated active learning by initializing one or more feature calculation models 310 with 6 training points and selecting q=4 from a random sampling pool of 40 candidate points in an iteration. The entire experiment is repeated 10 times. Figure 7 (B) shows that qNEHVI-DAG identifies more joint positives than qNEHVI and random overactivity learning iterations. Fig. 9 Depicted is a comparison between qNEHVI-DAG, qNEHVI, and random selection for final selection stacked across 10 replicates for target 1. Overall, qNEHVI-DAG identifies more examples to the right of the threshold (black dashed line) than qNEHVI and random, and the improvement is more pronounced for identifying joint positives (middle panel).
[0137] Example Task III: Antibody Design
[0138] The antibody design task is derived from a dataset of real-world antibody sequences and their in vitro measured properties in terms of affinity and expression. As in the toy problem, the hierarchy of properties is defined as {y 0.0}→{y 1.0}, where y 0.0 =expression("target 0") and y 1.0 = affinity("target 1"). Target 0 is a binary value (eg, expressed or not expressed), while target 1 is zero-inflated and real-valued.
[0139] The selection engine 120 performs 3 simulated active learning iterations and repeats the entire process 5 times. To simulate active learning, the entire dataset of 4022 variable-length protein sequences designed as antibodies against anonymous target antigen A is divided into 5 groups of sizes 1230, 736, 746, and 600, respectively. The first group is used as an initial training set for one or more property computational models 310, while the next three groups are used as "candidate pools" from which 200 molecular designs are selected during each iteration. The last remaining group is used as a retained test set. Fig.10 As shown in (A), qNEHVI-DAG again outperformed qNEHVI and random in terms of the number of combined positives. Fig.10 (B) The logarithmic posterior density evaluated at the affinity measurements for co-positives (expressed binders) in the test set is also highest for qNEHVI-DAG, indicating that the feature calculation model 310 from qNEHVI-DAG has the most accurate belief for co-positives after the final iteration.
[0140] In view of the above specific implementation of the subject matter, the present application discloses the following list of examples, wherein one feature of a single example or a combination of more than one feature of the example, and optionally a combination with one or more features of one or more other examples, are other examples that also fall within the disclosure scope of the present application:
[0141] Item 1: A computer-implemented method comprising: applying one or more trained property computational models to a first molecular design to determine a first probability that the first molecular design exhibits a first property and a second probability that the first molecular design exhibits a second property; determining, based at least on outputs of the one or more property computational models, a first plurality of samples associated with the first molecular design, each sample in the first plurality of samples comprising a first value of the first property exhibited by the first molecular design and a second value of the second property exhibited by the first molecular design having the first value for the first property; within the first plurality of samples, identifying a first group of samples in which the first value of the first property meets a first criterion; determining, based at least on the first group of samples, a first utility metric corresponding to a first expected improvement of the first property and the second property of the first molecular design relative to the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.
[0142] Item 2: A method according to Item 1, wherein the first utility metric is determined by applying expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO) or joint entropy search (JES).
[0143] Item 3: A method according to any one of Items 1 to 2, wherein the first property and the second property of the one or more baseline molecules are determined based on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecule designs.
[0144] Item 4: A method according to any one of Items 1 to 3, further comprising: retraining the one or more property computational models based at least on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecular designs; and applying the one or more retrained property computational models to determine the first property and the second property of the one or more baseline molecules.
[0145] Item 5: The method of any one of Items 1 to 4, further comprising: within the first plurality of samples, identifying a second set of samples in which the first value of the first characteristic fails to meet the first criterion.
[0146] Item 6: The method of item 5, wherein the first utility metric is determined to include a first contribution from the first set of samples and exclude a second contribution from the second set of samples.
[0147] Item 7: A method according to any one of Items 1 to 6, further comprising: applying the one or more trained property computational models to the first molecular design to determine a third probability that the first molecular design exhibits a third property; determining the first plurality of samples based at least on the output of the one or more property computational models to further include a third value of the third property exhibited by the first molecular design as part of each sample; identifying the first group of samples further based on the first value of the first property satisfying the first criterion and the second value of the second property satisfying the second criterion; and determining the first utility metric based at least on the first group of samples to further correspond to the first expected improvement in the first property, the second property and the third property of the first molecular design relative to the first property, the second property and the third property of the one or more baseline molecules.
[0148] Item 8: A method according to Item 7, wherein the first characteristic and the second characteristic occupy the same level in a hierarchy above the third characteristic, so that before the first molecular design is evaluated for the third characteristic, the first molecular design needs to meet the first criterion associated with the first characteristic and the second criterion associated with the second characteristic.
[0149] Item 9: A method according to Item 7, wherein the first characteristic and the second characteristic occupy different levels in a hierarchy above the third characteristic, so that before the first molecular design is evaluated for the second characteristic, the first molecular design needs to meet the first criterion associated with the first characteristic, and wherein before the first molecular design is evaluated for the third characteristic, the first molecular design needs to further meet the second criterion associated with the second characteristic.
[0150] Item 10: The method according to any one of Items 7 to 9, wherein each of the first property, the second property and the third property is a different property among expression, binding affinity, specificity and thermal stability.
[0151] Item 11: The method of Item 1, wherein the one or more property computational models include a first property computational model trained to determine the first probability that the first molecular design exhibits the first property.
[0152] Item 12: A method according to Item 11, wherein the first characteristic calculation model includes a trained first probability binary classifier, which is trained to determine the first probability that the first molecule exhibits the first characteristic, and wherein the first probability binary classifier outputs a first value when the first probability meets a second threshold, and outputs a second value when the first probability fails to meet the second threshold.
[0153] Item 13: The method of any one of Items 11 to 12, wherein the first property computational model comprises a first probabilistic regressor trained to determine the first value of the first property exhibited by the first molecular design.
[0154] Item 14: A method according to any one of Items 11 to 13, wherein the one or more property calculation models further include a second property calculation model trained to determine the second probability that the first molecule exhibits the second property.
[0155] Item 15: A method according to Item 14, wherein the second characteristic calculation model includes a second binary classifier trained to determine the second probability that the first molecule exhibits the second characteristic, and wherein the second binary classifier outputs a first value when the second probability satisfies a second threshold, and outputs a second value when the second probability fails to satisfy the second threshold.
[0156] Item 16: The method of any one of Items 14 to 15, wherein the second property computational model comprises a second regressor trained to determine the second value of the second property exhibited by the first molecular design.
[0157] Item 17: A method according to any one of Items 1 to 16, wherein the one or more property computational models include an integration of property computational models, and wherein the first probability that the first molecular design exhibits the first property and / or the second probability that the first molecular design exhibits the second property is determined at least based on the output of the integration of the property computational models.
[0158] Item 18: A method according to any one of Items 1 to 17, further comprising: applying the one or more property calculations to a second molecular design to determine a third probability that the second molecular design exhibits the first property and a fourth probability that the second molecular design exhibits the second property; determining a second plurality of samples associated with the second molecular design based at least on the output of the one or more property calculation models, each sample in the second plurality of samples comprising a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design; within the second plurality of samples, identifying a second group of samples in which the third value of the first property meets the first criterion; determining a second utility metric based at least on the second group of samples, the second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design relative to the first property and the second property of the one or more baseline molecular designs; and identifying the second molecular design as another candidate for synthesis based at least on the second utility metric of the second molecular design.
[0159] Item 19: A method according to Item 18, wherein the one or more baseline molecular designs are updated to include the first molecular design, so that the second expected improvement includes the expected improvement in the first characteristic and the second characteristic of the second molecular design relative to the first characteristic and the second characteristic of the first molecular design.
[0160] Item 20: A method according to Item 19, wherein the one or more baseline molecular designs are updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or the second property exhibited by the first molecular design.
[0161] Item 21: The method of any one of Items 19-20, wherein the one or more baseline molecular designs are updated to include an average of the first plurality of samples associated with the first molecular design.
[0162] Item 22: A method according to any one of Items 18 to 21, wherein the second utility metric is determined to include a first contribution from the second group of samples and exclude a second contribution from a third group of samples, wherein the third value of the first characteristic fails to meet the first criterion.
[0163] Item 23: The method according to any one of Items 18 to 22, wherein the first molecular design and the second molecular design are further identified as candidates for batch in vitro and / or in vivo evaluation.
[0164] Item 24: A method according to any one of Items 1 to 23, wherein each of the first probability that the first molecular design exhibits the first characteristic and / or the second probability that the first molecular design exhibits the second characteristic comprises (i) a first probability distribution over a first value indicating that the corresponding characteristic is present in the first molecular design and a second value indicating that the corresponding characteristic is not present in the first molecular design, and (ii) a second probability distribution over a range of possible values indicating the magnitude of the corresponding characteristic exhibited by the first molecular design.
[0165] Item 25: A method according to any one of Items 1 to 24, wherein the first molecular design is identified as a candidate for synthesis based at least on the first utility metric of the first molecular design satisfying one or more thresholds.
[0166] Item 26: A method according to any one of Items 1 to 25, further comprising: selecting a quantity of N number of molecular designs having the highest utility metric as candidates for synthesis, and identifying the first molecular design as a candidate for synthesis based at least on the first molecular design being one of the N number of molecular designs having the highest utility metric.
[0167] Item 27: A method according to any one of items 1 to 26, wherein the first value of the first characteristic satisfies the first criterion by satisfying a threshold, falling within one or more intervals of values, or being a member of a set.
[0168] Item 28: The method according to any one of Items 1 to 27, further comprising: identifying the first molecular design as the candidate for synthesis based at least on the presence or absence of one or more specific amino acid residues in the first molecular design.
[0169] Item 29: A method according to any one of Items 1 to 28, wherein the first plurality of samples includes a distribution of the second value of the second property exhibited by the first molecular design over the first value of the first property exhibited by the first molecular design.
[0170] Item 30: The method according to any one of Items 1 to 29, wherein the first molecular design is a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide and / or a glycolipid.
[0171] Item 31: The method according to any one of items 1 to 30, further comprising: applying a molecular design computational model to generate the first molecular design.
[0172] Item 32: A system comprising: at least one data processor, and at least one memory storing instructions which, when executed by the at least one data processor, cause operations including those of the method described in any one of Items 1 to 31.
[0173] Item 33: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including those of the method described in any one of Items 1 to 31.
[0174] Fig.11 A block diagram illustrating an example of a computing system 1100 is depicted in accordance with some example embodiments. Figures 1 to 11 The computing system 1100 may be used to implement the molecular design engine 110 , the selection engine 120 , the lab-device 130 , the client device 140 , and / or any components thereof.
[0175] like Fig.11 As shown, the computing system 1100 may include a processor 1110, a memory 1120, a storage device 1130, and an input / output device 1140. The processor 1110, the memory 1120, the storage device 1130, and the input / output device 1140 may be interconnected via a system bus 1150. The processor 1110 is capable of processing instructions for execution within the computing system 1100. Such executed instructions may implement, for example, one or more components of the molecular design engine 110, the selection engine 120, the laboratory equipment 130, the client device 140, etc. In some exemplary embodiments, the processor 1110 may be a single-threaded processor. Alternatively, the processor 1110 may be a multi-threaded processor. The processor 1110 is capable of processing instructions stored on the memory 1120 and / or the storage device 1130 to display graphical information for a user interface provided via the input / output device 1140.
[0176] The memory 1120 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within the computing system 1100. For example, the memory 1120 may store a data structure representing a configuration object database. The storage device 1130 is capable of providing persistent storage for the computing system 1100. The storage device 1130 may be a floppy disk device, a hard disk device, an optical disk device, a tape device, or other suitable persistent storage device. The input / output device 1140 provides input / output operations for the computing system 1100. In some exemplary embodiments, the input / output device 1140 includes a keyboard and / or a pointing device. In various specific implementations, the input / output device 1140 includes a display unit for displaying a graphical user interface.
[0177] According to some exemplary embodiments, the input / output device 1140 may provide input / output operations for network devices. For example, the input / output device 1140 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0178] In some exemplary embodiments, the computing system 1100 can be used to execute various interactive computer software applications that can be used to organize, analyze and / or store data in various formats. Alternatively, the computing system 1100 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, editing electronic spreadsheet documents, word processing documents and / or any other objects, etc.), computing functions, communication functions, etc. The application can include various additional functions or can be an independent computing product and / or function. After activation within the application, the function can be used to generate a user interface provided via the input / output device 1140. The user interface can be generated by the computing system 1100 and presented to the user (e.g., on a computer screen monitor, etc.).
[0179] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuits, integrated circuits, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include a specific implementation in one or more computer programs, which are executable and / or interpretable on a programmable system, which includes at least one programmable processor (which may be dedicated or general, coupled to receive data and instructions from it and send data and instructions to it), a storage system, at least one input device, and at least one output device. A programmable system or computing system may include a client and a server. Typically, the client and the server are remotely arranged from each other, and generally interact through a communication network. The relationship between the client and the server is generated by means of a computer program running on each computer and a client-server relationship between each other.
[0180] These computer programs may also be referred to as programs, software, software applications, applications, components or codes, including machine instructions for programmable processors, and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine languages. As used herein, the term "machine-readable medium" refers to any computer product, device and / or equipment (such as, for example, disks, optical disks, memories and programmable logic devices (PLDs)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. A machine-readable medium may store such machine instructions non-temporarily (such as, for example, a non-temporary solid-state memory or a magnetic hard drive or any equivalent storage medium). A machine-readable medium may store such machine instructions in a temporary manner (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores) alternatively or additionally.
[0181] To provide interaction with a user, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to a user) and a keyboard and a pointing device (such as, for example, a mouse or trackball, through which a user can provide input to the computer). Other types of devices may also be used to provide interaction with a user. For example, the feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; input from the user may be received in any form, including sound, voice, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive tracking pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0182] In the above description and claims, phrases such as "at least one" or "one or more" may appear, followed by a list of combinations of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to represent any element or feature listed alone, or any other described element or feature combined with any other described element or feature. For example, the phrase "at least one of A and B"; "one or more of A and B"; "A and / or B" are each intended to represent "single A, single B, or A and B together". A similar interpretation also applies to lists including three or more items. For example, the phrases "at least one of A, B, and C"; "one or more of A, B, and C" and "A, B, and / or C" are each intended to represent "single A, single B, single C, A and B together, A and C together, B and C together, or A and B and C together". The use of the term "based on" above and in the claims is intended to represent "based at least in part", so that undescribed features or elements are also permissible.
[0183] Depending on the desired configuration, the subject matter described herein may be embodied in systems, devices, methods and / or articles. The embodiments described in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Instead, they are only some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, in addition to those features and / or variations described herein, other features and / or variations may also be provided. For example, the above-mentioned specific implementations may be directed to various combinations and sub-combinations of the disclosed features and / or to combinations and sub-combinations of several further features disclosed above. In addition, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other specific implementations may be within the scope of the following claims.
Claims
1. A computer-implemented method, include: applying one or more trained property computational models to a first molecular design to determine a first probability that the first molecular design exhibits a first property and a second probability that the first molecular design exhibits a second property; determining, based at least on outputs of the one or more property computational models, a first plurality of samples associated with the first molecular design, each sample in the first plurality of samples comprising a first value of the first property exhibited by the first molecular design and a second value of the second property exhibited by the first molecular design having the first value for the first property; within the first plurality of samples, identifying a first set of samples in which the first value of the first characteristic meets a first criterion; determining, based at least on the first set of samples, a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design relative to the first property and the second property of one or more baseline molecular designs; as well as The first molecular design is identified as a candidate for synthesis based at least on the first utility metric of the first molecular design.
2. The method of claim 1, wherein the first utility metric is determined by applying expected hypervolume improvement (EHVI), noise expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), or joint entropy search (JES).
3. The method according to any one of claims 1 to 2, further comprising: include: retraining the one or more property computational models based at least on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecular designs; as well as One or more retrained property calculation models are applied to determine the first property and the second property of the one or more baseline molecules.
4. The method according to any one of claims 1 to 3, further comprising: include: identifying, within the first plurality of samples, a second set of samples in which the first value of the first characteristic fails to meet the first criterion; as well as The first utility metric is determined to include a first contribution from the first set of samples and to exclude a second contribution from the second set of samples.
5. The method according to any one of claims 1 to 4, further comprising: include: applying the one or more trained property computational models to the first molecular design to determine a third probability that the first molecular design exhibits a third property; determining, based at least on the outputs of the one or more property computation models, the first plurality of samples to further include as part of each sample a third value of the third property exhibited by the first molecular design; identifying the first set of samples further based on the first value of the first characteristic satisfying the first criterion and the second value of the second characteristic satisfying a second criterion; as well as Based at least on the first set of samples, the first utility metric is determined to further correspond to the first expected improvement of the first property, the second property, and the third property of the first molecule design relative to the first property, the second property, and the third property of the one or more baseline molecules.
6. The method of claim 5 , wherein the first characteristic and the second characteristic occupy the same level in a hierarchy above the third characteristic, such that the first molecular design is required to satisfy the first criterion associated with the first characteristic and the second criterion associated with the second characteristic before the first molecular design is evaluated for the third characteristic.
7. A method according to claim 5, wherein the first characteristic and the second characteristic occupy different levels in a hierarchy above the third characteristic, so that before the first molecular design is evaluated for the second characteristic, the first molecular design needs to meet the first criterion associated with the first characteristic, and wherein before the first molecular design is evaluated for the third characteristic, the first molecular design needs to further meet the second criterion associated with the second characteristic.
8. The method of any one of claims 1 to 7, wherein the one or more property computational models comprises a first property computational model trained to determine the first probability that the first molecular design exhibits the first property.
9. The method of claim 8, wherein the first property computation model comprises a first probabilistic binary classifier trained to output a first value when the first probability satisfies a second threshold and to output a second value when the first probability fails to satisfy the second threshold, and wherein the first property computation model further comprises a first probabilistic regressor trained to determine the first value of the first property exhibited by the first molecular design.
10. The method of any one of claims 8 to 9, wherein the one or more property calculation models further comprises a second property calculation model trained to determine the second probability that the first molecule exhibits the second property.
11. The method of claim 10, wherein the second property calculation model comprises a second binary classifier trained to output a first value when the second probability satisfies a second threshold and to output a second value when the second probability fails to satisfy the second threshold, and wherein the second property calculation model further comprises a second regressor trained to determine the second value of the second property exhibited by the first molecular design.
12. The method according to any one of claims 1 to 11, wherein the one or more property computational models comprise an integration of property computational models, and wherein the first probability that the first molecular design exhibits the first property and / or the second probability that the first molecular design exhibits the second property are determined based at least on an output of the integration of the property computational models.
13. The method according to any one of claims 1 to 12, further comprising: include: computationally applying the one or more properties to a second molecular design to determine a third probability that the second molecular design exhibits the first property and a fourth probability that the second molecular design exhibits the second property; determining, based at least on the outputs of the one or more property computational models, a second plurality of samples associated with the second molecular design, each sample in the second plurality of samples comprising a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design; within the second plurality of samples, identifying a second set of samples in which the third value of the first characteristic meets the first criterion; determining, based at least on the second set of samples, a second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design relative to the first property and the second property of the one or more baseline molecular designs; as well as Based at least on the second utility metric of the second molecular design, the second molecular design is identified as another candidate for synthesis.
14. The method of claim 13, wherein the one or more baseline molecular designs are updated to include the first molecular design such that the second expected improvement includes an expected improvement in the first property and the second property of the second molecular design relative to the first property and the second property of the first molecular design.
15. The method of claim 14, wherein the one or more baseline molecular designs are updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or the second property exhibited by the first molecular design.
16. The method of any one of claims 14 to 15, wherein the one or more baseline molecular designs are updated to include an average value of the first plurality of samples associated with the first molecular design.
17. The method of any one of claims 1 to 16, wherein each of the first probability that the first molecular design exhibits the first characteristic and / or the second probability that the first molecular design exhibits the second characteristic comprises (i) a first probability distribution over a first value indicating that the corresponding characteristic is present in the first molecular design and a second value indicating that the corresponding characteristic is not present in the first molecular design, and (ii) a second probability distribution over a range of possible values indicating the magnitude of the corresponding characteristic exhibited by the first molecular design.
18. The method of any one of claims 1 to 17, wherein the first molecular design is identified as the candidate for synthesis based at least on the first utility metric of the first molecular design satisfying one or more thresholds.
19. The method according to any one of claims 1 to 18, further comprising: include: N number of molecular designs having the highest utility metric are selected as candidates for synthesis, the first molecular design being identified as the candidate for synthesis based at least on the first molecular design being one of the N number of molecular designs having the highest utility metric.
20. The method according to any one of claims 1 to 19, further comprising: include: The first molecular design is identified as the candidate for synthesis based at least on the presence or absence of one or more specific amino acid residues in the first molecular design.
21. The method of any one of claims 1 to 20, wherein the first plurality of samples comprises a distribution of the second value of the second property exhibited by the first molecular design over the first value of the first property exhibited by the first molecular design.
22. A system, wherein include: at least one data processor; as well as At least one memory storing instructions which, when executed by said at least one data processor, result in operations including the method according to any one of claims 1 to 21.
23. A non-transitory computer readable medium storing instructions which, when executed by at least one data processor, result in operations including the method of any one of claims 1 to 21.