Molecular design involving multi-objective optimization of partially ordered and mixed variable molecular properties.

The molecular design system uses multi-objective active learning to optimize the selection of molecular designs, addressing computational intractability and resource limitations by prioritizing certain properties, enhancing the efficiency and effectiveness of in vitro and in vivo evaluations.

JP2025534388APending Publication Date: 2025-10-15GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025518818
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-10-03
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Designing molecules, particularly proteins, is computationally intractable due to vast combinatorial search spaces and resource limitations in wet laboratory evaluations, leading to inefficient selection of molecular designs for in vitro and in vivo testing.

Method used

A molecular design system employing multi-objective active learning techniques, including probabilistic surrogate models and utility functions, to select co-positive molecular designs that meet multiple property criteria, prioritizing certain properties over others, thereby optimizing the selection process for in vitro and in vivo evaluation.

Benefits of technology

Enhances the likelihood of selecting molecular designs with improved properties, reducing the risk of suboptimal candidates and increasing the efficiency of laboratory resource utilization by identifying candidates that meet multiple criteria across successive design iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025534388000001_ABST
    Figure 2025534388000001_ABST
Patent Text Reader

Abstract

One or more property computational models may be applied to determine a first probability of a molecular design exhibiting a first property and a second probability of a molecular design exhibiting a second property. A plurality of samples, each sample comprising a first value of the first property and a second value of the second property exhibited by the molecular design, may be determined based on the output of the property computational models. A set of samples may be identified in which the first value of the first property meets a criterion. A utility metric corresponding to an expected improvement of the first property and the second property of the molecular design over the first property and the second property of a baseline molecular design may be determined based on the set of samples. One or more molecular designs may be identified as candidates for synthesis based on the corresponding utility metric.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 378,186, filed October 3, 2022, entitled "Multi-Objective Active Learning for Molecular Design," and U.S. Provisional Application No. 63 / 385,609, filed November 30, 2022, entitled "Multi-Objective Active Learning for Molecular Design," the disclosures of which are incorporated herein by reference in their entireties.

[0002] Technical Field The subject matter described herein relates generally to molecular design, and more particularly to multi-objective active learning techniques for molecular design. [Background technology]

[0003] introduction A molecule is a group of two or more atoms held together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of the substance. An example of a molecule is a protein molecule, and examples of non-protein molecules include small molecules, ions, nucleic acids, polysaccharides, glycolipids, etc. The function and properties of a molecule can depend on its three-dimensional structure. For example, proteins are responsible for many essential cellular functions, including, for example, enzymatic reactions, molecular transport, regulation and execution of several biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, etc.

[0004] A protein structure may include one or more polypeptides, which are chains of amino acid residues linked together by peptide bonds. The sequence of amino acid residues in the polypeptide chains forming the protein structure determines the three-dimensional structure of the protein (e.g., the protein's tertiary structure). Furthermore, the sequence of amino acids in the polypeptide chains forming the protein determines the underlying function of the protein. Therefore, one goal of protein design may include constructing one or more sequences of amino acid residues that exhibit various desirable properties. For example, in the case of large molecule drug discovery, de novo protein design often seeks to identify sequences of amino acid residues (e.g., antibodies) that can bind to a target antigen (e.g., a viral antigen, a tumor antigen, etc.), including by adopting a three-dimensional structure that complements the three-dimensional structure of the target antigen. Summary of the Invention

[0005] overview Systems, methods, and articles of manufacture, including computer program products, are provided for molecular design with multi-objective active learning. In one aspect, a system for molecular design with multi-objective active learning is provided. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that, when executed by the at least one data processor, cause an operation. The operations may include applying one or more property computational models to the first molecular design, the property computational models being trained to determine a first probability of the first molecular design exhibiting a first property and a second probability of the first molecular design exhibiting a second property; determining a first plurality of samples associated with the first molecular design based at least on output of the one or more property computational models, each sample of the first plurality of samples including a first value for a first property exhibited by the first molecular design and a second value for a second property exhibited by the first molecular design having a first value for the first property; identifying a first set of samples within the first plurality of samples whose first values ​​for the first property meet a first criterion; determining a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design over the first property and the second property of one or more baseline molecular designs based at least on the first utility metric of the first molecular design; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.

[0006] In another aspect, a method for molecular design involving multi-objective optimization is provided. The method may include applying to the first molecular design one or more property computational models trained to determine a first probability of the first molecular design exhibiting a first property and a second probability of the first molecular design exhibiting a second property; determining a first plurality of samples associated with the first molecular design based at least on output of the one or more property computational models, each sample of the first plurality of samples comprising a first value for a first property exhibited by the first molecular design and a second value for a second property exhibited by the first molecular design having a first value for the first property; identifying a first set of samples within the first plurality of samples whose first values ​​for the first property meet a first criterion; determining, based at least on the first set of samples, a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design over the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.

[0007] In another aspect, a non-transitory computer program product for molecular design with multi-objective active learning is provided. The non-transitory computer program product may store instructions that, when executed by at least one data processor, cause operations to occur. The operations may include applying one or more property computational models to the first molecular design, the property computational models being trained to determine a first probability of the first molecular design exhibiting a first property and a second probability of the first molecular design exhibiting a second property; determining a first plurality of samples associated with the first molecular design based at least on output of the one or more property computational models, each sample of the first plurality of samples including a first value for a first property exhibited by the first molecular design and a second value for a second property exhibited by the first molecular design having a first value for the first property; identifying a first set of samples within the first plurality of samples whose first values ​​for the first property meet a first criterion; determining a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design over the first property and the second property of one or more baseline molecular designs based at least on the first utility metric of the first molecular design; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.

[0008] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may be included, optionally in any possible combination.

[0009] In some variations, the first utility metric may be determined by applying expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), or joint entropy search (JES).

[0010] In some variations, the one or more property computational models may be retrained based at least on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecular designs, and the one or more retrained property computational models may be applied to determine the first property and the second property of the one or more baseline molecules.

[0011] In some variations, a second set of samples may be identified within the first plurality of samples in which the first value of the first characteristic does not meet the first criterion, and a first utility metric may be determined to include a first contribution from the first set of samples and exclude a second contribution from the second set of samples.

[0012] In some variations, one or more property computational models may be applied to the first molecular design, the property computational models being trained to determine a third probability of the first molecular design exhibiting a third property. A first plurality of samples may be determined based at least on the output of the one or more property computational models, such that each sample further includes a third value of the third property exhibited by the first molecular design. A first set of samples may further be identified based on first values ​​of the first property that satisfy a first criterion and second values ​​of the second property that satisfy a second criterion. A first utility metric may be determined based at least on the first set of samples, further corresponding to a first expected improvement in the first property, the second property, and the third property of the first molecular design over the first property, the second property, and the third property of one or more baseline molecules.

[0013] In some variations, the first property and the second property may occupy the same level in the hierarchy above the third property, such that the first molecular design must satisfy a first criterion associated with the first property and a second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

[0014] In some variations, the first property and the second property may occupy different levels in a hierarchy above a third property, such that a first molecular design must satisfy a first criterion associated with the first property before the first molecular design is evaluated with respect to the second property. The first molecular design may further be required to satisfy a second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

[0015] In some variations, the one or more property computational models may include a first property computational model that is trained to determine a first probability of a first molecular design exhibiting a first property.

[0016] In some variations, the first property computational model may include a first probabilistic binary classifier trained to output a first value when the first probability meets a second threshold and to output a second value when the first probability does not meet the second threshold. The first property computational model may further include a first probabilistic regressor trained to determine a first value for the first property exhibited by the first molecular design.

[0017] In some variations, the one or more property computational models may further include a second property computational model trained to determine a second probability of the first molecule exhibiting a second property.

[0018] In some variations, the second property computational model may include a second binary classifier trained to output a first value when the second probability meets a second threshold and to output a second value when the second probability does not meet the second threshold. The second property computational model may further include a second regressor trained to determine a second value for a second property exhibited by the first molecular design.

[0019] In some variations, the one or more property computational models may include an ensemble of property computational models, and a first probability of the first molecular design exhibiting the first property and / or a second probability of the first molecular design exhibiting the second property may be determined based at least on an output of the ensemble of property computational models.

[0020] In some variations, one or more property calculations may be applied to the second molecular design to determine a third probability of the second molecular design exhibiting the first property and a fourth probability of the second molecular design exhibiting the second property. A second plurality of samples associated with the second molecular design may be determined based at least on the output of one or more property calculation models. Each sample of the second plurality of samples may include a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design. A second set of samples within the second plurality of samples may be identified in which the third value of the first property satisfies a first criterion. A second utility metric may be determined based at least on the second set of samples, the second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design over the first property and the second property of one or more baseline molecular designs. The second molecular design may be identified as an alternative candidate for synthesis based at least on the second utility metric of the second molecular design.

[0021] In some variations, one or more baseline molecular designs may be updated to include the first molecular design, such that the second expected improvement includes an expected improvement in the first property and the second property of the second molecular design over the first property and the second property of the first molecular design.

[0022] In some variations, one or more baseline molecular designs may be updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or second property exhibited by the first molecular design.

[0023] In some variations, one or more baseline molecular designs may be updated to include an average of the first plurality of samples associated with the first molecular design.

[0024] In some variations, the first probability of the first molecular design exhibiting the first property and / or the second probability of the first molecular design exhibiting the second property may each include (i) a first probability distribution over a first value indicating that the corresponding property is present in the first molecular design and a second value indicating that the corresponding property is absent in the first molecular design, and (ii) a second probability distribution over a range of possible values ​​indicating the magnitude of the corresponding property exhibited by the first molecular design.

[0025] In some variations, a first molecular design may be identified as a candidate for synthesis based at least on a first utility metric of the first molecular design satisfying one or more thresholds.

[0026] In some variations, the one with the highest usefulness metric TIFF2025534388000002.tif4170 molecular designs may be selected as candidates for synthesis. A first molecular design may be identified as a candidate for synthesis based at least on the first molecular design being one of the N molecular designs having the highest utility metric.

[0027] In some variations, a first molecular design may be identified as a candidate for synthesis based at least on the presence or absence of one or more particular amino acid residues in the first molecular design.

[0028] In some variations, the first plurality of samples may include a distribution of second values ​​of a second property exhibited by the first molecular design across a first value of a first property exhibited by the first molecular design.

[0029] Implementations of the present subject matter can include, but are not limited to, methods according to the description provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, or store one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods consistent with one or more implementations of the present subject matter can be implemented by one or more data processors present in a single computing system or in multiple computing systems. Such multiple computing systems can be connected, e.g., to exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of the multiple computing systems, etc.

[0030] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the presently disclosed subject matter are described for illustrative purposes in connection with the design of biological sequences such as protein molecules, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure define the scope of the protected subject matter. [Brief explanation of the drawings]

[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.

[0032] [Figure 1] 1 depicts a system diagram illustrating an example of a molecular design system, according to some exemplary embodiments.

[0033] [Figure 2A] 1 depicts a flowchart illustrating an example of a process for molecular design involving multi-objective optimization of partially ordered and mixed variable properties, according to some illustrative embodiments.

[0034] [Figure 2B] 1 depicts a flowchart illustrating another example of a process for molecular design involving multi-objective optimization of partially ordered and mixed variable properties, according to some illustrative embodiments.

[0035] [Figure 3] FIG. 1 depicts a block diagram showing a molecular design pipeline with partially ordered and mixed variable property multi-objective optimization, according to some exemplary embodiments.

[0036] [Figure 4] 1 depicts a schematic diagram showing an example of a hierarchy related to molecular design properties, according to some exemplary embodiments;

[0037] [Figure 5] 10 depicts a graph showing the effect of post-surrogate resampling on an example acquisition function, according to some exemplary embodiments.

[0038] [Figure 6A] 1 depicts a flowchart illustrating an example of a process for determining the probability of a molecular design exhibiting a property, according to some exemplary embodiments.

[0039] [Figure 6B] 1 depicts a flowchart illustrating another example of a process for determining the probability of a molecular design exhibiting a property, according to some illustrative embodiments.

[0040] [Figure 7] 1 depicts a graph showing the change in the number of joint positive molecular designs over multiple active learning iterations, according to some exemplary embodiments.

[0041] [Figure 8] 1 depicts a graph showing a pairwise Pareto front visualization of an example penicillin production task, according to some exemplary embodiments.

[0042] [Figure 9] 1 depicts a graph showing the distribution of example molecular designs selected as candidates for synthesis and testing, according to some exemplary embodiments.

[0043] [Figure 10] 1 depicts a graph showing the log posterior density of the number of co-positive molecular designs and binding affinity in an example antibody design task, according to some exemplary embodiments.

[0044] [Figure 11] 1 depicts a block diagram illustrating an example of a computing system according to some illustrative embodiments.

[0045] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION

[0046] Detailed Description Designing molecules, including biological sequences such as proteins or nonbiological small molecules, requires exploring a vast combinatorial design space. For example, de novo protein design aims to identify protein sequences (e.g., sequences of amino acid residues) that exhibit homogeneity of desired properties, such as expression, binding affinity for another molecule (e.g., a viral antigen, a tumor antigen, etc.), lack of nonspecificity, stability, lack of immunogenicity, humanness, and lack of self-association. De novo protein design is a particularly difficult and resource-intensive task, at least because the combinatorial search space of all possible substitutions of amino acid residues that can form protein structures is vast but sparsely populated by sequences of amino acid residues that actually correspond to functional proteins. That is, the majority of protein sequences in the combinatorial search space do not exhibit any function at all, let alone any combination of the aforementioned desired properties. Exploring this vast combinatorial search space becomes even more computationally intractable when considering candidate protein sequences of variable length (e.g., candidate protein sequences formed by different amounts of amino acid residues). Thus, even when performed in silico, brute force methods that indiscriminately examine every possible sequence of amino acid residues to identify sequences that exhibit desired properties are too computationally expensive to be a viable solution.

[0047] In some exemplary embodiments, instead of exploring a vast combinatorial space sparsely populated with functional molecules, a molecular design engine may generate one or more molecular designs, including, for example, protein molecules, small molecules, ions, nucleic acids, polysaccharides, glycolipids, etc., by sampling data distributions associated with various known molecules. For example, the molecular design engine may include a molecular design computational model trained using known molecules, including, for example, molecules known to exhibit specific functions and molecules without known functions. In doing so, the molecular design computational model may learn a data distribution corresponding to a reduced-dimensional representation of the composition and / or structure of the various known molecules. In the case of de novo protein design, for example, this data distribution may correspond to a reduced-dimensional representation of the sequences of amino acid residues that form the various known protein sequences. In some cases, the data distribution may occupy a topological space (e.g., a manifold) occupied by known molecules that describes the relationships that exist between them. While the high dimensionality of data relating to known molecules tends to obscure relationships between populations of molecules with compositional and / or structural similarity, data distributions learned by molecular design computational models may occupy a low dimensional space in which one or more populations of molecules with compositional and / or structural similarity may form identifiable clusters.

[0048] Nevertheless, while computational molecular design models, including the aforementioned computational molecular design models, can accelerate the initial molecular design process, limited wet laboratory resources still impose an obstacle on the speed at which candidate molecular designs can undergo in vitro and in vivo evaluation. In a typical drug development pipeline, before a molecular design can proceed to preclinical development and clinical trials, the molecular design must be validated in vitro and undergo multiple rounds of optimization, where the molecular performance is tested in vivo. While computational molecular design models, such as the aforementioned computational molecular design models, may be capable of generating a large number of molecular designs (e.g., on the order of millions of molecular designs), limited wet laboratory resources do not allow for the synthesis and in vitro evaluation of all molecular designs. Instead, a subset of the molecular designs generated by the computational molecular design models may be selected for in vitro and / or in vivo evaluation.

[0049] Indiscriminate selection of molecular designs for in vitro and / or in vivo evaluation may increase the likelihood of selecting designs with poor molecular properties, such as suboptimal pharmacological and physicochemical properties, which increases the likelihood of failure during subsequent preclinical development and clinical trials, while better candidates are overlooked. In particular, in some cases, molecular designs may be generated and evaluated over successive design iterations, with each design iteration (or design round) involving the generation of one or more new molecular designs to improve upon those from one or more previous design iterations. Thus, molecular designs selected for in vitro and / or in vivo evaluation during a current design iteration should exhibit better molecular properties than those from previous design iterations. Thus, as described in more detail below, the selection engine may perform multi-objective optimization over sets of partially ordered and mixed-variable molecular properties when selecting computationally generated molecular designs for further in vitro and / or in vivo evaluation. Doing so may increase the likelihood that better molecular designs, such as those exhibiting better molecular properties than molecular designs from previous design iterations, will be selected for in vitro and / or in vivo evaluation.

[0050] In some exemplary embodiments, the selection engine may perform multi-objective Bayesian optimization (BO), which utilizes one or more probabilistic surrogate models and utility functions to trade off exploration (evaluating highly uncertain molecular designs) and exploration (evaluating molecular designs that are believed to increase or maximize an objective) in a principled manner. For example, the selection engine may apply one or more property computational models trained to determine a first probability that a first molecular design generated by the molecular design computational model will exhibit a first property. Further, the selection engine may apply one or more property computational models to determine a second probability that the first molecular design will exhibit a second property. In this regard, the one or more property computational models may serve as in silico surrogates for in vitro and / or in vivo evaluations that would be too resource-intensive to apply to all molecular designs generated by the molecular design computational model.

[0051] In some exemplary embodiments, the one or more property computation models may be implemented as one or more zero-inflated probabilistic surrogate models, each including a probabilistic binary classifier and a probabilistic regression model. Thus, in some cases, the output of the one or more property computation models may include a first plurality of predicted samples associated with the first molecular design or predicted samples associated with the first molecular design. Each predicted sample associated with the first molecular design may include a first value for a first property exhibited by the first molecular design and a second value for a second property exhibited by the first molecular design having the first value for the first property. In some cases, multiple property computation models may be applied to generate the first plurality of predicted samples to account for uncertainty in the output of each property computation model. Uncertainty in this context may refer to a level of confidence that the output of a property computation model, such as the value of a property predicted by a property computation model for a molecular design, is accurate. For example, in some cases, multiple property computation models may be applied to determine a first value for a first property exhibited by the first molecular design, while multiple property computation models may be applied to determine a second value for a second property exhibited by the first molecular design. Thus, the first predicted sample may include output from a different property computational model than the property computational model used to generate the second predicted sample. In some cases, the first plurality of samples associated with the first molecular design may form a first distribution of second values ​​of a second property exhibited by the first molecular design across first values ​​of a first property exhibited by the first molecular design. In the case of an antibody design, the output of the one or more property computational models may include, for example, a distribution of levels of binding affinity exhibited by the first molecular design across expression levels exhibited by the first molecular design. That is, in the case of an antibody design, the output of the one or more property computational models may include one or more predicted samples, each including an expression level of the first molecular design and a corresponding binding affinity of the first molecular design.

[0052] In some exemplary embodiments, the selection engine may determine a first utility metric indicating a first magnitude by which a first molecular design improves one or more baseline molecular designs with respect to a first property and a second property. For example, in the case of an antibody design, the first utility metric of a first molecular design may indicate the magnitude by which the first molecular design improves the expression level and binding affinity of the baseline molecular design. Furthermore, in some cases, the selection engine may select a first molecular design as a candidate for synthesis and testing based at least on the first utility metric of the first molecular design. In doing so, the selection engine may ensure that the first molecular design selected as a candidate for synthesis and testing is a so-called co-positive molecular design, which refers to a molecular design that meets certain criteria with respect to a first property and a second property. As described in more detail below, the selection engine may perform multi-objective optimization to select such co-positive molecular designs. It should be understood that a co-positive molecular design is not necessarily a molecular design with the best values ​​in all molecular properties (e.g., maximum expression and maximum binding affinity), since at least such a molecular design may not exist at all. Alternatively, multi-objective optimization (MOO) in this context may involve identifying co-positive molecular designs that exhibit a first value of a first property that cannot be improved without worsening a second value of a second property.

[0053] In some exemplary embodiments, the selection engine may apply an active learning approach in which a first molecular design becomes one of the baseline molecular designs during subsequent design iterations. For example, the selection engine may determine to select a second molecular design generated by a molecular design computational model as the next candidate for synthesis based at least on a second utility metric indicating a second magnitude by which the second molecular design improves a first property and a second property of a baseline molecular design that includes the first molecular design. The values ​​of each of the first property and the second property exhibited by the first molecular design may be determined based on the output of one or more property computational models. Alternatively and / or additionally, the values ​​of each of the first property and the second property exhibited by the first molecule may be determined based on one or more in vitro measurements or in vivo characterizations associated with the first molecular design. In some cases, to account for noise (e.g., measurement error associated with laboratory equipment 130) that may be present in one or more in vitro measurements or in vivo characterizations associated with the first molecular design, the values ​​of each of the first and second characteristics exhibited by the first molecular design may be determined based on the output of one or more property computation models after the one or more property computation models have been updated, for example, by being retrained based on one or more in vitro measurements or in vivo characterizations associated with the first molecular design.

[0054] In some exemplary embodiments, the selection engine may impose a partial ordering, e.g., to prioritize a first property over a second property, when determining a first utility metric associated with a first molecular design. As described above, the first utility metric associated with a first molecular design may indicate a first magnitude by which the first and second properties of the first molecular design improve the first and second properties of one or more baseline molecules. In some cases, the partial ordering of the first and second properties may require that a first property of the first molecular design satisfy a first criterion before the first molecular design is evaluated to determine whether the second property of the first molecular design satisfies the second criterion. In the context of antibody design, for example, the selection engine may require that the expression level of a first molecular design meets one or more criteria before the first molecular design is evaluated for its binding affinity to a target antigen to reflect experimental and / or biological dependencies that may be required for the first molecular design to reach a particular expression level before sufficient quantities of the first molecular design can be synthesized and assayed for other properties, such as binding affinity to the target antigen. Thus, to determine a first utility metric associated with a first molecular design while imposing a partial ordering to favor the first property over the second property, the selection engine may include contributions from first samples in the first distribution where the first value of the first property satisfies one or more criteria, while excluding contributions from second samples in the first distribution where the first value of the first property does not satisfy one or more criteria. In doing so, the first utility metric associated with the first molecular design may indicate a first magnitude by which the first property and the second property of the first molecular design improve the first property and the second property of one or more baseline molecular designs, when the first property of the first molecular design satisfies the one or more criteria.

[0055] In some exemplary embodiments, the selection engine may further evaluate a third property of the first molecular design when selecting the first molecular design as a candidate for synthesis and in vitro measurement and / or in vivo characterization. For example, the selection engine may apply one or more property computation models to determine a third probability of the first molecular design exhibiting the third property. In this case, each predicted sample of the first plurality of predicted samples output by the one or more property computation models may include a third value of the third property exhibited by the first molecular design exhibiting a first value of the first property and a second value of the second property. Furthermore, a first utility metric associated with the first molecular design may indicate how much the first property, the second property, and the third property of the first molecular design improve the first property, the second property, and the third property of one or more baseline molecules.

[0056] In some exemplary embodiments, the partial ordering that prioritizes a first characteristic over a second characteristic may further include a third characteristic. In some cases, the selection engine may impose a partial ordering to prioritize a combination of a first characteristic and a third characteristic over a second characteristic. In this particular scenario, the first characteristic and the third characteristic of a first molecular design may be required to satisfy a first criterion before the first molecular design is evaluated to determine whether the second characteristic of the first molecular design satisfies the second criterion. Alternatively and / or additionally, the selection engine may impose a partial ordering to prioritize a first characteristic over a second characteristic, which is further prioritized over the third characteristic. In this case, the first characteristic of the first molecular design may be required to satisfy a first criterion before the first molecular design is evaluated with respect to the second characteristic, and the second characteristic of the first molecular design may be further required to satisfy a second criterion before the first molecular design is further evaluated to determine whether the third characteristic of the first molecular design satisfies the third criterion. Referring again to the antibody design example, the selection engine may require the expression level of a first molecular design to meet a first criterion before the first molecular design is evaluated for its binding affinity to the target antigen. Furthermore, the binding affinity of the first molecular design may further be required to meet a second criterion before the selection engine evaluates its various developable properties, such as specificity, thermal stability, etc. By including a third property, the selection engine may ensure that the first molecular design selected as a candidate for synthesis and testing is a co-positive molecular design that meets certain criteria for the first property, the second property, and the third property.

[0057] As previously mentioned, in some exemplary embodiments, the selection engine may perform multi-objective optimization, such as multi-objective Bayesian optimization, across multiple partially ordered and mixed variable properties (or objectives). This framework may reflect some scenarios in drug design where a molecular design may need to satisfy a first property (e.g., expression) before being evaluated with respect to a second property (e.g., affinity) and / or a third property (e.g., specificity). As described in more detail below, multi-objective optimization (e.g., multi-objective Bayesian optimization) may include imposing a partial ordering during drug design that prioritizes satisfaction of a first property (e.g., expression) over satisfaction of a second property (e.g., affinity) and / or a third property (e.g., specificity). For example, for each molecular design, the selection engine may modify the posterior probability distribution of each objective (e.g., determined by one or more probabilistic surrogate models) so that the properties exhibited by the molecular design are modeled as zero-inflated distributions (a mixture of zero values ​​and a continuous distribution of non-zero values), and certain properties are prioritized over others. In doing so, the selection engine can identify significantly more co-positive molecular designs, such as molecular designs that meet criteria across all properties that are improved, than conventional techniques such as standard Bayesian optimization. Thus, the selection engine increases or maximizes the likelihood that better molecular designs will be selected for in vitro and / or in vivo evaluation. In particular, the selection of molecular designs from the current design iteration may leverage experimental knowledge from prior design iterations, such that candidates with incrementally better properties are selected over successive design iterations.

[0058] FIG. 1 shows a system diagram illustrating an example of a molecular design system 100 according to some exemplary embodiments. Referring to FIG. 1 , the molecular design system 110 may include a molecular design engine 110, a selection engine 120, one or more wet lab instruments 130, and a client device 140. As shown in FIG. 1 , the molecular design engine 110, the selection engine 120, the one or more lab instruments 130, and the client device 140 may be communicatively coupled via a network 150. The one or more lab instruments 130 may include any wet lab equipment and dry lab equipment capable of performing in vitro measurements and / or in vivo characterization. Examples of the one or more lab instruments 130 may include a sequencer, a mass spectrometer, a centrifuge, etc. The client device 140 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. Network 150 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.

[0059] Referring again to FIG. 1 , the molecular design engine 110 may apply the molecular design computational model 115 to generate multiple molecular designs, including, for example, a first molecular design 160a, a second molecular design 160b, etc. For example, in some cases, the molecular design computational model 115 may be a machine learning model trained to learn a data distribution corresponding to a reduced-dimensional representation of the composition and / or structure of various known molecules, such as a protein sequence. In some cases, the data distribution may be a topological space (e.g., a manifold) occupied by the known molecules that describes the relationships that exist between them. The molecular design computational model 115 may generate each of the first molecular design 160a and the second molecular design 160b by sampling the data distribution (e.g., the topological space). For example, in some cases, the first molecular design 160a and the second molecular design 160b may each be a corresponding protein sequence, a reduced-dimensional representation of which occupies the data distribution (e.g., the topological space).

[0060] In some exemplary embodiments, the molecular design engine 110 can generate a large number of molecular designs, but not all molecular designs generated by the molecular design engine 110 undergo in vitro and in vivo evaluation. Instead, the selection engine 120 may perform one or more active learning iterations to identify one or more co-positive molecular designs that meet specific criteria for multiple properties for synthesis and testing by one or more laboratory instruments 130. In some cases, the one or more criteria associated with the properties may include property values ​​that meet one or more thresholds, fall within one or more value intervals, or are members of a set. For example, in the case of antibody design, the selection engine 120 may perform one or more active learning iterations to identify one or more molecular designs that exhibit sufficient expression levels and appropriate binding affinity to the target antigen. In some cases, the selection engine 120 may further perform one or more iterations of active learning to identify one or more molecular designs that, in addition to having sufficient expression levels and appropriate binding affinity, further exhibit certain developmental properties, such as specificity, thermostability, etc.

[0061] 2A shows a flowchart illustrating an example of a process 200 for molecular design involving multi-objective optimization of partially ordered and mixed variable properties, according to some exemplary embodiments. Referring to FIGS. 1-2, process 200 may be performed by molecular design engine 110 and selection engine 120 to, for example, identify a subset of molecular designs generated by molecular design engine 110 as candidates for in vitro and / or in vivo evaluation.

[0062] At 202, the molecular design engine 110 may generate a plurality of molecular designs. In some exemplary embodiments, the molecular design engine 110 may apply the molecular design computational model 115 to generate a plurality of molecular designs, including, for example, a first molecular design 160a, a second molecular design 160b, etc.

[0063] At 204, selection engine 120 may determine, for each molecular design of the plurality of molecular designs, a utility metric corresponding to the magnitude by which the combination of properties exhibited by each molecular design improves the combination of properties exhibited by one or more baseline molecular designs. In some exemplary embodiments, selection engine 120 may determine, for each of the plurality of molecular designs generated by molecular design engine 110, a corresponding utility metric indicating how much the combination of properties (e.g., expression, binding affinity, specificity, thermal stability, etc.) exhibited by each molecular design improves the same combination of properties exhibited by one or more baseline molecules. For example, in some cases, selection engine 120 may determine, for first molecular design 160a, a first utility metric indicating a first magnitude by which the combination of properties exhibited by first molecular design 160 improves the combination of properties exhibited by one or more baseline molecular designs. Additionally, the selection engine 120 may determine a second utility metric for the second molecular design 160b indicating a second magnitude by which the properties of the second molecular design 160b improve the properties of the baseline molecular design, which may in some cases include the first molecular design 160a.

[0064] At 206, the selection engine 120 may select one or more molecular designs as candidates for synthesis and testing based at least on a utility metric associated with each molecular design in the plurality of molecular designs. For example, in some cases, the selection engine 120 may identify the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing if their respective utility metrics meet one or more thresholds. Alternatively and / or additionally, the selection engine 120 may select the molecular design having the highest utility metric from the plurality of molecular designs generated by the molecular design engine 110 as a candidate for synthesis and testing. TIFF2025534388000003.tif4170 molecular designs may be selected, where the first molecular design 160a and / or the second molecular design 160b have the highest utility metric among the plurality of molecular designs generated by the molecular design engine 110. TIFF2025534388000004.tif4170 molecular designs, the first molecular design 160a and / or the second molecular design 160b may be selected as candidates for synthesis and testing. In some cases, in addition to the utility metric associated with each of the first molecular design 160a and the second molecular design 160b, the selection engine 120 may impose additional conditions when selecting the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing. For example, in the case of antibody designs, the selection engine 120 may further require the presence (or absence) of a particular amino acid residue (or sequence of amino acid residues) when selecting the first molecular design 160a and / or the second molecular design 160b as candidates for synthesis and testing.

[0065] 2B shows a flowchart illustrating another example of a process 250 for molecular design involving multi-objective optimization of partially ordered and mixed variable properties, according to some demonstrative embodiments. With reference to FIGS. 1 and 2A-2B, process 250 may be performed, for example, by selection engine 120 to determine a utility metric for each molecular design generated by molecular design engine 110. In some cases, process 250 may implement at least a portion of operation 204 of process 200 shown in FIG. 2A.

[0066] The selection engine 120 may receive the molecular design at 252. For example, in some cases, the selection engine 120 may receive from the molecular design engine 110 a first molecular design 160a generated by the molecular design computational model 115.

[0067] At 254, the selection engine 120 may apply one or more property computation models to determine a first probability of a molecular design exhibiting a first property and a second probability of a molecular design exhibiting a second property. For example, in some cases, a first property computation model may be applied to determine a first probability of a first molecular design 160a exhibiting a first property by at least enumerating the probability of occurrence of each possible value of a first property of the first molecular design 160a. In some cases, the first property computation model or a second property computation model may be applied to determine a second probability of a first molecular design 160a exhibiting a second property by at least enumerating the probability of occurrence of each possible value of a second property of the first molecular design 160a. In some cases, multiple property computation models (e.g., an ensemble of property computation models) may be applied to determine a value for each property of the first molecular design 160a.

[0068] At 256, the selection engine 120 may determine, based at least on the output of the one or more property computational models, a plurality of predicted samples associated with the molecular designs, each predicted sample including a first value for a first property exhibited by the molecular design and a second value for a second property exhibited by the molecular design having the first value for the first property. For example, in some cases, the output of the one or more property computational models may include a plurality of intermediate posterior samples, each including a first value for a first property exhibited by the first molecular design 160a and a second value for a second property exhibited by the first molecular design 160a. The intermediate posterior samples associated with the first molecular design 160a may correspond to a distribution of second values ​​for the second property exhibited by the first molecular design 160a across the first values ​​for the first property exhibited by the first molecular design 160a.

[0069] At 258, the selection engine 120 may identify, within the plurality of predicted samples, a set of predicted samples in which the first value of the first property satisfies a criterion. In some exemplary embodiments, the selection engine 120 may determine a utility metric indicating the extent to which the first property and the second property of the first molecular design 160a improve the first property and the second property of one or more baseline molecular designs. Furthermore, the selection engine 120 may impose a partial ordering when determining the utility metric associated with the first molecular design 160a, e.g., to prioritize the first property over the second property. To further illustrate, FIG. 4 illustrates a utility metric associated with the first property of the first molecular design 160a. TIFF2025534388000005.tif4170 is the second characteristic of the first molecular design 160a TIFF2025534388000006.tif4170 and third characteristics 4 shows a schematic diagram illustrating an example of a hierarchy 400 that takes precedence over TIFF2025534388000007.tif4170. Three features distributed across two levels indexed by TIFF2025534388000008.tif4170 TIFF2025534388000009.tif4170. The properties occupying the same level of hierarchy 400 are: Further indexed by TIFF2025534388000010.tif3170. First property Second property from TIFF2025534388000011.tif4170 TIFF2025534388000012.tif4170 and third characteristics The arrows to each of the TIFF2025534388000013.tif4170 may represent dependencies, such as experimental and / or biological dependencies, between them. Second property from TIFF2025534388000014.tif4170 TIFF2025534388000015.tif4170 and third characteristics The arrows to each of TIFF2025534388000016.tif4170 indicate the second characteristic TIFF2025534388000017.tif4170 and third characteristics The first characteristic for each of TIFF2025534388000018.tif4170 TIFF2025534388000019.tif4170. As explained in more detail below, the output of the probabilistic regression model 315 Arrows from TIFF2025534388000020.tif4170 and predicted samples output by the probabilistic binary classifier 313 Properties from TIFF2025534388000021.tif5170 The arrows point to TIFF2025534388000022.tif4170, respectively. TIFF2025534388000023.tif4170 is modeled as zero-inflated, where TIFF2025534388000024.tif5170 dominates zero events, TIFF2025534388000025.tif4170 governs consecutive non-zero events.

[0070] At 260, the selection engine 120 may determine, based at least on the set of prediction samples, a utility metric corresponding to an expected improvement in the first and second properties of the molecular design over the first and second properties of one or more baseline molecular designs. For example, in some cases, the selection engine 120 may determine a utility metric indicating the magnitude by which the first and second properties of the first molecular design 160a improve the first and second properties of one or more baseline molecular designs. Furthermore, the selection engine 120 may impose a partial ordering that favors the first property over the second property by subjecting intermediate posterior samples associated with at least the first molecular design 160a to resampling to generate a plurality of corresponding posterior samples. These posterior samples are then subjected to a multi-objective acquisition function to determine a utility metric for the first molecular design 160a. Examples of multi-objective acquisition functions applied to determine the utility metric of the first molecular design 160a may include expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc.

[0071] FIG. 3 illustrates a schematic diagram of an example molecular design pipeline 300, according to some demonstrative embodiments. As illustrated in the molecular design pipeline 300 illustrated in FIG. 3, the selection engine 120 may apply one or more property computation models 310 to, for example, determine a first probability of a first molecular design 160a exhibiting a first property and a second probability of a first molecular design 160b exhibiting a second property. In some cases, the one or more property computation models 310 may output a first probability distribution indicating the first probability of the first molecular design 160a exhibiting the first property by enumerating at least the probability of occurrence of each possible value of the first property exhibited by the first molecular design 160a. Furthermore, in some cases, the one or more property computation models 310 may output a second probability distribution indicating the second probability of the second molecular design 160a exhibiting the second property by enumerating at least the probability of occurrence of each possible value of the second property exhibited by the first molecular design 160a.

[0072] In some exemplary embodiments, the one or more property computation models 310 may be implemented as a zero-inflated probabilistic surrogate model including a probabilistic binary classifier 313 and a probabilistic regressor model 315. Further, in some cases, the one or more property computation models 310 may include separate property computation models trained to determine the probability that the first molecular design 160a exhibits each individual property. For example, the one or more property computation models 310 may include a first property computation model trained to determine a first probability of the first molecular design 160a exhibiting a first property and a second property computation model trained to determine a second probability of the first molecular design 160a exhibiting a second property. In some cases, the one or more property computation models 310 may include at least one property computation model trained to determine the probability that the first molecular design 160a exhibits multiple properties including, for example, a first property, a second property, etc. Further, in some cases, the one or more property computation models 310 may include multiple property computation models (e.g., an ensemble of property computation models) trained to determine the probability that the first molecular design 160a will exhibit the same property. For example, the one or more property computation models 310 may include a first property computation model and a third property computation model, each of which is trained to determine a first probability that the first molecular design 160a will exhibit a first property.

[0073] Including multiple property calculation models (or an ensemble of property calculation models) for a single property can compensate for at least some of the uncertainty that may exist in the output of the individual property calculation models. For example, in some cases, the output of an individual property calculation model may have low uncertainty (or high confidence that it is accurate) for some molecular designs encountered by the property calculation model and high uncertainty (or low confidence that it is accurate) for other molecular designs encountered by the property calculation model. Alternatively and / or additionally, the output of one individual property calculation model may have lower uncertainty (or high confidence that it is accurate) than the output of another individual property calculation model for the same molecular design. In this way, when multiple property calculation models (or an ensemble of property calculation models) are applied to determine the properties of a molecular design, low uncertainty in the output of some property calculation models can compensate for high uncertainty in the output of other property calculation models.

[0074] Referring again to FIG. 3 , the output of the one or more property computational models 310 may include a plurality of intermediate posterior samples 320. Each sample of the plurality of intermediate posterior samples 320 may include a first value of a first property exhibited by a first molecular design 160a and a second value of a second property exhibited by a first molecular design 160b. The plurality of intermediate posterior samples 320 may correspond to a distribution of second values ​​of a second property exhibited by a first molecular design 160a across the first value of the first property exhibited by the first molecular design 160a. For example, if the first property is an expression level and the second property is a binding affinity, each intermediate posterior sample may include a first value of an expression level exhibited by a molecular design and a second value of a binding affinity exhibited by a molecular design having the first value of the expression level. Thus, the plurality of intermediate posterior samples 320 may enumerate a distribution of different expression levels across the different binding affinities exhibited by the molecular designs.

[0075] In some exemplary embodiments, the selection engine 120 may impose a partial ordering that prioritizes a first property over a second property by subjecting at least the intermediate posterior samples 320 to resampling 330 to generate a plurality of posterior samples 340. A multi-objective acquisition function 350 may be applied to the posterior samples 340 to determine a utility metric for the first molecular design 160a. Examples of multi-objective acquisition functions 350 may include expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc. The resampling 330 may transform any intermediate posterior samples in which the first value of the first property does not satisfy one or more criteria that precludes these intermediate posterior samples from contributing to the utility metric for the first molecular design 160a.

[0076] To further illustrate, FIG. 5 shows the effect of resampling 330 on intermediate posterior samples 320 output by one or more property computational models 310. The dashed line in FIG. 5 represents a criterion (e.g., a threshold) for Objective 0 such that molecular designs selected as candidates for synthesis and testing increase or maximize Objective 1 while meeting this criterion for Objective 0. The dots in FIG. 5 constitute the individual baseline molecular designs that form the baseline Pareto front. Different shading within the grid indicates the hypervolume improvement (HVI) calculated from each posterior sample at a given location in objective space. Consider six intermediate posterior samples 320 (triangles) from the posterior (white outline) shown in the default state of graph 500 in FIG. 5(a). As shown in graph 550 in FIG. 5(b), resampling 330 can transform intermediate posterior samples 320 that do not meet the criterion for Objective 0 (e.g., below a threshold) so that their hypervolume improvement (HVI) contribution is zero. That is, the hypervolume improvement (HVI) values ​​of intermediate posterior samples 320 that do not meet the objective (0) (e.g., dotted vertical line), which are not zero by default as shown in graph 500 of FIG. 5(a), are set to zero as a result of resampling 330, as shown in graph 550 of FIG. 5(b).

[0077] In some exemplary embodiments, the utility metric output by the multi-objective acquisition function 350 for the first molecular design 160a may indicate how much the first and second properties of the first molecular design 160a improve over the first and second properties of one or more baseline molecular designs. Furthermore, as part of an active learning paradigm, the first molecular design 160a may become one of the baseline molecular designs for subsequent selection iterations. For example, the selection engine 120 may determine to select the second molecular design 160b generated by the molecular design engine 110 as another candidate for synthesis based at least on a second utility metric corresponding to an expected improvement in the first and second properties of the second molecular design 160b over the first and second properties of the baseline molecular design, which may be updated to include the first molecular design 160a in this selection iteration. The values ​​of the first and second properties associated with the first molecular design 160a may be determined empirically, for example, based on in vitro measurements and / or in vivo characterization. Alternatively and / or additionally, the values ​​of the first property and the second property of the first molecular design 160a may correspond to intermediate posterior samples 320 associated with the first molecular design 160a (e.g., an average of the values ​​of the first property and the second property included in the intermediate posterior samples 320).

[0078] As previously described, the selection engine 120 may determine a utility metric indicating the extent to which the first and second characteristics of the first molecular design 160a improve the first and second characteristics of one or more baseline molecular designs. Additionally, as shown in the molecular design pipeline 300 of FIG. 3, the selection engine 120 may impose a partial ordering that favors the first characteristic over the second characteristic by subjecting at least the intermediate posterior samples 320 to resampling 330 to generate a plurality of posterior samples 340, which are then subjected to a multi-objective acquisition function 350 to determine a utility metric for the first molecular design 160a. Examples of the multi-objective acquisition function 350 may include expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto-efficient global optimization (ParEGO), maximum entropy search method (MESMO), joint entropy search (JES), etc.

[0079] In some exemplary embodiments, the selection engine 120 may compensate for noise (e.g., measurement error associated with the laboratory equipment 130) that may be present in the observed properties of one or more baseline molecular designs. For example, in the example shown in FIG. 1 , where the selection engine 120 determines a utility metric indicating the extent to which a first property and a second property of a first molecular design 160 a improve the first property and the second property of one or more baseline molecular designs, the values ​​of the first property and the second property of the one or more baseline molecular designs may be observed in a wet laboratory and, therefore, may include at least some noise resulting from measurement error associated with the laboratory equipment 130. The effect of this noise may be reduced or minimized by determining the utility metric based on the output of one or more property computation models 310 that have been retrained based on the observed values ​​of the first property and the second property of the one or more baseline molecular designs. For example, the retrained property computation models 310 may be applied to determine the values ​​of the first property and the second property of the first molecular design 160 a and the values ​​of the first property and the second property of the one or more baseline molecular designs. The utility metric of the first molecular design 160a may be determined based on the output of the retrained property computational model 310 instead of the observed values ​​of the first and second properties of one or more baseline molecular designs.

[0080] To further illustrate, Figure 6A shows a flowchart illustrating an example of a process 600 for determining the probability of a molecular design exhibiting a property, according to some illustrative embodiments. With reference to Figures 1, 2A, and 6A, process 600 may be performed by selection engine 120 and may, for example, implement at least a portion of operation 204 of process 204 shown in Figure 2A.

[0081] At 602, the selection engine 120 may receive, for a baseline molecular design from a previous design iteration, observations about the properties of the baseline molecular design. In some exemplary embodiments, the molecular design from the previous design iteration may become the baseline molecular design for a subsequent design iteration. For example, one or more molecular designs from a previous design iteration may be selected for in vitro measurement and / or in vivo characterization of one or more desirable properties (e.g., expression, binding affinity for another molecule (e.g., a viral antigen, a tumor antigen, etc.), lack of non-specificity, stability, lack of immunogenicity, humanness, lack of self-association, etc.). As described in more detail below, the observations of the properties exhibited by these baseline molecular designs may be used during subsequent design iterations to identify one or more molecular designs whose combination of properties represents the greatest improvement over the properties of the baseline molecular design.

[0082] At 604, the selection engine 120 may retrain one or more property computational models 310 that were trained to determine probability distributions over different possible values ​​of the properties based at least on the observed values ​​of the properties of the baseline molecular design. The observed values ​​of the properties of the baseline molecular design received in operation 602 may contain at least some noise due, for example, to measurement error present in the laboratory equipment 130 used to generate the observed values. Thus, in some cases, the observed values ​​of the properties of the baseline molecular design are not directly used to identify molecular designs from subsequent design iterations whose combination of properties shows improvement or most improvement compared to the properties of the baseline molecular design. Instead, in some exemplary embodiments, the observed values ​​of the properties of the baseline molecular design are applied to retrain one or more property computational models 310. Doing so may generate posterior probability distributions for the corresponding properties, for example, by updating the prior probabilities of those properties based on the observed values.

[0083] In 606, the selection engine 120 may apply one or more retrained property computation models to determine a first value of a property of the baseline molecular design and a second value of a property of the molecular design from the subsequent design iteration. For example, in some cases, the retrained property computation model 310 may be applied to determine a value of a property of the baseline molecular design from a previous design iteration as well as a value of a property of a molecular design generated during the subsequent design iteration. In the case of expression levels, for example, the selection engine 120 may apply the retrained property computation model 310 to determine a first expression level of the baseline molecular design from a previous design iteration as well as a second expression level of the molecular design from the subsequent design iteration. Although the expression level of the baseline molecular design has been observed through wet laboratory experiments, the utility metrics of the molecular designs from the subsequent design iterations are not determined directly based on the observed expression level of the baseline molecular design. Instead, as described in more detail below, the utility metrics of the molecular designs from the subsequent design iterations may be determined based on the first expression level of the baseline molecular design determined by the retrained property computation model 310.

[0084] At 608, the selection engine 120 may determine, based at least on the first value of the property in the baseline molecular design and the second value of the property in the molecular design from the subsequent design iteration, a utility metric corresponding to the extent to which the combination of properties exhibited by the molecular design from the subsequent design iteration improves the combination of properties exhibited by the baseline molecular design. In some exemplary embodiments, the selection engine 120 may determine, for a molecular design from the subsequent design iteration, a utility metric corresponding to how much the properties of that molecular design improve the properties of the baseline molecular design from one or more previous design iterations. As mentioned above, even if observed values ​​for the properties of the baseline molecular design are available, the utility metric for the molecular design from the subsequent design iteration may instead be determined based on the output of the retrained property computation model 310 based on the observed values. Returning to the expression level example, the utility metric for the molecular design from the subsequent design iteration may be determined based on the first expression level of the baseline molecular design and the second expression level of the molecular design from the subsequent design iteration, where both the first expression level and the second expression level are determined by the retrained property computation model 310. In some cases, the utility metric may quantify the expected improvement (EI) for a combination of partially ordered and mixed-value properties exhibited by a molecular design from a subsequent design iteration relative to a baseline molecular design. Further, in some cases, the utility metric of a molecular design from a subsequent design iteration may be calculated by applying a utility function (e.g., multi-objective acquisition function 350) including, for example, expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), joint entropy search (JES), etc.

[0085] In some exemplary embodiments, the selection engine 120 may compensate for uncertainty that may exist in the output of each property calculation model 310. For example, the selection engine 120 may apply one or more property calculation models 310 to determine a first value of a first property and a second value of a second property exhibited by a molecular design. In some cases, the uncertainty in the output of each property calculation model 310 may be reduced or minimized by applying at least multiple property calculation models 310 (e.g., an ensemble of property calculation models) to determine each of the first value of the first property and the second value of the second property of the molecular design. For example, the selection engine 120 may apply a first property calculation model and a second property calculation model to evaluate the first property of the molecular design, and the first value of the first property may be determined based on the outputs of the first property calculation model and the second property calculation model. Similarly, the selection engine 120 may apply a third property calculation model and a fourth property calculation model to evaluate the second property of the molecular design, and the second value of the second property may be determined based on the outputs of the third property calculation model and the fourth property calculation model.

[0086] To further illustrate, Figure 6B shows a flowchart illustrating another example of a process 650 for determining the probability of a molecular design exhibiting a property, according to some illustrative embodiments. With reference to Figures 1, 2A, and 6A-6B, process 650 may be performed by selection engine 120 and may, for example, implement at least a portion of operation 254 of process 250 shown in Figure 2B or operation 606 of process 600 shown in Figure 6A.

[0087] At 652, the selection engine 120 may apply a first property computation model to determine a first value of a property exhibited by the molecular designs, and may apply a second property computation model to determine a second value of the property exhibited by the molecular designs. For example, in the case of expression levels, the selection engine 120 may apply an ensemble of property computation models trained to determine expression levels (e.g., a first property computation model, a second property computation model, etc.) to determine the expression level of each molecular design. Alternatively, in the case of binding affinity, the selection engine 120 may also apply an ensemble of property computation models trained to determine binding affinity (e.g., a third property computation model, a fourth property computation model, etc.) to determine the binding affinity of each molecular design. In some cases, the first property value and the second property value exhibited by the molecular designs may be expressed as probability distributions. For example, the output of a first property computational model may include a first probability distribution over the range of possible values ​​of the property (e.g., expression level, binding affinity, etc.), and the output of a second property computational model may include a second probability distribution over the range of possible values ​​of the property (e.g., expression level, binding affinity, etc.). In some cases, the outputs of the first property computational model and the second property computational model may be zero-inflated, meaning that the outputs include a first value (e.g., binary) indicating the presence (or absence) of the property (e.g., expression level, binding affinity, etc.) and a second value indicating the magnitude of the property exhibited by the molecular design.

[0088] At 654, the selection engine 120 may determine a third value of the property exhibited by the molecular design to calculate a utility metric for the molecular design based at least on the first value and the second value. The output of each probability computation model in the ensemble may be associated with at least some uncertainty. In particular, differences in architecture and / or training may result in a higher (or lower) property computation model when applied to different molecular designs. For example, a first property computation model may generate a more reliable output for the first molecular design than a second property computation model, while the second property computation model may generate a more reliable output for the second molecular design than the first property computation model. Thus, in some exemplary embodiments, the selection engine 120 may compensate for this uncertainty by determining at least the utility metric for the molecular design based on the output of multiple property computation models instead of the output of a single property computation model. In some cases, the value of a property (e.g., expression level, binding affinity, etc.) used to determine the utility metric for a molecular design may be determined based on multiple values ​​of the same property as determined by the ensemble of property computation models. For example, if a first property computation model is applied to determine a first value of the property and a second property computation model is applied to determine a second value of the property, the selection engine 120 may determine a third value of the property for determining a utility metric of the corresponding molecular design based at least on the first value and the second value. In some cases, the third value of the property may correspond to the mean, median, maximum, minimum, mode, and / or range of the first and second values.

[0089] As previously mentioned, in some exemplary embodiments, the selection engine 120 may perform multi-objective Bayesian optimization (BO) to trade off exploration (evaluating molecular designs that are highly uncertain) and search (evaluating molecular designs that are believed to increase or maximize an objective) by utilizing one or more probabilistic surrogate models (e.g., one or more property calculation models 310) and a utility function (e.g., multi-objective acquisition function 350). TIFF2025534388000026.tif5170 is a design space that is expensive to evaluate (e.g., in a wet laboratory) is a black box function in TIFF2025534388000027.tif4170, the objective of Bayesian optimization is to Design to increase or maximize TIFF2025534388000028.tif5170 TIFF2025534388000029.tif4170. Thus, Bayesian optimization (BO) in this context leverages one or more probabilistic surrogate models (e.g., one or more property calculation models 310) and a utility function (e.g., a multi-objective acquisition function 350) to efficiently identify molecular designs (e.g., TIFF2025534388000030.tif5170) to evaluate the design space (molecular designs with unknown potential to increase or maximize Search for TIFF2025534388000031.tif4170 and This may involve trading off the use of more specific molecular designs that are believed to increase or maximize the effectiveness of the TIFF2025534388000032.tif5170.

[0090] a probabilistic surrogate model (e.g., property computation model 310); TIFF2025534388000033.tif5170 is based on existing information We can build confidence about the probability distribution of TIFF2025534388000034.tif5170. For example, TIFF2025534388000035.tif5170 is the expression level of molecular design and probabilistic surrogate models (e.g., characteristic calculation model 310) The feature calculation model 310 may be trained based on web laboratory measurements of expression levels exhibited by molecular designs from previous design iterations. Assuming there is observation noise (e.g., measurement error associated with laboratory equipment 130), the feature calculation model 310 may be trained based on the observed expression levels exhibited by molecular designs from previous design iterations. It can be trained on the noisy dataset available up to TIFF2025534388000037.tif3170. In other words, at each iteration TIFF2025534388000038.tif4170, respectively TIFF2025534388000039.tif5170 Noisy observations of the dataset TIFF2025534388000040.tif5170 TIFF2025534388000041.tif6170. Thus, a probabilistic surrogate model (e.g., characteristic calculation model 310) TIFF2025534388000042.tif5170 is for representation purposes Posterior distribution quantifying the validity of TIFF2025534388000043.tif5170 TIFF2025534388000044.tif6170. For the example expression levels, the posterior distribution TIFF2025534388000045.tif6170 quantifies the probability distribution of possible expression levels exhibited by the next batch of molecular designs (e.g., generated by the molecular design computational model 115).

[0091] a utility function (e.g., a multi-objective acquisition function 350); TIFF2025534388000046.tif4170 is the posterior distribution determined by a probabilistic surrogate model (e.g., characteristic calculation model 310). Import TIFF2025534388000047.tif6170 and calculate the utility metrics for each molecular design. TIFF2025534388000048.tif3170. In some cases, a usefulness metric TIFF2025534388000049.tif3170 is the molecular design The utility of each molecular design may be quantified, as described in more detail below. The utility (or usefulness) of a molecular design may correspond to the likelihood of the molecular design exhibiting better properties than molecular designs from previous design iterations. For example, in some cases, the utility metric may be used to measure the likelihood of further in vitro measurements and / or in vivo characterization. Molecular designs that increase or maximize the posterior distribution of TIFF2025534388000052.tif3170 may be selected. That is, in some cases, molecular designs selected for in vitro measurement and / or in vivo characterization may be identified based on the difference between the expected value of a set of partially ordered and mixed variable properties (or objectives) exhibited by each molecular design (e.g., in molecular designs from previous design iterations) and the maximum value of the same properties (or objectives) observed to date. The expected value of each property (or objective) may be calculated based on the aforementioned posterior distribution inferred by the corresponding property computational model 310.

[0092] In some exemplary embodiments, a utility function (e.g., multi-objective acquisition function 350) may determine an expected improvement (EI) in properties exhibited by a molecular design in a current design iteration relative to properties exhibited by molecular designs from previous design iterations. Examples of TIFF2025534388000053.tif5170 may include Expected Hypervolume Improvement (EHVI), Noisy Expected Hypervolume Improvement (NEHVI), Pareto Efficient Global Optimization (ParEGO), Maximum Entropy Search Method (MESMO), Joint Entropy Search (JES), etc. For Expected Improvement (EI), the Expected Improvement (EI) acquisition function is: This can be obtained by taking TIFF2025534388000054.tif7170, in which case TIFF2025534388000055.tif7170. In some cases, the integral is a posteriori sample. It may be approximated by Monte Carlo (MC) integration using TIFF2025534388000056.tif5170. The maximizer of TIFF2025534388000057.tif3170 may be selected as a molecular design for in vitro measurement and / or in vivo characterization. As described in more detail below, the actual values ​​of the properties exhibited by this molecular design may be measured (e.g., in a web laboratory) before the observations are added to a dataset for retraining a corresponding probabilistic surrogate model (e.g., corresponding property computational model 310).

[0093] As mentioned above, the utility metric of the molecular design in the current design iteration TIFF2025534388000058.tif3170 may be calculated for property values ​​of molecular designs from previous design iterations. In some cases, property values ​​of molecular designs from previous design iterations may be determined based on wet laboratory measurements. Alternatively, property values ​​of molecular designs from previous design iterations may be determined by a corresponding probabilistic surrogate model (e.g., corresponding property calculation model 310) after the model has been retrained based on wet laboratory measurements of those properties. For example, in some cases, a probabilistic surrogate model in a molecular design TIFF2025534388000059.tif5170 (e.g., characteristic calculation model 310) is queried, for example, from wet laboratory measurements to obtain labeled pairs of the design. TIFF2025534388000060.tif5170 is obtained, and these values ​​are used in the next design iteration. TIFF2025534388000061.tif4170 dataset The probabilistic surrogate model (e.g., characteristic calculation model 310) may be added to the next design iteration. This expanded dataset before being applied to determine the posterior distribution of corresponding property values ​​in the molecular design from TIFF2025534388000063.tif4170 It can be retrained on TIFF2025534388000064.tif5170.

[0094] If there is a single objective (property) of interest, the best molecular design can be identified based on a ranking of the property values ​​(e.g., highest expression level, highest binding affinity, etc.). If there are multiple objectives (or properties) of interest, the best molecular design can be identified based on a ranking of the property values ​​(e.g., highest expression level, highest binding affinity, etc.). Since there may not be a single superior molecular design, the best molecular design may not have the best value for all objectives (or properties). TIFF2025534388000066.tif4 In a scenario with 170 objectives (or characteristics), at least TIFF2025534388000067.tif4170 probabilistic surrogate models (e.g., characteristic calculation 310) TIFF2025534388000068.tif5170 TIFF2025534388000069.tif4170. In some cases, instead of a single molecular design with the best values ​​for all properties, the goal of multi-objective optimization (MOO) may be to identify a set of Pareto-optimal tradeoffs, where improving one objective (or property) in the set worsens another objective (or property). For example, a Pareto-optimal tradeoff for optimizing across expression level and binding affinity may be a set of molecular designs where improved expression level is accompanied by a decrease in binding affinity.

[0095] To further illustrate, two molecular designs TIFF2025534388000070.tif5170 and Considering TIFF2025534388000071.tif5170, ( TIFF2025534388000072.tif4170 objectives) each design in objective space Solution for TIFF2025534388000073.tif3170 TIFF2025534388000074.tif5170. TIFF2025534388000075.tif6170 is all TIFF2025534388000076.tif 4170 purposes TIFF2025534388000077.tif6170 or In the case of TIFF2025534388000078.tif6170 Along with dominating TIFF2025534388000079.tif6170 TIFF2025534388000080.tif for at least one of the 4170 objectives It may be said to dominate TIFF2025534388000081.tif6170. TIFF2025534388000082.tif3170The Pareto Frontier (PF) can be defined as the set of non-dominated solutions, expressed as equation (1) below. TIFF2025534388000083.tif5170

[0096] The above formulation of the Pareto Frontier (PF) is a Pareto-optimal design This will result in a set of TIFF2025534388000084.tif5170. The size of TIFF2025534388000085.tif4170 is the number of iterations required for multi-objective optimization (MOO) to find a Pareto-optimal design. This is why we determine a finite approximation of the set of TIFF2025534388000086.tif4170. One way to measure the quality of the approximate Pareto Frontier (PF) is by looking at molecular designs from previous design iterations. TIFF2025534388000087.tif4170 desired values, etc. TIFF2025534388000088.tif4170 is governed by the specified reference point Hypervolume of a polytope bounded from below by TIFF2025534388000089.tif5170 TIFF2025534388000090.tif5170 is calculated. Represented as TIFF2025534388000091.tif5170, the design iteration Existing baseline Pareto front at time TIFF2025534388000092.tif3170 Expected improvements, such as expected hypervolume improvement (EHVI), for TIFF2025534388000093.tif5170 are: TIFF2025534388000094.tif5170. As previously mentioned, the expected hypervolume improvement (EHVI) acquisition function is one example of a multi-objective acquisition function 350 that may be applied by the selection engine 120. Other examples of the multi-objective acquisition function 350 may include noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), joint entropy search (JES), etc., which are described in more detail below. According to equation (2) below, the expected hypervolume improvement (EHVI) acquisition function is calculated by multiplying the posterior distribution output by the probabilistic surrogate model (e.g., the characteristic calculation model 310) by Hypervolume improvement for TIFF2025534388000095.tif5170 The integral of equation (2) is the posterior probability of the proxy. It can be evaluated by subtracting a sample (e.g., a Monte Carlo sample) from TIFF2025534388000097.tif5170. TIFF2025534388000098.tif5170

[0097] In a noise-free setting, the observed baseline Pareto front (PF) may be the true baseline Pareto front (e.g., TIFF2025534388000099.tif5170 where, However, this is not the case in many practical applications where wet laboratory measurements carry noise, for example in the form of measurement errors associated with the laboratory equipment 130. For example, the noise covariance Molecular design given a zero-mean Gaussian measurement process with TIFF2025534388000101.tif4170 Feedback for TIFF2025534388000102.tif3170 TIFF2025534388000103.tif5170, TIFF2025534388000104.tif5170 itself. To account for this noise, the baseline Pareto front (PF) associated with molecular designs from previous design iterations may correspond to property values ​​determined by an updated probabilistic surrogate model (e.g., property calculation model 310) retrained on wet laboratory measurements of those molecular designs. That is, the noisy expected excess volumetric improvement (NEHVI) is calculated from the previously observed points according to equation (3) below: It should be understood that noisy expected improvement (EI) extends expected improvement to settings with observation noise (e.g., measurement error associated with laboratory equipment 130), while noisy expected hypervolume improvement (NEHVI) extends noisy expected improvement to optimization of multiple objectives (or properties). TIFF2025534388000106.tif5170

[0098] In sequential optimization, or single molecule design per iteration Querying TIFF2025534388000107.tif5170 may be impractical for many applications due to latency in feedback: for example, in protein engineering, it may be necessary to select a batch of molecular designs at a given iteration and wait several months to receive measurements. TIFF2025534388000108.tifFrom a large pool of 5170 candidates The collaborative selection of a batch of 4170 molecular designs may require combinatorial evaluation of a utility function (e.g., multi-objective acquisition function 350). In the context of optimizing molecular designs based on the gradient of a utility function (e.g., multi-objective acquisition function 350), the gradient of the utility function (e.g., multi-objective acquisition function 350) at a given design iteration may be used. TIFF2025534388000110.tif4 Successive greedy selection of 170 molecular designs in various utility functions (e.g., multi-objective acquisition function 350) TIFF2025534388000111.tif4 can achieve performance equivalent to joint selection of 170 candidates.

[0099] Many molecular design applications require the implementation of some hierarchy among multiple desired properties of interest. In some cases, partial ordering can arise from experimental dependencies, where a molecular design must meet one or more criteria (e.g., pass a certain threshold) in one property before other properties can be measured. In the context of antibody design, a design candidate is a sequence of amino acids representing an antibody that must first be expressed in cell culture. If the expression level does not exceed a certain threshold in mass per volume, the laboratory cannot produce it in viable quantities and assay for other properties, such as binding affinity to the target antigen. Partial ordering that captures this experimental dependency can take the following form: Expression → Affinity. Such experimental dependencies create an asymmetry between the objectives. Non-expressing designs cannot provide binding measurements, thus reducing the information content of non-expressing molecular designs much more than non-binding designs.

[0100] Alternatively, the partial ordering can encode preferences for molecular design types. One or more properties can be prioritized, for example, so that molecular designs that perform poorly in these properties can be rejected regardless of how well they perform in other properties. In the context of antibody design, if a molecular design does not bind to the target antigen, it has failed in its primary function, and there is little interest in examining its developability properties, such as specificity for the target antigen and thermal stability, although unlike implicit properties, these developability properties often remain measurable. A partial ordering to capture this preference could take the form: expression → affinity → {specificity, thermal stability}.

[0101] Referring again to the example hierarchy 400 shown in FIG. 4, a partial ordering of properties, where some properties are prioritized over others, can be expressed as an ordered set of properties. TIFF2025534388000112.tif9170 where, TIFF2025534388000113.tif4170 is at level 400 TIFF2025534388000114.tif5170 shows the characteristics, TIFF2025534388000115.tif7170 is at the same level TIFF2025534388000116.tif4170 TIFF2025534388000117.tif7170 and its index among sibling properties. Various examples of multi-objective optimization (MOO) described herein, including the Bayesian optimization discussed above, are used at the molecular design level. TIFF2025534388000118.tif4170 If a characteristic does not meet the corresponding criteria, the level TIFF2025534388000119.tif4170 levels so that molecular designs that meet their corresponding criteria are rejected. TIFF2025534388000120.tif4170 characteristics of subsequent levels This may include taking precedence over the properties of TIFF2025534388000121.tif4170.

[0102] Biological properties tend to have an excess of zero or null values. That is, zero may be the most common value for biological properties such as expression, binding affinity, lack of nonspecificity, stability, lack of immunogenicity, humanness, lack of self-association, etc. For example, in the case of antibody design, a large proportion of molecular designs may not express at all, which contributes to a high incidence of zero values ​​(or null values) for that particular property. The zero-inflated nature of biological properties encourages the use of statistical models that account for the large incidence of zeros. For example, in some exemplary embodiments, one or more property computational models 310 may be implemented as zero-inflated probabilistic surrogate models. Thus, for each objective (or property), For TIFF2025534388000122.tif4170, the probabilistic binary classifier 313 of the feature calculation model 310 calculates the feature A binary random variable to indicate the presence (or absence) of TIFF2025534388000123.tif4170 TIFF2025534388000124.tif5170 and therefore the desired TIFF2025534388000125.tif4170 may generate zero values, while the probabilistic regression model 115 may generate zero values ​​with the same characteristics Corresponding to the remaining variance of consecutive non-zero values ​​for TIFF2025534388000126.tif4170 TIFF2025534388000127.tif5170. In some cases, the probabilistic binary classifier 313 may generate A binary random variable based on whether TIFF2025534388000128.tif4170 satisfies one or more thresholds TIFF2025534388000129.tif5170 can be assigned. Instead of indicating whether TIFF2025534388000130.tif4170 is completely present (or absent), the output of the probabilistic binary classifier 313 is It should be appreciated that this may indicate whether TIFF2025534388000131.tif4170 is present in a sufficient amount or level.

[0103] In one example where the feature computation model 310 is trained to predict the expression level of a molecular design, the output of the feature computation model 310 may include a first value determined by the probabilistic binary classifier 313 to indicate whether the expression level of the molecular design meets one or more thresholds. Additionally, the output of the feature computation model 310 may include a second value determined by the probabilistic regressor model 115 that indicates the expression level of the molecular design if the expression level of the molecular design is determined to meet one or more thresholds (e.g., assigned a value of 1 by the probabilistic binary classifier 313).

[0104] To explain further, Assume that TIFF2025534388000132.tif5170 is non-negative (equivalently, that it is bounded from below). Dataset available at TIFF2025534388000133.tif3170 Given TIFF2025534388000134.tif5170, the probabilistic binary classifier 313 finds the objective (or feature) Peripheral prediction posterior for TIFF2025534388000135.tif4170 TIFF2025534388000136.tif5170 can be modeled. TIFF2025534388000137.tif7170 where, TIFF2025534388000138.tif5170 corresponds to the cumulative distribution function (CDF) of the standard normal distribution. The first term of the integral shown in equation (4) may be a Bernoulli distribution that governs allelic uncertainty, and the second term may govern epistatic uncertainty.

[0105] On the other hand, the stochastic regression model 315 has the characteristic We can model the marginal predicted posterior of the non-zero modes (or non-zero values) of TIFF2025534388000139.tif4170. TIFF2025534388000140.tif6170

[0106] The stochastic regressor model 315 is It may be trained separately on the non-zero examples in TIFF2025534388000141.tif4170. Gaussian processes (GPs) may be used as a surrogate for Bayesian optimization (BO), and because the common Gaussian process assumption fails for sparse multimodal data, isolating the non-zero modes of the data can improve posterior inference.

[0107] Considering the above, each purpose (or characteristic) The marginal predicted posterior for TIFF2025534388000142.tif4170 may be a weighted mixture of the delta function and equation (5) above, with the relative weight of the latter given by the Bernoulli parameter in equation (4). This relationship is TIFF2025534388000143.tif5170, it can be expressed as the following equation (6). Zero-inflation modeling is a single-objective While described for TIFF2025534388000144.tif4170, zero-inflation modeling can be used to model multiple objectives (or properties), e.g., using a multitasking Gaussian process (GP). It should be understood that the TIFF2025534388000145.tif5170 file may be expanded post-hoc. TIFF2025534388000146.tif12170

[0108] The framework presented above for zero-inflated continuous-valued objectives (mixtures of delta functions at zero and continuous distributions) is very small, respectively. TIFF2025534388000147.tif3170 and accompanied by TIFF2025534388000148.tif7170 It applies to binary and continuous valued objectives without zero expansion, which can be considered as a particular case of taking TIFF2025534388000149.tif5170.

[0109] In some example embodiments, via resampling 330 shown in FIG. 3, the selection engine 120 may modify multiple intermediate posterior samples 320 output from one or more trait computation models 310 (e.g., zero-inflated probabilistic surrogate models) to further enforce hierarchical (e.g., parent-child) relationships between various traits. TIFF2025534388000150.tif4170 and its predecessor or parent node Consider TIFF2025534388000151.tif5170. Per TIFF2025534388000152.tif5170 TIFF2025534388000153.tif5170 and Consider one intermediate posterior sample from multiple intermediate posterior samples 320 output by one or more property calculation models for single molecule design with TIFF2025534388000154.tif5170. Without any modifications, Equation (7) is given by Sample of TIFF2025534388000155.tif4170 Resulting in TIFF2025534388000156.tif4170. TIFF2025534388000157.tif12170

[0110] However, the sample TIFF2025534388000158.tif4170 does not depend on any hierarchy, such as hierarchy 400 shown in FIG. 4, in which some properties take precedence over others. In some cases, dependencies that exist in hierarchy 400 may be passed to the parent node The output of the probabilistic binary classifier from TIFF2025534388000159.tif5170 is compared with the probabilistic binary classifier and the characteristics TIFF2025534388000160.tif4170. The output of the probabilistic binary classifier from TIFF2025534388000161.tif5170 is "0", therefore the corresponding molecular design is not a parent node Intermediate posterior sample if it shows that certain characteristics cannot be accounted for TIFF2025534388000162.tif5170 TIFF2025534388000163.tif4170 are excluded from contributing to the utility metrics (e.g., expected excess improvement (EHVI)) of the corresponding molecular design.

[0111] To further explain, the selection engine 120 may instead start, for example, at the top level of the hierarchy 400 and work down levels therein to select the characteristics. TIFF2025534388000164.tif5170 and its predecessor or parent property A dependency may be imposed between TIFF2025534388000165.tif5170. If TIFF2025534388000166.tif4170 is the top level characteristic, TIFF2025534388000167.tif5170 and TIFF2025534388000168.tif5170. If not, TIFF2025534388000169.tif4170 has parent properties, TIFF2025534388000170.tif5170 is defined as follows: TIFF2025534388000171.tif12170

[0112] Then, as follows: Valid sample of TIFF2025534388000172.tif4170 Binary sample modified to get TIFF2025534388000173.tif5170 TIFF2025534388000174.tif7170 can be used. TIFF2025534388000175.tif12170

[0113] TIFF2025534388000176.tif5170 and TIFF2025534388000177.tif5170, TIFF2025534388000178.tif5170, the sample-level transformation described in equations (8) and (9) is performed. Let's show it as TIFF2025534388000179.tif4170.

[0114] The selection engine 120 may repeat the resampling 330 for other intermediate posterior samples 320 associated with the molecular design. TIFF2025534388000180.tif5170 and each corresponding corrected sample vector TIFF2025534388000181.tif5170 may be used to evaluate a multi-objective acquisition function 350 (e.g., expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI) (Equation (4)), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), joint entropy search (JES), etc.), for example, via Monte Carlo (MC) integration. More precisely, TIFF2025534388000182.tif4170 intermediate posterior samples are design candidates TIFF2025534388000183.tif4170 (reflecting alexigenic and epistatic uncertainties) and previously observed designs Assume it was drawn parallel to TIFF2025534388000184.tif7170 (reflecting dyslexic uncertainty), Regarding TIFF2025534388000185.tif4170, each figure TIFF2025534388000186.tif5170 and TIFF2025534388000187.tif9170. A Monte Carlo approximation of a multi-objective acquisition function (e.g., possibly noisy expected excess improvement (NEHVI)) can then be efficiently evaluated as follows. TIFF2025534388000188.tif27170

[0115] Example Tasks for Multi-Objective Optimization (MOO)

[0116] The performance of the selection engine 120 was evaluated through simulated active experiments on two synthesis tasks and one real-world antibody design task. In these exemplary use cases, the noisy expected hypervolume improvement (NEHVI) (Equation (4)) was used as the multi-objective acquisition function 350. Furthermore, the multi-objective acquisition function 350 was evaluated through Monte Carlo integration (Equation (10)). Each experiment tested three types of acquisition: (1) batch multi-objective Bayesian optimization (BO) with partial ordering of features ("qNEHVI-DAG"), (2) batch multi-objective Bayesian optimization without ordering of features ("qNEHVI"), and (3) random. The primary performance metric, as previously described, is the number of "co-positive" molecular designs obtained, which refers to molecular designs that meet a specific criterion (e.g., exceed a selected threshold) in all objectives according to the specified partial ordering of features. Here, the size of each batch is Shown as TIFF2025534388000189.tif4170.

[0117] Exemplary Task I: Penicillin Production

[0118] The first penicillin production task is based on a penicillin production simulator. The two objectives in this example are to keep the fermentation time below a set threshold and to keep the yield above a set threshold. The latter two objectives can be defined as reducing or minimizing carbon dioxide (CO2) by-product emissions while ensuring that the target value exceeds the target value. This may be negated to assume TIFF2025534388000191.tif5170, where tiff2025534388000192.tif4170 yield("objective 0"), TIFF2025534388000193.tif4170 negative fermentation time ("Objective 1"), and TIFF2025534388000194.tif4170 Negative CO2 byproduct ("Objective 2"). Zero-mean Gaussian noise was added to the input.

[0119] the purpose An exact Gaussian process (GP) model is used for each image separately. TIFF2025534388000196.tif4170 may be fitted, and an approximate Gaussian process is modeled with the variational evidence lower bound (ELBO). TIFF2025534388000197.tif5170. In this exemplary task, 512 posterior samples were taken to assess qNEHVI.

[0120] Ten rounds of simulated active learning initialize one or more feature computation models 310 with eight training points, and in each iteration select a feature from a pool of 80 randomly sampled candidate points. This can be performed by selecting TIFF2025534388000198.tif5170. The three acquisition modes (qNEHVI-DAG, qNEHVI, and Random) were subjected to the same initial training points and candidate pool each time. The entire experiment was repeated five times. Figure 7(A) shows that qNEHVI-DAG identifies significantly more co-positives than qNEHVI and Random overactive learning iterations. Figure 8 compares qNEHVI and qNEHVI-DAG selections for each pair of objectives, with the final selection (after iteration 10) stacked across five iterations. For all objectives, qNEHVI-DAG identifies more examples to the right of the threshold (dashed black line) than qNEHVI and Random.

[0121] Type II Example: Branin-Currin

[0122] The next task is TIFF2025534388000199.tif4170 and Based on the analytical Branin-Currin test function from TIFF2025534388000200.tif4170. The Branin-Currin task is configured to simulate the antibody design task in a controlled environment, where the hierarchy of properties is: TIFF2025534388000201.tif4170 dimension 0 ("Objective 0") and TIFF2025534388000202.tif4170 Dimension 1 ("Objective 1") It may be defined as TIFF2025534388000203.tif5170. Objective 0 was converted to binary using a set threshold, and Objective 1 was zero-inflated and real-valued. Posterior inference was performed following a procedure similar to that for the penicillin production task.

[0123] Here, the selection engine 120 initializes one or more feature computation models 310 with six training points and selects from a pool of 40 randomly sampled candidate points in an iteration. Twenty rounds of simulated active learning were performed by selecting TIFF2025534388000204.tif5170. The entire experiment was repeated 10 times. Figure 7(B) shows that qNEHVI-DAG identifies significantly more co-positives than qNEHVI and Random overactive learning iterations. Figure 9 shows a comparison of qNEHVI-DAG, qNEHVI, and Objective 1 random selections for the final selections stacked across 10 iterations. Overall, qNEHVI-DAG identified more examples to the right of the threshold (dashed black line) than qNEHVI and Random, with the improvement being more pronounced for identifying co-positives (middle panel).

[0124] Exemplary Task III: Antibody Design

[0125] The antibody design task is derived from a real-world dataset of antibody sequences and their measured in vitro properties for affinity and expression. Similar to the toy problem, the hierarchy of properties is: TIFF2025534388000205.tif4170 formula ("Objective 0") and TIFF2025534388000206.tif4170 with affinity ("Objective 1") TIFF2025534388000207.tif5170. Objective 0 was binary (e.g., expressed or not), while objective 1 was zero-inflated and real-valued.

[0126] The selection engine 120 performed three iterations of simulated active learning, repeating the entire procedure five times. To simulate active learning, the entire dataset of 4,022 variable-length protein sequences designed as antibodies against the anonymized target antigen A was divided into five groups of sizes 1230, 736, 746, and 600. The first group served as the initial training set for one or more property calculation models 310, while the next three groups served as a "candidate pool" from which 200 molecular designs were selected during each iteration. The final remaining group served as a preliminary test set. As shown in Figure 10(A), qNEHVI-DAG outperformed qNEHVI and Random in terms of the number of co-positives. The log posterior density evaluated by affinity measurements for co-positives (expressing binders) in the test set shown in Figure 10(B) was also highest for qNEHVI-DAG, indicating that the feature calculation model 310 from qNEHVI-DAG had the most accurate confidence for co-positives after the final iteration.

[0127] In view of the foregoing embodiments of the subject matter, the present application discloses the following list of examples, which are further examples that may be included in the disclosure of this application by combining one feature of a single example or two or more features of said examples, optionally in combination with one or more features of one or more additional examples.

[0128] Item 1: A computer-implemented method comprising: applying to a first molecular design one or more property computational models trained to determine a first probability of the first molecular design exhibiting a first property and a second probability of the first molecular design exhibiting a second property; determining a first plurality of samples associated with the first molecular design based at least on output of the one or more property computational models, each sample of the first plurality of samples comprising a first value for a first property exhibited by the first molecular design and a second value for a second property exhibited by the first molecular design having a first value for the first property; identifying a first set of samples within the first plurality of samples whose first values ​​for the first property meet a first criterion; determining, based at least on the first set of samples, a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design over the first property and the second property of one or more baseline molecular designs; and identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design.

[0129] Term 2: The method described in Term 1, wherein the first utility metric is determined by applying expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), or joint entropy search (JES).

[0130] Item 3: The method of items 1 or 2, wherein the first property and the second property of the one or more baseline molecules are determined based on one or more in vitro measurements and / or in vivo characterizations associated with one or more baseline molecular designs.

[0131] Item 4: The method of any of items 1 to 3, further comprising retraining one or more property calculation models based at least on one or more in vitro measurements and / or in vivo property evaluations associated with one or more baseline molecular designs, and applying the one or more retrained property calculation models to determine first and second properties of the one or more baseline molecules.

[0132] Item 5: The method of any of items 1 to 4, further comprising identifying a second set of samples within the first plurality of samples in which the first value of the first characteristic does not satisfy the first criterion.

[0133] Term 6: The method of Term 5, wherein a first utility metric is determined to include a first contribution from a first set of samples and exclude a second contribution from a second set of samples.

[0134] Clause 7: The method of any of clauses 1 to 6, further comprising: applying to the first molecular design one or more property computation models trained to determine a third probability of the first molecular design exhibiting a third property; determining a first plurality of samples to further include as part of each sample a third value of the third property exhibited by the first molecular design based at least on the output of the one or more property computation models; identifying a first set of samples further based on a first value of the first property that satisfies a first criterion and a second value of the second property that satisfies a second criterion; and determining a first utility metric based at least on the first set of samples to further correspond to a first expected improvement in the first property, the second property, and the third property of the first molecular design over the first property, the second property, and the third property of one or more baseline molecules.

[0135] Item 8: The method of Item 7, wherein the first property and the second property occupy the same level in the hierarchy above the third property, such that the first molecular design must satisfy a first criterion associated with the first property and a second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

[0136] Item 9: The method of Item 7, wherein the first property and the second property occupy different levels in a hierarchy above the third property, such that the first molecular design must satisfy a first criterion associated with the first property before the first molecular design is evaluated with respect to the second property, and the first molecular design must further satisfy a second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

[0137] Paragraph 10: The method of any of paragraphs 7 to 9, wherein each of the first property, the second property, and the third property is a different one of expression, binding affinity, specificity, and thermostability.

[0138] Item 11: The method described in Item 1, wherein the one or more property computational models include a first property computational model trained to determine a first probability of a first molecular design exhibiting a first property.

[0139] Item 12: The method of Item 11, wherein the first characteristic computation model includes a first probabilistic binary classifier trained to determine a first probability of a first molecule exhibiting the first characteristic, and the first probabilistic binary classifier outputs a first value when the first probability meets a second threshold and outputs a second value when the first probability does not meet the second threshold.

[0140] Item 13: The method of items 11 or 12, wherein the first property computational model includes a first stochastic regressor trained to determine a first value of a first property exhibited by the first molecular design.

[0141] Paragraph 14: A method according to any of paragraphs 11 to 13, wherein the one or more property computation models further include a second property computation model trained to determine a second probability of the first molecule exhibiting a second property.

[0142] Item 15: The method of Item 14, wherein the second characteristic computation model includes a second binary classifier trained to determine a second probability of the first molecule exhibiting the second characteristic, and the second binary classifier outputs a first value when the second probability meets a second threshold and outputs a second value when the second probability does not meet the second threshold.

[0143] Item 16: The method of items 14 or 15, wherein the second property computational model includes a second regressor trained to determine a second value of the second property exhibited by the first molecular design.

[0144] Item 17: The method of any of items 1 to 16, wherein the one or more property calculation models include an ensemble of property calculation models, and the first probability of the first molecular design exhibiting the first property and / or the second probability of the first molecular design exhibiting the second property is determined based at least on an output of the ensemble of property calculation models.

[0145] Item 18: The method of any of items 1 to 17, further comprising: applying one or more property calculations to the second molecular design to determine a third probability of the second molecular design exhibiting the first property and a fourth probability of the second molecular design exhibiting the second property; determining a second plurality of samples associated with the second molecular design based at least on the output of the one or more property calculation models, each sample of the second plurality of samples comprising a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design; identifying a second set of samples within the second plurality of samples for which the third values ​​of the first property satisfy a first criterion; determining a second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design over the first property and the second property of one or more baseline molecular designs based at least on the second set of samples; and identifying the second molecular design as another candidate for synthesis based at least on the second utility metric of the second molecular design.

[0146] Item 19: The method of Item 18, wherein one or more baseline molecular designs are updated to include the first molecular design, such that the second expected improvement includes an expected improvement in the first property and the second property of the second molecular design over the first property and the second property of the first molecular design.

[0147] Item 20: The method of Item 19, wherein the one or more baseline molecular designs are updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or second property exhibited by the first molecular design.

[0148] Paragraph 21: The method of paragraph 19 or 20, wherein the one or more baseline molecular designs are updated to include an average of a first plurality of samples associated with the first molecular design.

[0149] Clause 22: The method of any of clauses 18 to 21, wherein the second usefulness metric is determined to include first contributions from the second set of samples and exclude second contributions from the third set of samples where a third value of the first characteristic does not satisfy the first criterion.

[0150] Paragraph 23: The method of any of paragraphs 18 to 22, wherein the first molecular design and the second molecular design are further identified as candidates in batch in vitro and / or in vivo evaluation.

[0151] Item 24: The method of any of items 1 to 23, wherein each of the first probability of the first molecular design exhibiting the first property and / or the second probability of the first molecular design exhibiting the second property comprises (i) a first probability distribution over a first value indicative of the corresponding property being present in the first molecular design and a second value indicative of the corresponding property not being present in the first molecular design, and (ii) a second probability distribution over a range of possible values ​​indicative of the magnitude of the corresponding property exhibited by the first molecular design.

[0152] Paragraph 25: The method of any of paragraphs 1 to 24, wherein the first molecular design is identified as a candidate for synthesis based at least on a first utility metric of the first molecular design satisfying one or more thresholds.

[0153] Item 26: has the highest usefulness metric 26. The method of any of items 1 to 25, further comprising selecting 4170 molecular designs as candidates for synthesis, wherein the first molecular design is identified as a candidate for synthesis based at least on the first molecular design being one of the N molecular designs having the highest utility metric.

[0154] Clause 27: The method of any of clauses 1 to 26, wherein the first value of the first characteristic satisfies the first criterion by satisfying a threshold, falling within one or more value intervals, or being a member of a set.

[0155] Item 28: The method of any of claims 1 to 27, further comprising identifying the first molecular design as a candidate for synthesis based at least on the presence or absence of one or more particular amino acid residues in the first molecular design.

[0156] Item 29: The method of any of items 1 to 28, wherein the first plurality of samples comprises a distribution of second values ​​of a second property exhibited by the first molecular design across a first value of a first property exhibited by the first molecular design.

[0157] Item 30: The method according to any one of Items 1 to 29, wherein the first molecular design is a protein molecule, a small molecule, an ion, a nucleic acid, a polysaccharide, and / or a glycolipid.

[0158] Item 31: The method of any of items 1 to 30, further comprising applying a molecular design computational model to generate a first molecular design.

[0159] Clause 32: A system comprising at least one data processor and at least one memory having stored thereon instructions that, when executed by the at least one data processor, result in operations including the method of any of clauses 1 to 31.

[0160] Clause 33: A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one data processor, result in operations including the method described in any of clauses 1 to 31.

[0161] 11 depicts a block diagram of an example computing system 1100 according to some exemplary embodiments. Referring to FIGS. 1-11, computing system 1100 may be used to implement molecular design engine 110, selection engine 120, laboratory equipment 130, client device 140, and / or any components therein.

[0162] 11 , computing system 1100 may include a processor 1110, a memory 1120, a storage device 1130, and an input / output device 1140. The processor 1110, the memory 1120, the storage device 1130, and the input / output device 1140 may be interconnected via a system bus 1150. The processor 1110 is capable of processing instructions for execution within the computing system 1100. Such executed instructions may implement one or more components, such as, for example, the molecular design engine 110, the selection engine 120, the laboratory equipment 130, the client device 140, etc. In some exemplary embodiments, the processor 1110 may be a single-threaded processor. Alternatively, the processor 1110 may be a multi-threaded processor. The processor 1110 is capable of processing instructions stored in the memory 1120 and / or the storage device 1130 to display graphical information for a user interface provided via the input / output device 1140.

[0163] Memory 1120 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 1100. Memory 1120 may store, for example, data structures representing a configuration object database. Storage device 1130 may provide persistent storage for computing system 1100. Storage device 1130 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Input / output device 1140 provides input / output operations for computing system 1100. In some exemplary embodiments, input / output device 1140 includes a keyboard and / or a pointing device. In various embodiments, input / output device 1140 includes a display device for displaying a graphical user interface.

[0164] According to some demonstrative embodiments, input / output devices 1140 may provide input / output operations for network devices. For example, input / output devices 1140 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0165] In some exemplary embodiments, computing system 1100 can be used to execute various interactive computer software applications that can be used for organizing, analyzing, and / or storing various types of data. Alternatively, computing system 1100 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications can include various add-in functions or can be standalone computing products and / or features. When active within an application, functionality can be used to generate a user interface that is provided via input / output devices 1140. The user interface can be generated by computing system 1100 and presented to a user (e.g., on a computer screen monitor, etc.).

[0166] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0167] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), and includes a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.

[0168] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, by which the user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software.

[0169] In the above description and in the claims, phrases such as "at least one of" or "one or more of" may appear followed by a list of conjunctive elements or features. The term "and / or" may also appear with a list of two or more elements or features. Unless implicitly or explicitly contradicted by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B;," "one or more of A and B;," and "A and / or B" are intended to mean "A alone, B alone, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, and C;," "one or more of A, B, and C;," and "A, B, and / or C" are intended to mean "A alone, B alone, C alone, A and B, A and C, B and C, or A, B, and C," respectively. Use of the term "based on" above and in the claims is intended to mean "based at least in part on," allowing for unrecited features or elements.

[0170] The subject matter described herein may be implemented in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, these implementations are merely some examples consistent with aspects related to the described subject matter. While several variations have been described in detail above, other modifications or additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features described above. Furthermore, the logic flow illustrated in the accompanying drawings and / or described herein does not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.

Claims

1. applying to a first molecular design one or more property computational models trained to determine a first probability of the first molecular design exhibiting a first property and a second probability of the first molecular design exhibiting a second property; determining a first plurality of samples associated with the first molecular design based at least on output of the one or more property computational models, each sample of the first plurality of samples including a first value for the first property exhibited by the first molecular design and a second value for the second property exhibited by the first molecular design having the first value for the first property; identifying a first set of samples within the first plurality of samples in which the first value of the first property meets a first criterion; determining, based at least on the first set of samples, a first utility metric corresponding to a first expected improvement in the first property and the second property of the first molecular design over the first property and the second property of one or more baseline molecular designs; identifying the first molecular design as a candidate for synthesis based at least on the first utility metric of the first molecular design; 10. A computer-implemented method comprising:

2. 2. The method of claim 1, wherein the first utility metric is determined by applying expected hypervolume improvement (EHVI), noisy expected hypervolume improvement (NEHVI), Pareto efficient global optimization (ParEGO), maximum entropy search (MESMO), or joint entropy search (JES).

3. retraining the one or more property computational models based at least on one or more in vitro measurements and / or in vivo characterizations associated with the one or more baseline molecular designs; applying the one or more retrained property computational models to determine the first property and the second property of the one or more baseline molecules; The method of claim 1 or 2, further comprising:

4. identifying a second set of samples within the first plurality of samples in which the first value of the first characteristic does not meet the first criterion; determining the first utility metric to include a first contribution from the first set of samples and to exclude a second contribution from the second set of samples; The method of claim 1 , further comprising:

5. applying the one or more property computational models to the first molecular design, the models being trained to determine a third probability of the first molecular design exhibiting a third property; determining the first plurality of samples to further include as part of each sample a third value of the third property exhibited by the first molecular design based at least on the output of the one or more property computational models; identifying the first set of samples further based on the first values ​​of the first characteristic satisfying a first criterion and the second values ​​of the second characteristic satisfying a second criterion; determining, based at least on the first set of samples, the first utility metric to further correspond to the first expected improvement in the first property, the second property, and the third property of the first molecular design over the first property, the second property, and the third property of the one or more baseline molecules; The method of claim 1 , further comprising:

6. 6. The method of claim 5, wherein the first property and the second property occupy the same level in the hierarchy above the third property, such that the first molecular design must satisfy the first criterion associated with the first property and the second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

7. 6. The method of claim 5, wherein the first property and the second property occupy different levels in a hierarchy above the third property, such that the first molecular design must satisfy the first criterion associated with the first property before the first molecular design is evaluated with respect to the second property, and the first molecular design must further satisfy the second criterion associated with the second property before the first molecular design is evaluated with respect to the third property.

8. 8. The method of claim 1, wherein the one or more property computational models include a first property computational model trained to determine the first probability of the first molecular design exhibiting the first property.

9. 9. The method of claim 8, wherein the first property computational model comprises a first probabilistic binary classifier trained to output a first value when the first probability meets a second threshold and to output a second value when the first probability does not meet the second threshold, and wherein the first property computational model further comprises a first probabilistic regressor trained to determine the first value of the first property exhibited by the first molecular design.

10. 10. The method of claim 8 or 9, wherein the one or more property computational models further comprise a second property computational model trained to determine the second probability of the first molecule exhibiting the second property.

11. 11. The method of claim 10, wherein the second property computational model comprises a second binary classifier trained to output a first value when the second probability meets a second threshold and to output a second value when the second probability does not meet the second threshold, and wherein the second property computational model further comprises a second regressor trained to determine the second value of the second property exhibited by the first molecular design.

12. 12. The method of claim 1, wherein the one or more property computational models comprise an ensemble of property computational models, and wherein the first probability of the first molecular design exhibiting the first property and / or the second probability of the first molecular design exhibiting the second property is determined based at least on an output of the ensemble of property computational models.

13. applying the one or more property calculations to the second molecular design to determine a third probability of the second molecular design exhibiting the first property and a fourth probability of the second molecular design exhibiting the second property; determining a second plurality of samples associated with the second molecular design based at least on the output of the one or more property computational models, each sample of the second plurality of samples including a third value of the first property exhibited by the second molecular design and a fourth value of the second property exhibited by the second molecular design; identifying a second set of samples within the second plurality of samples in which the third value of the first property satisfies the first criterion; determining, based at least on a second set of samples, a second utility metric corresponding to a second expected improvement in the first property and the second property of the second molecular design over the first property and the second property of the one or more baseline molecular designs; identifying the second molecular design as another candidate for synthesis based at least on the second utility metric of the second molecular design; 13. The method of claim 1, further comprising:

14. 14. The method of claim 13, wherein the one or more baseline molecular designs are updated to include the first molecular design, such that the second expected improvement comprises an expected improvement in the first property and the second property of the second molecular design over the first property and the second property of the first molecular design.

15. 15. The method of claim 14, wherein the one or more baseline molecular designs are updated to include one or more in vivo measurements and / or in vivo characterizations of the first property and / or the second property exhibited by the first molecular design.

16. 16. The method of claim 14 or 15, wherein the one or more baseline molecular designs are updated to include an average of the first plurality of samples associated with the first molecular design.

17. 17. The method of any one of claims 1 to 16, wherein the first probability of the first molecular design exhibiting the first property and / or the second probability of the first molecular design exhibiting the second property each comprise: (i) a first probability distribution spanning a first value indicative of a corresponding property being present in the first molecular design and a second value indicative of the corresponding property being absent in the first molecular design; and (ii) a second probability distribution spanning a range of possible values ​​indicative of a magnitude of the corresponding property exhibited by the first molecular design.

18. 18. The method of any one of claims 1 to 17, wherein the first molecular designs are identified as candidates in the synthesis based at least on the first utility metric of the first molecular designs satisfying one or more thresholds.

19. Has the highest usefulness metric selecting a first molecular design as a candidate for synthesis, wherein the first molecular design is identified as a candidate for synthesis based at least on the first molecular design being one of the N molecular designs having a highest utility metric; 19. The method of any one of claims 1 to 18, further comprising:

20. identifying the first molecular design as a candidate for synthesis based at least on the presence or absence of one or more particular amino acid residues in the first molecular design; 20. The method of any one of claims 1 to 19, further comprising:

21. 21. The method of any one of claims 1 to 20, wherein the first plurality of samples comprises a distribution of the second values ​​of the second property exhibited by the first molecular design across the first values ​​of the first property exhibited by the first molecular design.

22. at least one data processor; at least one memory having stored thereon instructions which, when executed by said at least one data processor, result in operations comprising the method of any one of claims 1 to 21; A system comprising:

23. 22. A non-transitory computer readable medium having stored thereon instructions which, when executed by at least one data processor, result in operations comprising the method of any one of claims 1 to 21.