Conformal Inference for Optimization

CI-OPT enhances biopolymer sequence optimization by integrating conformal inference with Bayesian optimization, addressing the limitations of Gaussian process priors and neural network uncertainty, achieving superior performance in high-dimensional discrete spaces.

JP7698654B2Active Publication Date: 2025-06-25FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022546359
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2021-01-29
Publication Date
2025-06-25
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Existing machine learning models for optimizing biopolymer sequences face challenges in providing accurate predictions with less training data, particularly in high-dimensional discrete spaces like biopolymer sequences, due to the limitations of Gaussian process priors and the intractability of full Bayesian treatment of uncertainty in neural networks.

Method used

The integration of conformal inference optimization (CI-OPT) with Bayesian optimization, using a neural network as a surrogate function and conformal scoring functions to calculate confidence intervals, allows for improved function estimation and uncertainty quantification, enabling better optimization of biopolymer sequences.

Benefits of technology

CI-OPT provides more accurate function estimation and well-calibrated uncertainty, outperforming traditional Gaussian process-based methods in both synthetic and real-world protein datasets, especially in high-dimensional discrete spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698654000016
    Figure 0007698654000016
  • Figure 0007698654000017
    Figure 0007698654000017
  • Figure 0007698654000018
    Figure 0007698654000018
Patent Text Reader

Abstract

A computer-implemented method for optimizing the design of biopolymer sequences is provided. Accurate function estimates and well-calibrated uncertainty are important for Bayesian optimization (BO). The most theoretical guarantee for BO is established for modeling the objective function using a surrogate derived from a Gaussian process (GP) prior. GP priors are poorly suited to discrete, high-dimensional combinatorial spaces such as biopolymer sequences. Using neural networks (NNs) as surrogate functions can yield more accurate function estimates. The use of NNs allows for arbitrarily complex models, eliminates GP prior assumptions, and allows for easy pre-training, which is beneficial in low-data BO regimes. However, fully Bayesian treatment of uncertainty in NNs remains intractable, and existing approximation methods such as Monte Carlo dropout and variational inference can significantly miscalibrate uncertainty estimates. Conformal inference optimization (CI-OPT) uses confidence intervals calculated using conformal inference as a surrogate for posterior uncertainty in a specific BO acquisition function. A conformal scoring function with properties suitable for optimization has been validated on standard BO datasets and real-world protein datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Patent Application No. 62 / 967,941, filed on January 30, 2020. The entire disclosure of the above application is incorporated herein by reference.

Background Art

[0002] Machine learning generally employs statistical models that can be utilized by computer-implemented methods to perform a given task. Often, the statistical models employed by machine learning methods detect patterns and use the patterns to predict future behavior. The statistical models and neural networks employed by machine learning methods are typically trained using real-world data, and the machine learning methods utilize the real-world data to predict future behavior.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Therefore, there is a need for an improved machine learning model that provides better predictions of data using less training data. Accurate function estimation and well-calibrated uncertainty are important for Bayesian optimization (BO). The most theoretical guarantee of BO is established for the method of modeling the objective function using a surrogate derived from a Gaussian process (GP) prior. The GP prior is not very suitable for discrete high-dimensional combinatorial spaces such as biopolymer sequences. By using a neural network (NN) as a surrogate function, more accurate function estimation can be obtained. Using an NN enables any complex model, eliminates the GP prior assumption, and allows for easy pre-training that is beneficial in the low-data BO regime. However, the full Bayesian treatment of uncertainty in an NN remains intractable, and recent results have shown that approximate inference can lead to estimates that do not approximate the true posterior well. Conformal inference optimization (CI-OPT) uses conformal inference as a substitute for the posterior uncertainty in a specific BO acquisition function and uses the computed confidence intervals. Current methods do not combine conformal inference with BO due to intractability, but the applicant discloses a conformal scoring function having properties suitable for optimization that is effective in synthetic optimization tasks, standard BO data sets, and real-world protein data sets.

Means for Solving the Problems

[0004] In one aspect, a computer-implemented method for optimizing the design of a biopolymer sequence can include training a machine learning model using an observed biopolymer sequence and a labeled biopolymer sequence corresponding to each observed biopolymer sequence. The labeled sequence is a sequence associated with a real number that measures some property of interest. The method can further include determining a biopolymer sequence candidate and observing the one having the highest predicted value of the labeled biopolymer sequence based on the machine learning model. The biopolymer sequence candidate can include either a known sequence (e.g., a sequence faced previously, a sequence observed previously, or a natural sequence) or a newly designed sequence. The method can further include, for each biopolymer sequence candidate, specifying a conformal inference interval that represents the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence. The method can further include selecting at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence.

[0005] In one aspect, the value of the labeled sequence is the number used as a label as described above. Thus, the predicted value of the sequence is the predicted label of the sequence. One of ordinary skill in the machine learning art would be able to understand such a definition of a label. The sequence or data point is the machine learning input (x), and the prediction / measurement / optimization is the label (y).

[0006] In an aspect, the conformal inference interval includes a central value and an interval range. The central value can be the average value.

[0007] In one aspect, the machine learning model is a neural network fine-tuned using observed biopolymer sequences and their labels. The fine-tuned neural network is a neural network pre-trained on a large dataset that uses those weights as initial weights for a smaller dataset. Fine-tuning can accelerate training and overcome small dataset sizes. In one aspect, identifying the conformal inference interval is based on a second set of observed biopolymer sequences. The second set of sequences is the set of sequences used to adjust the conformal score.

[0008] In one aspect, identifying the conformal inference interval can further include calculating a residual interval based on each output of the machine learning model for each of the second set of observed biopolymer sequences and the corresponding labeled biopolymer sequences corresponding to each of the second set of biopolymer sequences. Identifying the conformal inference interval can further include calculating the average distance to a plurality of nearest neighbor sequences of the observed biopolymer sequence in the metric space for each output of the machine learning model. Identifying the conformal inference interval can further include calculating a conformal score based on the ratio of the residual to the sum of the average distance and a constant. As described below, the metric space is a set of possible sequences. An example of a metric can be the Levenshtein distance. In an aspect, the constant can vary with each iteration.

[0009] In one aspect, selecting at least one biopolymer sequence candidate includes calculating the average distance in the metric space to a plurality of nearest neighbor sequences in the metric space, generating a confidence interval based on the at least one biopolymer sequence candidate and the average distance, and selecting the at least one biopolymer sequence candidate based on the confidence interval.

[0010] In an embodiment, the conformational interval can be at least 50% and at most 99%. The biopolymer sequence can include at least one of an amino acid sequence, a nucleic acid sequence, and a carbohydrate sequence. The nucleic acid sequence can be a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence. The amino acid sequence can be any sequence that includes all proteins such as, for example, enzymes, growth factors, cytokines, hormones, signaling proteins, conformational proteins, motility proteins, antibodies (including both immunoglobulin-based molecules and alternative molecular scaffolds), and combinations of the above including fusion proteins and conjugates.

[0011] In one embodiment, a computer-implemented method and corresponding system for optimizing the design of a polymer sequence can include training a model to approximate a labeled biopolymer sequence of an initial sample from a plurality of observed sequences. The method can further include selecting, from the plurality of observed sequences, at least one sequence that optimizes a combination of a labeled polymer sequence generated by the trained model and a conformational interval for a particular batch of the plurality of observed sequences having the conformational intervals of the labeled biopolymer sequence and each observed sequence generated by the trained model. The method can further include recalculating the conformational intervals of the remaining sequences.

[0012] In an embodiment, the method can further include repeating, for each of the plurality of batches, selecting at least one sequence and recalculating the conformational interval. In an embodiment, the method can further include identifying an optimal number of batch experiments to perform in parallel. In an embodiment, identifying can be based on the optimization of wet lab resources.

[0013] In one aspect, a computer-implemented method can include training a machine learning model using data points in a metric space and functional values corresponding to each observed data point. The functional value is a real number that measures some property of interest of the data point. The method can further include determining data point candidates and observing those having the highest predicted functional value based on the machine learning model. The data point candidates can include known data points (e.g., data points faced previously, data points observed previously, or natural data points) or newly designed data points. The method can further include, for each data point candidate, specifying a conformal inference interval that represents the likelihood that the data point candidate has the predicted functional value of the data point. The method can further include selecting at least one data point candidate having an optimal linear combination of the conformal inference interval and the predicted functional value of the data point. One skilled in the art can recognize that this can include images, videos, audio, other media, and other data that can be interpreted by a machine learning model.

[0014] In one aspect, a computer-implemented method and corresponding system can include training a model to approximate functional value data points of an initial sample from a plurality of observed data points. The method can further include, for a particular batch of the plurality of observed data points having functional values generated by the trained model and conformal intervals for each observed data point, selecting at least one arrangement that optimizes a combination of labeled data points generated by the trained model and conformal intervals from among the plurality of data points. The method can further include recomputing the conformal intervals of the remaining data points.

[0015] In one aspect, a computer-implemented method for optimizing a design based on the distribution of data includes training a machine learning model using a plurality of observation data and labeled data corresponding to each observation data. The method can further include determining a plurality of data candidates and observing those having the highest predicted value of the labeled data based on the machine learning model. The method can further include, for each data candidate, specifying a conformal inference interval representing the likelihood that the data candidate has the predicted value of the labeled data. The method can further include selecting at least one data candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled data.

[0016] In one aspect, the method further includes providing at least one selected biopolymer sequence to means for synthesizing the selected biopolymer sequence, and optionally, the at least one selected biopolymer sequence is synthesized.

[0017] In one aspect, the method further includes synthesizing at least one selected biopolymer sequence.

[0018] In one aspect, the method further includes assaying (qualitatively or quantitatively by chemical assay) at least one selected biopolymer sequence.

[0019] In one aspect, a non-transitory computer-readable medium is configured to store instructions for optimizing the design of a biopolymer sequence. When executed by a processor, the instructions cause the processor to use a plurality of observed biopolymer sequences and the labeled biopolymer sequences corresponding to each observed biopolymer sequence to train a machine learning model, determine a plurality of biopolymer sequence candidates, and observe those having the highest predicted value of the labeled biopolymer sequence based on the machine learning model. For each biopolymer sequence candidate, identify a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence, and select at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence.

[0020] In one aspect, a system for optimizing the design of a biopolymer sequence includes a processor and a memory storing computer code instructions. The processor and the memory are configured to use the computer code instructions to cause the system to train a machine learning model using a plurality of observed biopolymer sequences and the labeled biopolymer sequences corresponding to each observed biopolymer sequence, determine a plurality of biopolymer sequence candidates, and observe those having the highest predicted value of the labeled biopolymer sequence based on the machine learning model. For each biopolymer sequence candidate, identify a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence, and select at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence.

[0021] In one aspect, disclosed herein is one or more selected biopolymer sequences obtainable by the method according to any one of the preceding claims.

[0022] In some embodiments, one or more selected biopolymer sequences are produced by in vitro methods of chemical synthesis. In other embodiments, one or more selected biopolymer sequences are produced by biosynthesis using, for example, cell-based systems such as bacterial, fungal, or animal (e.g., insect or mammalian) systems. For example, in some embodiments, one or more selected biopolymer sequences are one or more selected polypeptide sequences. In certain more specific embodiments, one or more selected polypeptide sequences are produced by chemical synthesis, for example, in a peptide synthesizer. In other more specific embodiments, one or more selected biopolymer sequences are synthesized by a biological system, which comprises, for example, providing one or more nucleic acid sequences (in an expression vector) to a biological system (e.g., a host cell or an in vitro translation system such as a transcription and translation system), culturing the biological system under conditions that promote the synthesis of one or more selected polypeptide sequences, and isolating the synthesized one or more selected polypeptide sequences from the system.

[0023] In some embodiments, the composition comprises one or more selected polymer sequences, optionally comprising a pharmaceutically acceptable excipient.

[0024] In some embodiments, the method comprises contacting the composition or the selected biopolymer sequence recited in any one of the preceding claims with one or more of a test compound, a biological fluid, a cell, a tissue, an organ, or an organism.

[0025] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings are requested and are provided by the Patent Office upon payment of the necessary fee.

[0026] The foregoing will be apparent from the following more particular description of example embodiments, as illustrated in the accompanying drawings, wherein like reference characters refer to the same parts throughout the various views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.

Brief Description of the Drawings

[0027]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

Figure 1G

Figure 1H

Figure 1I

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0028] The description of the example embodiments is as follows.

[0029] Bayesian optimization (BO) is a popular technique for optimizing black-box functions. The uses of BO include, in particular, experimental design, hyperparameter tuning, and control systems. Conventional BO methods rely on the ex-post well-calibrated uncertainty induced by the observation of the objective function or the true function. The objective function is the property to be optimized. For example, if the system is optimizing a biopolymer, the objective function can optimize the properties of the biopolymer. By using uncertainty to guide the decision, BO becomes particularly powerful in low-data situations. Current embodiments such as those shown in "Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling", arXiv preprint arXiv:1802.09127, (2018) (hereinafter "Riquelme") show that both accurate function estimation and well-calibrated uncertainty are important for strong performance in real-world problems.

[0030] The most theoretical guarantees of BO are established for methods of modeling the objective function using a surrogate derived from a Gaussian process (GP) prior. If the function deviates significantly from the GP prior, the resulting posterior probability can be a poor estimate of the true function, can be mis-calibrated uncertainty, or both. This is important, especially when the design space is discrete and combinatorial (e.g., a biopolymer sequence such as a protein sequence), because most GP priors are designed for low-dimensional continuous spaces and may not be good surrogates for these types of spaces.

[0031] One way to obtain a more accurate function estimate is to use a neural network as a surrogate function. A surrogate function is a function that models the objective / true function. In addition to enabling any complex model and eliminating the GP prior assumption, the use of a neural network allows for pre-training, which can be particularly beneficial in the low-data BO regime. However, full Bayesian treatment of uncertainty in neural networks, such as using Hamiltonian Monte Carlo to estimate the posterior, remains computationally intractable, and recent results have shown that approximate inference can produce estimates that do not reflect the true posterior well. An alternative is to use Bayesian linear regression on top of a neural network as a surrogate function. Riquelme's method compares the performance of different approximate Bayesian uncertainty quantification methods in the BO task.

[0032] Conformal inference collectively refers to a family of uncertainty quantification methods. Conformal inference methods provide valid calibrated prediction intervals under the assumption that the data are exchangeable. One skilled in the art can recognize that exchangeable data satisfy the equation p(x1,x2,...x n )=p(x s1 ,x s2 ,...,x sn ) for any permutation of the indices. Unlike Bayesian methods such as GP models, conformal inference does not rely on strong fundamental assumptions about the data or the objective function. Conformal inference can also be applied on top of any machine learning model, thereby enabling the construction of valid prediction intervals on modern deep learning techniques such as large pre-trained models for which Bayesian inference is not appropriate.

[0033] In one aspect of the present disclosure, the method and corresponding system employ conformal confidence intervals with a Bayesian optimization method. The combination of conformal confidence intervals with the Bayesian optimization method is hereinafter referred to as Conformal Inference Optimization (CI-OPT). CI-OPT employs confidence intervals calculated using conformal inference as a drop-in replacement for the posterior uncertainty in a particular BO acquisition function.

[0034] At a high level, the problem to be solved can be described as initially aiming to find the maximum of a function f(x) over a certain set of decisions

Number

[0035] As further described in "Taking the Human out of the Loop: A Review of Bayesian Optimization" by Shahriari et al., Proceedings of the IEEE, 104(1): 148-175 (2015) (hereinafter "Shahriari"), current Bayesian optimization techniques and methods are initiated by placing a prior on f. At time step t+1, the previous observations Y t ={x1,...,x t} at perhaps noisy locations X t ={y1,...,y t} induce the posterior distribution of f. The acquisition function a(x,t) determines which point in X to query next via the proxy optimization x t+1 =argmax x a(x|D t ,t), where Dt ={X t , y t}}. The acquisition function uses the posterior over f to balance the use of information obtained from previous queries and the use of regions with high uncertainty.

[0036] The Gaussian process is a common choice for function priors (see, e.g., Williams et al., “Gaussian processes for machine learning,” Volume 2. MIT press Cambridge, MA, 2006 (hereinafter “Williams”)). A GP is an infinite collection of random variables such that every finite subset of the random variables has a multivariate Gaussian distribution. The GP model assumes that the unknown true function is derived from the GP prior, and then the GP model uses the observations to compute the posterior over the function. The main advantage of the GP model is that there is a simple closed-form solution for the posterior, which makes the GP model one of the most popular theoretical tools for Bayesian optimization. The GP posterior at each step is marginalized in closed form to arrive at the predictive mean μ t (x) and standard deviation σ t (x),

Number

[0037] Notable acquisition functions include the following. a) As shown by Jones et al. in “Efficient Global Optimization of Expensive Black-Box Functions”, Journal of Global optimization, 13(4):455 - 492, 1998 (hereinafter “Jones”), the expected improvement:

Number

Number

Number

[0038] More generally, the observations can be queried in batches instead of being strictly sequential. In the batch setting, at time t + 1, a set of B items x r ,...,x r+B-1 is the previous (possibly noisy) observation y t ={x1,...,x t} at location X t ={y1,...,y tIt is selected for the query based on {}. Generally, B can be adaptively selected in every iteration (see, for example, "Parallelizing exploration-exploitation tradeoffs in gaussian process bandit optimization" by Desautels et al., Journal of Machine Learning Research, 15:3873-3923, 2014 (hereinafter "Desautels")), but in this example, a setting where the batch size is fixed is explored. Many batch Bayesian optimization methods use the uncertainty in the acquisition function to generally generate diverse batches.

[0039] For example, the method of Desautels generalizes the GP-UCB acquisition function that batches queries by updating t after each selection in the batch so that the selection is queried and observed as if its mean posterior value. Alternatively, as shown in "Maximizing Acquisition Functions for Bayesian Optimization" by Wilson, Advances in Neural Information Processing Systems, pp. 9884-9895, 2018 (hereinafter "Wilson") and "Sampling Acquisition Functions for Batch Bayesian Optimization" by De Palma et al., arXiv preprint arXiv:1903.09434, 2019 (hereinafter "De Palma"), the acquisition function can be sampled from the GP posterior to generate diverse batches.

[0040] Conformal inference is an auxiliary method that provides a strict finite-sample 1-ε prediction interval for any underlying machine learning model, as shown in "Transduction with Confidence and Credibility" by Saunders et al., 1999 (hereinafter "Saunders"). Exchangeable samples

Number

Number

Number

[0041] As shown in “Regression Conformal Prediction With Nearest Neighbours” by Papadopoulos et al., Journal of Artificial Intelligence Research, 40:815 - 840, 2011 (hereinafter “Papadopoulos”), conformal regression aims to find non - uniform variance confidence intervals. In one example, a regressor tr trained on Z tr = {X tr , Y tr}

Number

Number

[0042] An exemplary conformal scoring function C with properties suitable for optimization is described here. Conformal prediction intervals (e.g., 95% conformal intervals) can be used as drop-in replacements for t in Bayesian optimization procedures. However, one of ordinary skill in the art can understand that it is also possible to use other conformal intervals. In one example, conformal intervals in the range of 50% to 99% (including fractions) can be used.

[0043] At step t, the regressor h is trained with {X t , y t}, the conformal scoring function g, the sensitivity parameter β, and the 95 percentile calibration score c S . Then the equation μ t,CI (x * ) = h(x *) (Formula 6) and

number

[0044] The choice of g is crucial for deriving an interval that balances exploration and exploitation. Ideally, the interval should be narrow enough to accommodate densely sampled regions. For example, a common approach is to choose g for x, y ∈ Z. c This g is basically a model trained to predict the residual |h(x)-y| for x, y ... c but

number

number

number

[0045] Intuitively, this can be related to two sources of uncertainty in the Gaussian process posterior: the residual |h(x)-y| resembles the homoscedastic noise variance, while g kNN is similar to the hard-thresholded stationary GP covariance function. In other words, Equation 9 includes a conformal score that explicitly estimates the heteroscedastic uncertainties.

[0046] Figures 1G to 1E compare the uncertainties calculated from conformal inference using a neural network residual estimator and conformal inference using scaled k-nearest neighbors (Equation 9) after GP. The shaded regions are ±2 standard deviations in GP (Figure 1G) and 95% in conformal inference (Figures 1H and 1I).

[0047] Figure 1G shows the uncertainty calculated from the GP posterior (squared exponential kernel, which is a hyperparameter estimated by maximizing the marginal likelihood). Figures 1H and 1I show the uncertainty calculated from conformal inference using the sigmoid non-linearity and training the calibration set on top of a three-layer fully-connected neural network. Figure 1H shows the conformal interval generated for g using a neural network residual estimator. Figure 1I shows the conformal interval generated for g using Equation 9 with k = 5 and β = 0.001 for both conformal plots. By using a neural network residual estimator for g, a wider prediction interval is obtained in the densely sampled region of X tr which is a problem that worsens by setting Z c =Z tr Using those intervals in the acquisition function results in an optimizer that is more likely to get stuck at a local optimum as the weight on uncertainty increases.

[0048] The nearest neighbors can be determined by the distance to x in the metric space in the training set. The metric space is a set of possible arrays or data points. An example of a metric can be the Levenshtein distance.

[0049] Selecting items to query using the conformal score violates the assumption of exchangeable data in the GP model. Further, in the small data regime such as at the start of an optimization run, the calibration score is Z trIt may be necessary to calculate with. Therefore, both will be prediction intervals that do not have a strict finite sample coverage guarantee. However, these intervals remain useful for the trade-off between exploration and exploitation during optimization.

[0050] In other words, the applicant's method includes: (1) calculating a conformal interval of a prediction from a neural network fine-tuned using nearest neighbors in the array space; and (2) performing batch optimization using the conformal interval calculated in (1).

[0051] FIG. 2 shows an example embodiment of calculating a conformal interval of a prediction from a neural network fine-tuned using nearest neighbors in the array space as in the present disclosure. To calculate the conformal interval, the method uses the following: a) f(x): A neural network fine-tuned. b) X t : The array used for fine-tuning f(x). c) X c : The array (202) used for adjusting the conformal score. d) y c : The true function value (204) corresponding to X c . e) n: The number of nearest neighbors to consider. f) b: A hyperparameter. g) alpha: The desired confidence value, and h) X test : The new array to be predicted.

[0052] Then, for each x, y in X c , y c , the method calculates the residual: r = |f(x) - y| (206). For each x in X c , the method calculates the average distance to the n nearest neighbors in X t and assigns it to d (208). For x in X c , calculate the conformal score:

Equation

[0053] Figure 3 shows an example embodiment of a method for optimizing a batch using the above conformal interval. The method uses the following: a) B: batch size, b) N: number of iterations, c) n init : initial number of samples, d) X: possible arrays, and e) C: constant.

[0054] The method then evaluates n init arrays from X to determine an output y (302) and trains a model to f(x) to approximate y (304). Using X as X t and X c the method obtains the remaining conformal intervals of X from the conformal inference calculation. For each b in B, the method selects an x in X that maximizes f(x)+C * interval(x) (308) and recalculates the conformal interval as if the selected x had been observed (310). The method then determines whether any b that has not yet been evaluated in 308 or 310 remains in B (312). If b remains, the method repeats using the unevaluated b in B. Otherwise, the method determines whether any further iterations are required (314), and after N iterations, the method terminates (316).

[0055] The above method can be used for general data and data points, but the applicant has noted that the above method can be used for optimizing the design of biopolymer sequences. Examples of biopolymer sequences include amino acid sequences, nucleotide sequences, and carbohydrate sequences.

[0056] The amino acid sequence can include canonical amino acids, non-canonical amino acids, or combinations thereof, and can further include L-amino acids and / or D-amino acids. The amino acid sequence can also include amino acid derivatives and / or modified amino acids. Non-limiting examples of amino acid modifications include amino acid linkers, acylation, acetylation, amidation, methylation, terminal modification factors (such as cyclization modification), and N-methyl-α-amino group substitution.

[0057] The nucleotide sequence can include naturally occurring ribonucleotides or deoxyribonucleotide monomers and non-naturally occurring nucleotide derivatives and analogs thereof. Thus, nucleotides can include, for example, nucleotides containing naturally occurring bases (such as A, G, C, or T) and nucleotides containing modified bases (such as 7-deazaguanosine, inosine, or methylated nucleotides such as 5-methyldCTP and 5-hydroxymethylcytosine).

[0058] Examples of the properties (such as functional values) of the biopolymer sequences (such as amino acid sequences) analyzed by the model are the binding affinity, binding specificity, catalytic (such as enzyme) activity, fluorescence, solubility, thermal stability, structure, immunogenicity, and any other functional properties of the biopolymer sequence.

[0059] Described herein are devices, software, systems, and methods for evaluating input data containing protein or polypeptide information such as amino acid sequences (or nucleic acid sequences encoding amino acid sequences) and predicting one or more specific functions or properties based on the input data. The elucidation of specific functions or properties of amino acid sequences (e.g., proteins) has long been a goal of molecular biology. Thus, the devices, software, systems, and methods described herein utilize the capabilities of artificial intelligence or machine learning techniques for polypeptide or protein analysis to make predictions about structure and / or function. The machine learning techniques described herein enable the generation of models with increased predictive power compared to standard non-ML approaches.

[0060] In some aspects, the input data includes the primary amino acid sequence of a protein or polypeptide. In some cases, the model is trained using a labeled dataset that includes the primary amino acid sequence. For example, the dataset can include the amino acid sequences of fluorescent proteins labeled based on fluorescence intensity. However, other types of proteins labeled based on other properties can also be employed. Thus, the model can be trained on this dataset using machine learning methods to generate predictions of the fluorescence intensity for amino acid sequence inputs. In some aspects, the input data includes, in addition to the primary amino acid sequence, information such as, for example, surface charge, hydrophobic surface area, measured or predicted solubility, or other relevant information. In some aspects, the input data includes multi-dimensional input data that includes multiple types or categories of data.

[0061] In some aspects, the devices, software, systems, and methods described herein utilize data augmentation to enhance the performance of predictive models. Data augmentation involves training on similar but different examples or variations of a training dataset. As an example, in image classification, image data can be augmented by slightly changing the orientation of the image (e.g., a slight rotation). In some aspects, data inputs (e.g., primary amino acid sequences) are augmented by random mutations and / or biologically informed mutations to the primary amino acid sequence, multiple sequence alignments, contact maps of amino acid interactions, and / or tertiary protein structures. Additional augmentation strategies include the use of known and predicted isoforms from alternative splicing transcripts. For example, input data can be augmented by including isoforms of alternative splicing transcripts that correspond to the same function or property. Thus, data about isoforms or mutations can enable the identification of portions or features of the primary sequence that do not significantly affect the predicted function or property. This allows the model to, for example, enhance, reduce, or not be affected by predicted protein properties such as stability and take into account information such as amino acid mutations. For example, the data input can include sequences with random substitution amino acids at positions known to have no effect on function. This allows the model trained on this data to learn that the predicted function is invariant with respect to those specific mutations.

[0062] The devices, software, systems, and methods described herein can be used to generate a wide variety of predictions. The predictions can include protein functions and / or properties (e.g., enzyme activity, binding properties, stability, etc.). Protein stability can be predicted according to various measures such as, for example, thermal stability, oxidative stability, or serum stability. In some aspects, the predictions include one or more structural features such as, for example, secondary structure, tertiary protein structure, quaternary structure, or any combination thereof. The secondary structure can include an indication of whether the sequence of amino acids within an amino acid or polypeptide has an alpha-helix structure, a beta-sheet structure, or a disordered or loop structure. The tertiary structure can include the location or position of the portions of the amino acid or polypeptide in three-dimensional space. The quaternary structure can include the location or position of the multiple polypeptides that form one protein. In some aspects, the predictions include one or more functions. The functions of a polypeptide or protein can belong to various categories including metabolic reactions, DNA replication, providing structure, transport, antigen recognition, intracellular or extracellular signaling, and other functional categories. In some aspects, the predictions include enzyme functions such as, for example, catalytic efficiency (e.g., specificity constant k cat / K M ) or catalytic specificity.

[0063] In some aspects, the predictions include the enzyme function of a protein or polypeptide. In some aspects, the protein function is an enzyme function. Enzymes can perform various enzyme reactions and can be classified as transferases (e.g., transfer functional groups from one molecule to another), oxidoreductases (e.g., catalyze oxidation-reduction reactions), hydrolases (e.g., cleave chemical bonds via hydrolysis), lyases (e.g., generate double bonds), ligases (e.g., link two molecules via covalent bonds), and isomerases (e.g., catalyze structural changes from one isomer to another within a molecule).

[0064] In some aspects, protein functions include enzymatic functions, binding (e.g., DNA / RNA binding, protein binding, antibody-antigen binding, etc.), immune functions (e.g., antibodies, cytokines, checkpoint molecules, etc.), contraction (e.g., actin, myosin), and other functions. In some aspects, the output includes values related to protein functions such as, for example, the kinetics of enzymatic functions or binding. Such output can include measures of affinity, specificity, and reaction rate.

[0065] In some aspects, the machine learning methods described herein include supervised machine learning. Supervised machine learning includes classification and regression. In some aspects, the machine learning methods include unsupervised machine learning. Unsupervised machine learning includes clustering, autoencoding, variational autoencoding, protein language models (e.g., a model that predicts the next amino acid in a sequence given access to previous amino acids), and correlation rule mining.

[0066] In some embodiments, the prediction includes a classification such as binary, multi-label, or multi-class classification. Classification is generally used to predict discrete classes or labels based on input parameters. Binary classification predicts which of two groups a polypeptide or protein belongs to based on the input. In some embodiments, binary classification includes a positive or negative prediction about the properties or functions of a protein or polypeptide sequence. In some embodiments, binary classification includes any quantitative readout that undergoes a threshold process, such as binding to a DNA sequence above a certain affinity level, catalyzing a reaction above a certain threshold of kinetic parameters, or exhibiting thermal stability above a certain melting temperature. Examples of binary classification include positive / negative predictions that a polypeptide sequence exhibits autofluorescence, is a serine protease, or is a GPI-anchored transmembrane protein. In some embodiments, the classification is multi-class classification. For example, multi-class classification can classify an input polypeptide into one of more than two groups. Alternatively, the prediction can include multi-label classification. Multi-class classification classifies the input into one of mutually exclusive categories, while multi-label classification classifies the input into multiple labels or groups. For example, multi-label classification can label a polypeptide as both an intracellular protein (compared to extracellular) and a protease. By comparison, multi-class classification can include classifying an amino acid as belonging to one of an alpha helix, beta sheet, or disordered / loop peptide sequence.

[0067] In some embodiments, the prediction includes a regression that provides a continuous variable or value, such as the intensity of autofluorescence or the stability of a protein. In some embodiments, the prediction includes a continuous variable or value of any of the properties or functions described herein. As an example, the continuous variable or value can indicate the target specificity of a matrix metalloprotease for a particular substrate extracellular matrix component. Additional examples include various quantitative readouts such as target molecule binding affinity (e.g., DNA binding), the reaction rate of an enzyme, or thermal stability.

[0068] To show the effectiveness of the method described above, consider a comparison of CI-OPT using nearest neighbor conformational scores in two synthetic Bayesian optimization tasks and two empirically determined protein fitness datasets with Gaussian process-based optimization. The protein datasets have high-dimensional discrete spaces where any GP using conventional kernels is expected to be strongly misspecified.

[0069] The following is an overview of the methods to be evaluated below. a) GP is Bayesian optimization using either a Gaussian process surrogate function and a UCB or MI acquisition function. b) GP-CI: CI-OPT using a Gaussian process to calculate μ according to Equation 6 t,CI and using conformal inference to calculate σ according to Equations 7 and 9, with either a UCB or MI acquisition function. t,CI c) NN-CI: CI-OPT using a neural network to calculate μ according to Equation 6 and using conformal inference to calculate σ according to Equations 7 and 9, with either a UCB or MI acquisition function. t,CI t,CI

[0070] The Branin or Branin-Hoo function is a common black-box optimization benchmark with three global optima in the 2D square [-5, 10] × [0, 15]. One example of a black-box optimization benchmark has an output that is normalized to have approximately mean 0 and variance 1 for numerical stability, as described in "Botorch: Programmable Bayesian Optimization in pytorch" by Balandat et al., arXiv preprint arXiv:1910.06403, 2019 (hereinafter "Botorch" or "Balandat").

[0071] The Hartmann function is another common black-box optimization benchmark. According to the Botorch literature, the 6D version is [0, 1]6 It is evaluated in. The Hartmann function has six local maxima and one global maximum.

[0072] The GB1 dataset contains the measured fitness values of most of the sequences in the four-site site-saturation library of protein G domain B1 with a total of 160,000 sequences, as described in "Adaptation In Protein Fitness Landscapes Is Facilitated By Indirect Paths", Elife, 5:e16965, 2016 by Wu et al. (hereinafter "Wu"). For missing sequences, the values ascribed by Wu can be used. The dataset is designed to capture the non-linear interactions between positions and amino acids.

[0073] The FITC dataset consists of the binding affinities Adams (2016) of thousands of variants of a well-studied scFv antibody to fluorescein isothiocyanate (FITC). Mutations were made in the CDR1H and CDR3H regions. The binding constant k D The lower it is, the stronger the binding, so in this case, the task is to maximize -logk D is.

[0074] For the synthetic task, CI-OPT using the UCB acquisition function and the GP surrogate model or the neural network surrogate model is compared with GP-UCB using the same GP model. GP and GP-UCB in the synthetic task according to the default in Botorch (e.g., the Matern kernel with ν = 2.5 with a strong prior in noise and length scale) are run using the reparameterization implementation in Botorch. The neural network included two hidden layers of dimension 256 connected with ReLU activation. The weights were optimized using Adam by Kingma et al. in arXiv preprint arXiv:1412.6980 (hereinafter "Adam" or "Kingma"), i.e., the stochastic optimization method, and L 2 The weight decay is 1e -3 is set to.

[0075] In each run, the method is initialized with 10 randomly selected observations. The experiment is repeated 64 times with different initializations. Conformal inference uses β = 1e -2 , Euclidean distance, and 5 nearest neighbors. The GP is retrained at each iteration. The neural network is first trained with 1000 mini - batches and then fine - tuned with an additional 100 mini - batches after each observation.

[0076] The demonstration of the effectiveness of the system and method using several real - world protein datasets is further described below. In the protein task, CI - OPT using the MI acquisition function is compared with GP - MI in sequential and batch settings. The GP in the protein task uses a squared - exponential kernel with hyperparameters chosen to maximize the marginal likelihood. CI - OPT is pre - trained with proteins from UniProt as disclosed in “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences” by Rives, bioRxiv, pp.622803, 2019 (hereinafter “Rives”) and then fine - tuned with observations, and uses the Transformer language model as described in “Attention is All You Need” by Vaswani, Advances in Neural Information Processing Systems, pp.5998 - 6008, 2017 (hereinafter “Vaswani”). In both datasets, CI - OPT adopts the Hamming distance and 5 nearest neighbors to calculate the conformal score. CI - OPT and the greedy method are repeated 10 times with different initial points, while the GP is repeated 25 times.

[0077] In the biological optimization problem, the goal is to find good rewards as quickly as possible. However, usually there is no penalty for evaluating inputs that lead to bad rewards along the way. Therefore, as will be further explained below, the method is evaluated by comparing the maximum rewards found by each method at iteration t instead of the average regret.

[0078] Figures 1A and 1B are graphs showing the results of sequential optimization in two synthetic tasks. In the 2D Branin task, GP-UCB, GP-CI, and NN-CI all quickly find the global maximum. In the 6D Hartmann task, GP-CI is on par with GP-UCB, but the performance of NN-CI drops. However, these results were obtained using neural networks without adjusting the hyperparameters of the neural networks.

[0079] Figures 1C and 1E are graphs showing the results of sequential optimization in the protein dataset. In these high-dimensional discrete spaces, NN-CI consistently outperforms GP-based methods. This performance is due to both the fact that pre-trained neural networks are much more accurate than GP and the GP uncertainty is miscalibrated, eliminating their theoretical advantages.

[0080] Figures 1D and 1F are graphs showing similar results in batch optimization in the protein dataset. Optimization with large batches is extremely difficult because each batch has to balance exploration and exploitation to maximize the acquisition function. The batch size 100 used here for GB1 is much larger than the batch sizes typically seen in Bayesian optimization experiments. For example, Wilson considered a maximum batch size of 16. However, 100 is a realistic batch size in protein engineering experiments.

[0081] Conformal inference optimization uses the prediction interval induced by the nearest neighbor-based conformal score for regression as a drop-in replacement for the GP posterior uncertainty in the acquisition function based on the upper confidence bound in black-box function optimization. This method is more suitable by leveraging large pre-trained neural networks in the optimization loop than the conventional BO method based on GP. CI-OPT is on par with GP-based Bayesian optimization on synthetic tasks and outperforms the GP-based method on two different protein optimization datasets.

[0082] FIG. 4 is a flowchart 400 illustrating an example of one aspect of the present disclosure. In one aspect, a computer-implemented method for optimizing the design of a biopolymer sequence can include training a machine learning model (402) using an observed biopolymer sequence and a labeled biopolymer sequence corresponding to each observed biopolymer sequence. The labeled sequence is a sequence associated with a real number that measures some property of interest. The method can further include determining a biopolymer sequence candidate and observing (404) the one having the highest predicted value of the labeled biopolymer sequence based on the machine learning model. The biopolymer sequence candidate can include either a known sequence (e.g., a sequence faced previously, an observed sequence previously, or a natural sequence) or a newly designed sequence. The method can further include, for each biopolymer sequence candidate (408), specifying (406) a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence. The method can further include selecting (410) at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence. In one aspect, the value of the labeled sequence is the number used as a label as described above. Thus, the predicted value of the sequence is the predicted label of the sequence. One skilled in the art of machine learning can understand such a definition of the label. The sequence or data point is the machine learning input (x), and the prediction / measurement / optimization is the label (y).

[0083] FIG. 5 is a flowchart 500 showing an example of an aspect of the present disclosure. In one aspect, a computer-implemented method and corresponding system for optimizing the design of a polymer sequence trains a model to approximate a labeled biopolymer sequence of an initial sample from a plurality of observed sequences (502). The method can further include selecting (504) at least one sequence from the plurality of observed sequences for a particular batch of the plurality of observed sequences to optimize a combination of the labeled polymer sequence generated by the trained model and the conformational interval. The batch has the labeled biopolymer sequence generated by the trained model and the conformational interval of each observed sequence. If the entire batch has not been analyzed (506), the method selects the next sequence (504). If the entire batch has been analyzed (506), the method can further include recalculating the conformational intervals of the remaining sequences (508).

[0084] FIG. 6 shows a computer network or similar digital processing environment in which aspects of the invention may be implemented.

[0085] Client computer / devices 50 and server computer 60 provide processing, storage, and input / output devices for executing application programs and the like. Client computer / devices 50 can also link through communication network 70 to other computing devices including other client devices / processes 50 and server computer 60. Communication network 70 can be a part of a remote access network, a global network (e.g., the Internet), a worldwide collection of computers, a local area or wide area network, and a gateway that communicates with each other using current protocols (TCP / IP, Bluetooth®, etc.). Other electronic device / computer network architectures are also suitable.

[0086] FIG. 7 is a diagram of an example of the internal structure of a computer (e.g., client processor / device 50 or server computer 60) in the computer system of FIG. 6. Each computer 50, 60 includes a system bus 79, which is a set of hardwired lines used for data transfer between components of a computer or processing system. The system bus 79 is basically a shared conduit that connects different elements of the computer system (e.g., processor, disk storage, memory, input / output ports, network ports, etc.) to enable information transfer between the elements. Attached to the system bus 79 is an I / O device interface 82 for connecting various input and output devices (e.g., keyboard, mouse, display, printer, speaker, etc.) to the computers 50, 60. The network interface 86 enables the computer to connect to various other devices attached to a network (e.g., network 70 of FIG. 5). The memory 90 provides volatile storage of computer software instructions 92 and data 94 (e.g., the Bayesian optimization module and the conformal inference module code detailed above) used in the implementation of one aspect of the present invention. The disk storage 95 provides non-volatile storage of computer software instructions 92 and data 94 used in the implementation of one aspect of the present invention. The central processing unit 84 is also attached to the system bus 79 and provides the execution of computer instructions.

[0087] In one aspect, the processor routine 92 and the data 94 include a computer program product (e.g., a removable storage medium such as one or more flash memories, DVD-ROM, CD-ROM, floppy disk, tape, etc.) that provides at least a portion of the software instructions of the system of the present invention, a non-transitory computer-readable medium (generally referred to as 92). The computer program product 92 can be installed by any suitable software installation procedure, as is well known in the art. In another aspect, at least a portion of the software instructions can also be downloaded via cable communication and / or wireless communication. In other aspects, the program of the present invention is a computer program propagation signal product implemented by a propagation signal in a propagation medium (e.g., radio waves, microwaves, infrared waves, laser waves, sound waves, or radio waves propagated via a global network such as the Internet or other networks). Such a carrier medium or signal can be employed to provide at least a portion of the software instructions of the routine / program 92 of the present invention.

[0088] All teachings of the patents, published applications, and cited references described herein are hereby incorporated by reference in their entirety.

[0089] Although example embodiments have been specifically shown and described, it will be understood by those skilled in the art that various changes in form and detail can be made without departing from the scope of the embodiments encompassed by the appended claims. Note that the present invention includes the following content as an aspect. 〔Aspect 1〕 A computer-implemented method for optimizing the design of a biopolymer sequence, training a machine learning model using a plurality of observed biopolymer sequences and the labeled biopolymer sequences corresponding to each observed biopolymer sequence, determining a plurality of biopolymer sequence candidates and observing the one having the highest predicted value of the labeled biopolymer sequence based on the machine learning model, for each biopolymer sequence candidate, specifying a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence, selecting at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence, A computer-implemented method comprising. 〔Aspect 2〕 The computer-implemented method according to Aspect 1, wherein the conformal inference interval includes a central value and an interval range. 〔Aspect 3〕 The computer-implemented method according to Aspect 2, wherein the central value is an average value. 〔Aspect 4〕 The computer-implemented method according to Aspect 1, wherein the machine learning model is a neural network fine-tuned using the observed biopolymer sequences and their labels. 〔Aspect 5〕 The computer-implemented method according to Aspect 4, wherein specifying the conformal inference interval is based on a second set of observed biopolymer sequences. 〔Aspect 6〕 Specifying the conformal inference interval calculating a residual interval based on each output of the machine learning model for each of the second set of observed biopolymer sequences and the corresponding labeled biopolymer sequences corresponding to each of the second set of biopolymer sequences, calculating the average distance to a plurality of nearest neighbor sequences of the observed biopolymer sequences in the metric space for each output of the machine learning model, calculating a conformal score based on the ratio of the residual to the sum of the average distance and a constant in the metric space, The computer-implemented method according to Aspect 5, further comprising. 〔Aspect 7〕 Selecting the at least one biopolymer sequence candidate calculating the average distance in the metric space to a plurality of nearest neighbor sequences in the metric space, generating a confidence interval based on the at least one candidate biopolymer sequence and the average distance; selecting at least one candidate biopolymer sequence based on the confidence interval; The computer-implemented method according to aspect 5, comprising: [Aspect 8] The method according to aspect 1, wherein the conformational interval is at least 50% and at most 99%. [Aspect 9] The method according to aspect 1, wherein the biopolymer sequence comprises at least one of an amino acid sequence, a nucleic acid sequence, and a carbohydrate sequence. [Aspect 10] The method according to aspect 9, wherein the nucleic acid sequence is a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence. [Aspect 11] The method according to aspect 1, wherein the predicted value is a functional value of the biopolymer sequence, and the function is one or more of binding affinity, binding specificity, catalytic activity, enzyme activity, fluorescence, solubility, thermal stability, structure, immunogenicity, and functional properties of the biopolymer sequence. [Aspect 12] The method according to aspect 1, wherein selecting the at least one candidate biopolymer sequence has increased performance compared to Bayesian optimization that does not decompose the specified conformational inference interval. [Aspect 13] A computer-implemented method for optimizing the design of a biopolymer sequence, comprising: training a model to approximate a labeled biopolymer sequence of an initial sample from a plurality of observed sequences; for a particular batch of the plurality of observed sequences having a labeled biopolymer sequence generated by the trained model and a conformational interval of each observed sequence, selecting at least one sequence from the plurality of observed sequences to optimize the combination of the labeled polymer sequence and the conformational interval generated by the trained model; recalculating the conformational intervals of the remaining sequences; The computer-implemented method comprising: [Aspect 14] The computer-implemented method according to aspect 13, further comprising repeating selecting the at least one sequence and recalculating the conformational interval for each of a plurality of batches. [Aspect 15] The method according to aspect 13, further comprising identifying an optimal number of batch experiments to run in parallel. [Aspect 16] The method according to aspect 15, wherein identifying is based on optimization of wet lab resources. [Aspect 17] A computer-implemented method for optimizing a design based on a distribution of data, comprising: Training a machine learning model using a plurality of observation data and the labeled data corresponding to each observation data, Determining a plurality of data candidates and observing those having the highest predicted value of the labeled data based on the machine learning model, For each data candidate, specifying a conformal inference interval representing the likelihood that the data candidate has the predicted value of the labeled data, Selecting at least one data candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled data, A computer-implemented method comprising. [Aspect 18] The method according to any one of Aspects 1 to 17, further comprising providing the at least one selected biopolymer sequence to means for synthesizing the selected biopolymer sequence. [Aspect 19] The method according to Aspect 18, wherein the at least one selected biopolymer sequence is synthesized. [Aspect 20] The method according to any one of Aspects 1 to 19, further comprising synthesizing the at least one selected biopolymer sequence. [Aspect 21] The method according to Aspect 18 or 20, further comprising assaying the at least one selected biopolymer sequence, for example, in a qualitative or quantitative chemical assay. [Aspect 22] A non-transitory computer-readable medium storing instructions for optimizing the design of a biopolymer sequence, the instructions, when executed by a processor, cause the processor to Train a machine learning model using a plurality of observed biopolymer sequences and the labeled biopolymer sequences corresponding to each observed biopolymer sequence, Determine a plurality of biopolymer sequence candidates and observe those having the highest predicted value of the labeled biopolymer sequence based on the machine learning model, For each biopolymer sequence candidate, specify a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence, Select at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence, A non-transitory computer-readable medium that causes the above to be performed. [Aspect 23] A system for optimizing the design of a biopolymer sequence, A processor, A memory storing computer code instructions, comprising, the processor and the memory using the computer code instructions to cause the system to train a machine learning model using a plurality of observed biopolymer sequences and labeled biopolymer sequences corresponding to each of the observed biopolymer sequences; determine a plurality of biopolymer sequence candidates and observe those having the highest predicted value of the labeled biopolymer sequence based on the machine learning model; for each biopolymer sequence candidate, identify a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the predicted value of the labeled biopolymer sequence; select at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the predicted value of the labeled biopolymer sequence; A system configured to perform the above. 〔Aspect 24〕 One or more selected biopolymer sequences obtainable by the method according to any one of Aspects 1 to 21. 〔Aspect 25〕 The one or more selected biopolymer sequences are the one or more selected polypeptide sequences produced by a method of culturing a host cell containing one or more nucleic acids encoding the one or more selected polypeptide sequences under conditions that promote the synthesis of the one or more selected polypeptide sequences and isolating the one or more selected polypeptide sequences. The one or more selected biopolymer sequences according to Aspect 24. 〔Aspect 26〕 A composition comprising the one or more selected biopolymer sequences according to Aspect 24 or 25, wherein the one or more selected biopolymer sequences comprise a pharmaceutically acceptable excipient. 〔Aspect 27〕 A method comprising contacting the composition or the selected biopolymer sequence according to any one of Aspects 24 to 26 with one or more of a test compound, a biological fluid, a cell, a tissue, an organ, or an organism.

Claims

1. A computer-implemented method for optimizing the design of a biopolymer sequence, comprising: training a machine learning model using a plurality of observed biopolymer sequences and labeled biopolymer sequences corresponding to each of the observed biopolymer sequences; determining a plurality of biopolymer sequence candidates and observing, based on the machine learning model, those having the highest predicted value of the labeled biopolymer sequence; for each biopolymer sequence candidate, identifying a conformal inference interval representing the likelihood that the biopolymer sequence candidate has the highest predicted value of the labeled biopolymer sequence; selecting at least one biopolymer sequence candidate having an optimized linear combination of the conformal inference interval and the highest predicted value of the labeled biopolymer sequence; A computer-implemented method comprising the above.

2. The computer-implemented method according to claim 1, wherein the conformal inference interval includes a central value and an interval range.

3. The computer-implemented method according to claim 2, wherein the central value is an average value.

4. The computer-implemented method according to claim 1, wherein the machine learning model is a neural network fine-tuned using the observed biopolymer sequences and their labels.

5. The computer-implemented method according to claim 4, wherein identifying the conformal inference interval is based on a second set of observed biopolymer sequences.

6. Identifying the conformal inference interval further includes: calculating a residual interval for each output of the machine learning model for each corresponding labeled biopolymer sequence corresponding to the second set of observed biopolymer sequences and the second set of biopolymer sequences; calculating, for each output of the machine learning model, an average distance to a plurality of nearest neighbor sequences of the observed biopolymer sequences in a metric space; calculating a conformal score based on a ratio of the residual to a sum of the average distance and a constant; The computer-implemented method according to claim 5, further comprising the above.

7. Selecting the at least one biopolymer sequence candidate further includes: calculating an average distance in the metric space to a plurality of nearest neighbor sequences in the metric space; generating a confidence interval based on the at least one biopolymer sequence candidate and the average distance; selecting the at least one candidate biopolymer sequence based on the generated confidence interval; The computer-implemented method according to claim 5, comprising: **Claim 8** The computer-implemented method according to claim 1, wherein the identified conformational inference interval is at least 50% and at most 99%. **Claim 9** The computer-implemented method according to claim 1, wherein the at least one selected candidate biopolymer sequence comprises at least one of an amino acid sequence, a nucleic acid sequence, and a carbohydrate sequence. **Claim 10** The computer-implemented method according to claim 9, wherein the nucleic acid sequence is a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence. **Claim 11** The computer-implemented method according to claim 1, wherein the highest predicted value is a functional value of the biopolymer sequence, and the functional value is one or more of binding affinity, binding specificity, catalytic activity, enzyme activity, fluorescence, solubility, thermal stability, structure, immunogenicity, and functional properties of the biopolymer sequence. **Claim 12** The computer-implemented method according to claim 1, wherein selecting the at least one candidate biopolymer sequence has increased performance compared to Bayesian optimization without decomposing the identified conformational inference interval. **Claim 13** A computer-implemented method for optimizing the design of a biopolymer sequence, comprising: training a model to approximate a labeled biopolymer sequence of an initial sample from a plurality of observed sequences; selecting, from the plurality of observed sequences, at least one sequence that optimizes a combination of the labeled polymer sequence and the conformational interval generated by the trained model for a particular batch of the plurality of observed sequences having the conformational intervals of the labeled biopolymer sequences and each observed sequence generated by the trained model; recalculating the conformational intervals of the remaining sequences; The computer-implemented method comprising: **Claim 14** The computer-implemented method according to claim 13, further comprising repeating selecting the at least one sequence and recalculating the conformational intervals for each of a plurality of batches. **Claim 15** The computer-implemented method according to claim 13, further comprising identifying an optimal number of batch experiments to run in parallel. **Claim 16** Said identifying is a computer-implemented method according to claim 15, based on the optimization of wet lab resources.

17. A computer-implemented method for optimizing a design based on the distribution of data, comprising: training a machine learning model using a plurality of observed data and labeled data corresponding to each observed data; determining a plurality of data candidates and observing those having the highest predicted value of the labeled data based on the machine learning model; for each data candidate, identifying a conformal inference interval representing the likelihood that the data candidate has the highest predicted value of the labeled data; selecting at least one data candidate having an optimized linear combination of the conformal inference interval and the highest predicted value of the labeled data, wherein the at least one selected data candidate corresponds to at least one selected biopolymer sequence candidate; A computer-implemented method comprising.

18. The computer-implemented method according to any one of claims 1 to 17, further comprising providing the at least one selected biopolymer sequence candidate to means for synthesizing the at least one selected biopolymer sequence candidate.

19. The computer-implemented method according to claim 18, wherein the at least one selected biopolymer sequence candidate is synthesized.

20. The computer-implemented method according to any one of claims 1 to 19, further comprising synthesizing the at least one selected biopolymer sequence candidate.

21. The computer-implemented method according to claim 18 or 20, further comprising assaying the at least one selected biopolymer sequence candidate in a qualitative or quantitative chemical assay.

22. A non-transitory computer-readable medium storing instructions for optimizing the design of a biopolymer sequence, which instructions, when executed by a processor, cause the processor to: train a machine learning model using a plurality of observed biopolymer sequences and labeled biopolymer sequences corresponding to each observed biopolymer sequence; determine a plurality of biopolymer sequence candidates and observe those having the highest predicted value of the labeled biopolymer sequence based on the machine learning model; For each candidate biomacromolecule sequence, identifying a conformal inference interval representing the likelihood that the candidate biomacromolecule sequence has the highest predicted value of the labeled biomacromolecule sequence; selecting at least one candidate biomacromolecule sequence having an optimized linear combination of the conformal inference interval and the highest predicted value of the labeled biomacromolecule sequence; A non-transitory computer-readable medium that causes the above to be performed.

23. A system for optimizing the design of a biomacromolecule sequence, comprising: a processor; a memory storing computer code instructions; The processor and the memory are configured to cause the system to: train a machine learning model using a plurality of observed biomacromolecule sequences and labeled biomacromolecule sequences corresponding to each observed biomacromolecule sequence; determine a plurality of candidate biomacromolecule sequences and observe those having the highest predicted value of the labeled biomacromolecule sequence based on the machine learning model; for each candidate biomacromolecule sequence, identifying a conformal inference interval representing the likelihood that the candidate biomacromolecule sequence has the highest predicted value of the labeled biomacromolecule sequence; selecting at least one candidate biomacromolecule sequence having an optimized linear combination of the conformal inference interval and the highest predicted value of the labeled biomacromolecule sequence; A system configured to cause the above to be performed.