Generative protein design with smoothed energy-based models
The protein design computational model addresses overfitting and limited exploration by training on a noisy dataset, generating diverse and therapeutically viable protein sequences with improved drug-like properties through gradient-based sampling and denoising.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2026-03-10
AI Technical Summary
Existing computational protein design methods face challenges in efficiently generating diverse and therapeutically viable protein sequences due to overfitting and limited exploration of the high-dimensional data distribution, leading to repetitive and less diverse output sequences.
A protein design computational model utilizing an energy-based model (EBM) is trained on a noisy training set to approximate the data distribution, employing gradient-based Markov chain Monte Carlo sampling to generate protein sequences by manipulating noisy embeddings and denoising them to increase diversity and likelihood of desirable properties.
The model effectively generates protein sequences with enhanced diversity and increased likelihood of exhibiting drug-like properties, such as expression, affinity, and stability, by navigating a smoothed energy landscape and capturing three-dimensional structural information.
Smart Images

Figure 2026508122000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 482,756, entitled "GENERATIVE PROTEIN DESIGN WITH SMOOTHED ENERGY-BASED MODELS," filed February 1, 2023, U.S. Provisional Application No. 63 / 502,497, entitled "GENERATIVE PROTEIN DESIGN WITH SMOOTHED ENERGY-BASED MODELS," filed May 16, 2023, and U.S. Provisional Application No. 63 / 588,437, entitled "GENERATIVE PROTEIN DESIGN WITH SMOOTHED ENERGY-BASED MODELS," filed October 6, 2023, the disclosures of which are incorporated herein by reference in their entireties.
[0002] The subject matter described herein relates generally to computational protein design, and more particularly to energy-based models (EBMs) for generating protein sequences. [Background technology]
[0003] Proteins are genetically encoded macromolecules with incredible diversity in size and chemical composition. By regulating biological systems, proteins facilitate many essential cellular functions, including, for example, enzymatic reactions, molecular transport, regulation and execution of numerous biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and the like. Protein structures can include one or more polypeptides, each of which is a chain of amino acid residues linked to each other by peptide bonds (e.g., covalent peptide bonds). Unlike non-canonical amino acid residues, there are 20 canonical amino acid residues that are directly encoded by genetic coding. Each canonical amino acid residue consists of the same backbone atom (e.g., an amino group (NH), an α-carbon ( TIFF2026508122000002.tif5170), and a carboxylic acid group (COOH).
[0004] The primary structure of a protein molecule refers to the arrangement of amino acid residues in each polypeptide chain that forms the protein structure. The backbone atoms of adjacent amino acid residues involved in peptide bonds (e.g., covalent peptide bonds) between them form a repeating sequence of atoms known as the polypeptide backbone (or skeleton) of the protein molecule. The secondary structure of a protein molecule refers to the local folded structures (e.g., α-helices, β-pleated sheets, etc.) that form within individual polypeptide chains due to interactions between backbone atoms (e.g., amino hydrogen atoms, carboxyl oxygen atoms, etc.). Additional interactions (e.g., non-covalent bonds such as hydrogen bonds, ionic bonds, dipole-dipole interactions, and van der Waals forces) between the side chains (or R groups) of the amino acid residues of a protein molecule can cause folding within individual polypeptide chains, thus forming the tertiary structure of the protein molecule. The tertiary structure of a protein molecule is also known as the conformation or three-dimensional structure of the protein molecule. In protein molecules with multiple polypeptide chains, the protein molecule may also exhibit quaternary structure, which is formed when the polypeptide chains are packed and held together by hydrogen bonds and van der Waals forces (e.g., between nonpolar side chains).
[0005] The function of a protein molecule can depend on the sequence of amino acid residues in the polypeptide chain that forms the protein molecule, as well as the three-dimensional structure adopted by the polypeptide chain. For example, the primary structure of a protein molecule can determine the three-dimensional structure assumed by the protein molecule through the folding of the constituent polypeptide chains. In some cases, the binding affinity of a protein molecule to a target molecule, such as a virus or tumor antigen, can depend on whether the polypeptide chain of the protein molecule can assume a three-dimensional structure that complements the three-dimensional structure of the target molecule and is stable enough to allow a binding interaction between the two molecules. Therefore, one prominent goal of computational protein design is to construct one or more protein sequences (e.g., antibodies) that exhibit specific desirable properties. For example, in the case of large molecule drug discovery (LMDD), computational protein design may seek to identify therapeutically viable protein sequences (e.g., antibodies) that possess various desired properties, such as expression, binding affinity for the target molecule, binding specificity, stability, non-immunogenicity, humanity, absence of self-association (or non-aggregation), and lack of chemical cleavage (e.g., aspartate isomerization, oxidation, deamidation). Summary of the Invention
[0006] Systems, methods, and articles of manufacture, including computer program products, are provided for generative protein design that applies an energy-based model (EBM) to generate protein sequences. In one aspect, a system for generative protein design is provided, including at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include generating a first training set including a plurality of noisy sample sequences, where each noisy sample sequence in the first training set is generated by at least adding noise to a corresponding sample sequence from a first data distribution; training a protein design computational model by applying the protein design computational model to generate at least one or more output sequences and adjusting the protein design computational model to reduce differences between the one or more generated output sequences and the plurality of noisy sample sequences of the first training set; and applying the trained protein design computational model to generate output sequences by modifying at least input sequences.
[0007] In another aspect, a method for generative protein design is provided. The method may include generating a first training set including a plurality of noisy sample sequences, where each noisy sample sequence in the first training set is generated by at least adding noise to a corresponding sample sequence from a first data distribution, training a protein design computational model by applying the protein design computational model to generate at least one or more output sequences and adjusting the protein design computational model to reduce differences between the one or more generated output sequences and the plurality of noisy sample sequences of the first training set, and applying the trained protein design computational model to generate output sequences by modifying at least input sequences.
[0008] In another aspect, a computer program product is provided that includes a non-transitory computer-readable medium storing instructions. The instructions may cause operations to be executed by at least one data processor. The operations may include generating a first training set including a plurality of noisy sample sequences, where each noisy sample sequence in the first training set is generated by at least adding noise to a corresponding sample sequence from a first data distribution; training a protein design computational model by applying the protein design computational model to generate at least one or more output sequences and adjusting the protein design computational model to reduce differences between the one or more generated output sequences and the plurality of noisy sample sequences of the first training set; and applying the trained protein design computational model to generate output sequences by modifying at least input sequences.
[0009] In some variations, one or more of the features disclosed herein may be included, in any operable combination, including the following features:
[0010] In some variations, the protein design computational model includes a first energy-based model (EBM).
[0011] In some variations, training the protein design computational model includes adjusting a plurality of parameters of the first energy-based model that parameterize an energy function of the first energy-based model.
[0012] In some variations, the plurality of parameters are adjusted such that energy values determined by the energy function correspond to likelihoods of the one or more generated output sequences within the first data distribution.
[0013] In some variations, the plurality of parameters are adjusted such that the energy function outputs a lower energy value for a first generated output sequence that is similar to the plurality of noisy samples of the first training set than for a second generated output sequence that is dissimilar to the plurality of noisy samples of the first training set.
[0014] In some variations, training the protein design computational model includes applying a first energy-based model with a first adjustment to generate a first modified sequence, applying the first energy-based model with a second adjustment to generate a second modified sequence, and upon determining that the first modified sequence is more similar to the plurality of noisy samples of the first training set than the second modified sequence, further modifying the first energy-based model with the first adjustment instead of the second adjustment.
[0015] In some variations, the first energy-based model is further adjusted until one or more criteria are met, including at least one of (i) performing a threshold number of iterations of adjustments to the first energy-based model, and (ii) the second modified sequence exhibiting a threshold similarity to a plurality of noisy samples of the first training set.
[0016] In some variations, the protein design computational model further includes a second energy-based model (EBM).
[0017] In some variations, a second training set including a plurality of sample sequences from a second data distribution may be generated. A first adjustment to the first energy-based model may be determined to reduce a first difference between a first output sequence generated by the first energy-based model and the plurality of noisy sample sequences of the first training set. A second adjustment to the second energy-based model may be determined to reduce a second difference between a second output sequence generated by the second energy-based model and the plurality of sample sequences of the second data distribution. The first energy-based model may be trained by applying to the first energy-based model at least a third adjustment determined based on the first adjustment and the second adjustment.
[0018] In some variations, the third adjustment corresponds to a sum or weighted sum of the first adjustment and the second adjustment.
[0019] In some variations, each sample sequence from the first data distribution may be encoded to generate an embedding for each sample sequence. A plurality of noisy sample sequences of the first training set may be generated by adding noise to at least the embedding for each sample sequence.
[0020] In some variations, each sample sequence from the first data distribution is encoded by augmenting it with additional information.
[0021] In some variations, the additional information includes structural information that identifies, for each constituent amino acid residue, one or more adjacent amino acid residues in three-dimensional space.
[0022] In some variations, the trained protein design computational model generates the output sequence by at least generating a noisy input sequence by adding noise to at least the input sequence, applying an energy-based model to generate a noisy output sequence by at least modifying the noisy input sequence based at least on an energy function of the energy-based model, and generating the output sequence by denoising at least the modified noisy output sequence generated by the energy-based model.
[0023] In some variations, the trained protein design computational model generates the output sequence by at least: generating an embedding of the input sequence by at least encoding the input sequence; generating a noisy embedding of the input sequence by at least adding noise to the embedding of the input sequence; applying an energy-based model to at least modify the noisy embedding of the input sequence based at least on an energy function of the energy-based model to generate a modified noisy embedding; denoising the noisy embedding to generate a denoised embedding; and generating the output sequence by at least denoising the noisy embedding.
[0024] In some variations, the embedding of the input sequence is generated by at least generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
[0025] In some variations, the embedding of the input sequence is generated by at least generating, for at least one amino acid residue of the input sequence, one or more structural tokens that identify one or more adjacent amino acid residues in three-dimensional space.
[0026] In some variations, the trained computational protein design model modifies the input sequence by at least one of (i) inserting amino acid residues, (ii) deleting amino acid residues, and (iii) changing the identity of amino acid residues in the input sequence.
[0027] In some variations, a fixed-length representation of the input sequence may be generated, and the trained protein design computational model may be applied to generate an output sequence by modifying at least the fixed-length representation of the input sequence.
[0028] In some variations, the fixed length representation of the input sequence is generated by at least aligning each amino acid residue in the input sequence to a fixed set of structural roles such that each amino acid residue in the input sequence is assigned an integer position corresponding to the amino acid residue's structural role, and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue with a corresponding structural role.
[0029] In some variations, the difference between the one or more generated output sequences and the plurality of noisy sample sequences is quantified by one or more of an antibody similarity metric, an edit distance, and a naturalness metric.
[0030] In another aspect, a system for generative protein design is provided, including at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include identifying an input sequence having a plurality of amino acid residues; generating a noisy embedding of the input sequence by at least adding noise to the input sequence; modifying the noisy embedding of the input sequence by applying a protein design computational model trained to at least approximate a data distribution of protein sequences exhibiting one or more desirable properties, where the protein design computational model modifies the noisy embedding of the input sequence to increase the likelihood that the resulting modified noisy embedding is present in the data distribution; and generating an output sequence by denoising at least the modified noisy embedding generated by the protein design computational model.
[0031] In another aspect, a method for generative protein design is provided, which may include identifying an input sequence having a plurality of amino acid residues, generating a noisy embedding of the input sequence by at least adding noise to the input sequence, modifying the noisy embedding of the input sequence by applying a protein design computational model trained to at least approximate a data distribution of protein sequences exhibiting one or more desirable properties, where the protein design computational model modifies the noisy embedding of the input sequence to increase the likelihood that the resulting modified noisy embedding is present in the data distribution, and generating an output sequence by denoising at least the modified noisy embedding generated by the protein design computational model.
[0032] In another aspect, a computer program product is provided that includes a non-transitory computer-readable medium storing instructions. The instructions can cause operations to be executed by at least one data processor. The operations can include identifying an input sequence having a plurality of amino acid residues; generating a noisy embedding of the input sequence by at least adding noise to the input sequence; modifying the noisy embedding of the input sequence by applying a protein design computational model trained to at least approximate a data distribution of protein sequences exhibiting one or more desirable properties, where the protein design computational model modifies the noisy embedding of the input sequence to increase the likelihood that the resulting modified noisy embedding is present in the data distribution; and generating an output sequence by denoising at least the modified noisy embedding generated by the protein design computational model.
[0033] In some variations, one or more of the features disclosed herein may be included, in any operable combination, including the following features:
[0034] In some variations, the input sequence is encoded to generate an embedding of the input sequence. The noisy embedding of the input sequence is generated by adding noise to at least the embedding of the input sequence. The output sequence is generated by decoding the denoised embedding produced by denoising the modified noisy embedding.
[0035] In some variations, the input sequence is encoded by generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
[0036] In some variations, the input sequence is encoded by at least generating one or more tokens that encode the relative position of each amino acid residue in the input sequence.
[0037] In some variations, the input sequence is encoded by generating, for at least one amino acid residue in the input sequence, one or more structural tokens that identify one or more adjacent amino acid residues in three-dimensional space.
[0038] In some variations, modifying the noisy embedding includes applying an energy-based model (EBM) trained to approximate the data distribution to modify the noisy embedding of the input sequence and generate a first modified noisy embedding; applying the energy-based model (EBM) to modify the noisy embedding of the input sequence and generate a second modified noisy embedding; applying an energy function parameterized by the energy-based model (EBM) to determine a first energy value of the first modified noisy embedding and a second energy value of the second modified noisy embedding; and applying the energy-based model (EBM) to further modify the first modified noisy embedding instead of the second modified noisy embedding based at least on the first energy value and the second energy value.
[0039] In some variations, an energy-based model (EBM) is applied to further modify the first modified noisy embedding until one or more criteria are met.
[0040] In some variations, the one or more criteria include at least one of (i) performing a threshold number of iterations of modifications to the noisy embedding of the input sequence, and (ii) a first energy value of the first modified noisy embedding that satisfies one or more thresholds.
[0041] In some variations, an energy-based model (EBM) is applied to further modify the first modified noisy embedding in place of the second modified noisy embedding based at least on the first energy value and the second energy value indicating that the first modified noisy embedding has a higher likelihood in the data distribution than the second modified noisy embedding.
[0042] In some variations, an energy-based model (EBM) is applied to further modify the first modified noisy embedding in place of the second modified noisy embedding based at least on the first energy value and the second energy value indicating that the first modified noisy embedding samples from a higher density region of the data distribution than the second modified noisy embedding.
[0043] In some variations, a fixed-length representation of the input sequence is generated. A noisy embedding of the input sequence is generated based on at least the fixed-length representation of the input sequence.
[0044] In some variations, the fixed length representation of the input sequence is generated by at least aligning each amino acid residue in the input sequence to a fixed set of structural roles such that each amino acid residue in the input sequence is assigned an integer position corresponding to the amino acid residue's structural role, and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue with a corresponding structural role.
[0045] In some variations, the protein design computational model modifies the noisy embedding of the input sequence by at least one of: changing the identity of one or more amino acid residues of the input sequence; deleting amino acid residues occupying positions in the fixed length representation of the input sequence by replacing the amino acid residues with gap characters; and inserting amino acid residues at positions in the fixed length representation of the input sequence by replacing at least the gap residues occupying the positions with the amino acid residues.
[0046] In some variations, the one or more desirable properties include at least one of expression, affinity, specificity, stability, non-immunogenicity, humanity, absence of self-association, and lack of chemical liability.
[0047] In some variations, the input sequence is a known protein sequence or a noise sequence comprising a random sequence of amino acid residues.
[0048] Implementations of the present subject matter may include, but are not limited to, methods according to the descriptions provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described that may include one or more processors and one or more memories coupled to the one or more processors. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, etc., one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods according to one or more implementations of the present subject matter may be implemented by one or more data processors in a single computing system or in multiple computing systems. Such multiple computing systems may be connected, for example, via one or more connections, including connections over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via a direct connection between one or more of the multiple computing systems, and may exchange data and / or instructions or other instructions, etc.
[0049] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the presently disclosed subject matter are described for illustrative purposes in connection with the computational design of protein molecules, including protein-based therapeutics such as antibodies, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure define the scope of the protected subject matter. [Brief explanation of the drawings]
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.
[0051] [Figure 1] FIG. 1 shows a system diagram illustrating an example of a protein design system, according to some exemplary embodiments.
[0052] [Figure 2A] 1 depicts a flowchart showing an example of a process for computational protein design, according to some exemplary embodiments.
[0053] [Figure 2B] 1 depicts a flowchart showing another example of a process for computational protein design, according to some exemplary embodiments.
[0054] [Figure 3A] 1 depicts a flowchart illustrating an example of a process for training a protein design computational model, according to some exemplary embodiments.
[0055] [Figure 3B] 1 depicts a flowchart illustrating another example of a process for training a protein design computational model, according to some exemplary embodiments.
[0056] [Figure 4A] 1 depicts a schematic diagram illustrating an example of sampling from a noisy data distribution, according to some exemplary embodiments;
[0057] [Figure 4B] 1 depicts a schematic diagram illustrating a comparison of density estimation of a clean data distribution without any noise disturbances and density estimation of a noisy data distribution, according to some exemplary embodiments;
[0058] [Figure 5A] 1 depicts a schematic diagram illustrating an example of sampling from a smoothed discrete space, according to some exemplary embodiments;
[0059] [Figure 5B] 1 depicts a block diagram illustrating an example of a discrete energy-based model (dEBM), in accordance with some demonstrative embodiments.
[0060] [Figure 6] 1 depicts a schematic diagram showing an example of sampling from a smoothed latent space, according to some exemplary embodiments;
[0061] [Figure 7A] 1 depicts a graph showing validation loss over successive training steps (or gradient updates) for different noise levels in the training data, in accordance with some exemplary embodiments.
[0062] [Figure 7B] 1 depicts a graph showing the similarity distribution of antibody heavy and light chains generated by a protein design computational model, according to some exemplary embodiments.
[0063] [Figure 7C] 1 depicts a graph showing the distribution of naturalities of antibody heavy and light chains generated by a protein design computational model, according to some exemplary embodiments.
[0064] [Figure 8] 1 depicts a schematic diagram showing distributional fitness score-based evaluation of in silico protein designs generated by a protein design computational model against a reference set of validation samples, according to some exemplary embodiments.
[0065] [Figure 9] 1 depicts a block diagram illustrating an example of a computing system, according to some illustrative embodiments.
[0066] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION
[0067] Computational protein design aims to generate protein sequences that exhibit various desirable properties. In the context of large molecule drug discovery (LMDD), for example, whether a protein sequence exhibits specific drug-like properties may determine its viability as a protein-based therapeutic, such as an antibody, enzyme, growth factor, hormone, interferon, interleukin, or thrombolytic agent. Therefore, in some cases, the drug development pipeline may involve evaluating candidate protein sequences for the presence of drug-like properties. For example, candidate protein sequences that successfully pass in vitro validation can then undergo preclinical development and clinical trials, in which the performance of the candidate protein sequence is tested in vivo. However, the significant consumption of wet-lab resources means that only a limited number of candidate protein sequences can proceed to in vitro and in vivo evaluation. Therefore, one important goal of computational protein design is to increase (or maximize) the likelihood that computationally generated candidate protein sequences exhibit the drug-like properties necessary for successful in vitro and in vivo testing. For example, as a protein-based therapeutic, the protein sequence can be computationally engineered to ensure that it exhibits sufficient expression, affinity and targeting, in vivo stability, pharmacokinetics, cell permeability, and non-immunogenicity.
[0068] Computational protein design is a formidable and resource-intensive task, at least in part because there are a large number of possible mutations in protein sequence and conformation (or three-dimensional structure), but only a small fraction of these variants have any therapeutic value. For example, a selection of 20 canonical amino acid residues is needed. TIFF2026508122000003.tif6170-Amount of amino acid residues formed Of the possible protein sequences in TIFF2026508122000004.tif6170, few possess the combination of drug-like properties (e.g., affinity, specificity, biological activity, and generativity) required for protein-based therapeutics. Therefore, increasing (or maximizing) the likelihood that computationally generated protein sequences submitted as candidates for in vitro and / or in vivo evaluation will exhibit drug-like properties is crucial. However, it may be necessary to evaluate at least some of the possible protein sequences in TIFF2026508122000005.tif6170 for the presence of drug-like properties. Evaluating any subset of possible protein sequences may inadvertently overlook at least some with superior drug-like properties. In contrast, a brute-force evaluation of all possible protein sequences, even when performed in silico, would be computationally too expensive to be a viable solution. Thus, in some exemplary embodiments, a protein design computational model may explore the vast combinatorial space of possible protein sequences in a principled manner to identify candidate protein sequences that are more likely to exhibit a desired combination of drug-like properties.
[0069] In some exemplary embodiments, a protein design computational model may generate one or more protein sequences by at least sampling a data distribution filled by known protein sequences (e.g., from the Protein Data Bank (PDB)) or a specific subset thereof (e.g., the observed antibody space (OAS)). In some cases, the known protein sequences (or a subset thereof) may exhibit one or more desirable properties, including, for example, drug-like properties such as expression, affinity, specificity, stability, non-immunogenicity, humanity, absence of self-association, lack of chemical liability, etc. Furthermore, in some cases, high-density regions of the data distribution may be filled by protein sequences similar to known protein sequences that exhibit one or more desirable properties, while low-density regions of the data distribution may be filled by protein sequences that are different from the known protein sequences. Thus, in some cases, one or more protein sequences may be generated by sampling from high-density regions of the data distribution, which are more likely to be filled by protein sequences similar to the known protein sequences. To that end, the protein design computational model may be trained to determine an energy function that approximates the data distribution. For example, in some cases, the data distribution, particularly the gradient of an energy function that approximates different densities across the data distribution, can be determined by Bayesian inference. The protein design computational model can then sample the data distribution based on the gradient of the energy function so that protein sequences are sampled from high-density regions of the data distribution instead of low-density regions of the data distribution, thus increasing the likelihood that the protein sequences will exhibit one or more desirable properties.
[0070] Training a protein design computational model to approximate a data distribution and efficiently sample from it to generate novel, unique, diverse, and therapeutically actionable protein sequences poses many unique challenges. First, data distributions of protein sequences are high-dimensional (e.g., length Protein sequence of TIFF2026508122000007.tif6170 While the data distribution is relatively large (dimensions 6170), there are disproportionately few known protein sequences available to train computational protein design models. Known protein sequences may identify some regions of high density within the data distribution, but the density of the regions in between remains unknown. Therefore, protein design computational models may be prone to overfitting if they are unable to generate protein sequences that are sufficiently diverse from known protein sequences. In this case, the phenomenon of overfitting may result from the protein design computational model learning an energy function based on available known protein sequences that fails to accurately capture the different densities of the data distribution between known protein sequences. For example, in some cases, the energy function may approximate a jagged energy landscape because, at least, the slope of the energy function exhibits sharp changes corresponding to the distinct differences in density that exist between regions filled by known protein sequences where the density of the data distribution is uncertain. The approximation of the energy function to a jagged energy landscape can prevent full exploration of the data distribution when a computational protein design model is subsequently applied to samples from the data distribution based on the gradient of the energy function. For example, protein sequences generated by sampling the data distribution can be repetitive and limited in diversity, at least because the gradient of the energy function restricts the computational protein design model to samples from the immediate vicinity around known protein sequences.
[0071] In some exemplary embodiments, an energy function that approximates the density across the data distribution of known protein sequences may be determined based on a noisy training set of sample sequences, each sample sequence being a known protein sequence contaminated with noise (e.g., isotropic Gaussian noise, etc.). That is, instead of training a protein design computational model to approximate the data distribution of known protein sequences directly based on the known protein sequences, the protein design computational model may be trained to approximate the data distribution of known protein sequences based on the noisy training set. Doing so may reduce the jaggedness of the energy landscape approximated by the energy function so that the protein design computational model can efficiently sample across the data distribution of known protein sequences. For example, in some cases, the protein design computational model may include at least one energy-based model (EBM). The protein design computational model is trained to approximate the data distribution of known protein sequences based on the noisy training set to determine an energy function that approximates different densities across the data distribution. It should be understood that the energy function may be parameterized by parameters of the energy-based model. For example, if the energy-based model (EBM) is implemented in an artificial neural network (e.g., a convolutional neural network), the parameters of the energy function may correspond to the weights and / or biases applied by the neurons of each successive layer of the artificial neural network. Training a protein design computational model to learn a data distribution may include adjusting the parameters of the energy-based model (EBM) to increase the similarity between protein sequences generated by the energy-based model (EBM) and sample sequences of a noisy training set. In doing so, the parameters of the energy function may also be adjusted so that the energy function outputs, for each protein sequence, an energy that indicates whether the protein sequence is within or outside the data distribution.
[0072] In some exemplary embodiments, training of a protein design computational model may include gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Markov chain Monte Carlo sampling with Langevin dynamics, etc.), in which parameters of the energy-based model (EBM) and parameters of a corresponding energy function are adjusted over successive sampling iterations to increase the similarity between protein sequences generated by the energy-based model (EBM) sampling from a data distribution and sample sequences in a noisy training set. For example, gradient-based Markov chain Monte Carlo may include modifying an input sequence, which may be a known protein sequence or a noise sequence (e.g., a sequence of random amino acid residues), and applying the energy-based model (EBM) to generate a first modified sequence, before applying the energy-based model (EBM) to further modify the first modified sequence to generate a second modified sequence. In some cases, the second modified sequence may be further modified and the energy-based model (EBM) may be applied again to generate a third modified sequence. The parameters of the energy-based model (EBM) may be adjusted so that the second modified sequence is more similar to the sample sequences of the noisy training set than the first modified sequence. In some cases, the parameters of the energy-based model (EBM) may be further modified so that the third modified sequence is more similar to the sample sequences of the noisy training set than the second modified sequence. As described above, adjusting the parameters of the energy-based model (EBM) also adjusts the parameters of the corresponding energy function. For example, in some cases, the parameters of the energy function may be continuously adjusted to reduce the energy values output by the energy function for protein sequences that fall within the data distribution of known protein sequences.
[0073] In some exemplary embodiments, to avoid overfitting the protein design computational model to known protein sequences, the protein design computational model may be trained based on a noisy training set of known protein sequences, rather than on the known protein sequences themselves. For example, in some cases, training the protein design computational model may involve gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling, etc.), in which the parameters of an energy-based model (EBM) are adjusted over successive sampling iterations to increase the similarity between the protein sequences sampled from the data distribution and the sample sequences in the noisy training set. The energy function derived in this manner based on the noisy training set may be parameterized to obtain a smoothed energy landscape, which mitigates the phenomenon of mode collapse, in which the energy-based model (EBM) is relatively non-robust and can only generate a limited selection of protein sequences (e.g., those in the immediate vicinity of the known protein sequences in the data distribution). As described in more detail below, during inference, in which a trained energy-based model (EBM) is applied to generate output sequences by sampling from a data distribution, the trained energy-based model (EBM) may "walk" the smoothed energy landscape of the noisy protein sequences (e.g., the noisy data distribution) and draw output sequences therefrom, e.g., via one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.), toward incrementally denser regions of the data distribution, before "jumping" to the true data distribution by denoising the noisy output sequences.
[0074] In some exemplary embodiments, a protein design computational model may manipulate noisy embeddings and inference of protein sequences during training. As described above, in some cases, an energy-based model (EBM) may learn and sample from a noise-contaminated data distribution of known protein sequences. The energy function of this data distribution may obtain a smoothed energy landscape with fewer sharp gradient changes, limiting the diversity of output protein sequences sampled from the data distribution. In some cases, each known protein sequence may be encoded to generate a corresponding sequence embedding before noise is added to each sequence embedding. For example, in some cases, an energy-based model (EBM) may be trained based on a noisy embedded training set of noisy embeddings of known protein sequences to learn the corresponding data distribution. During inference, the noisy sequence embeddings may be sampled from the data distribution before being denoised and decoded to generate output protein sequences. As described in more detail below, encoding of protein sequences may include the addition of information, such as structural and / or environmental information of the protein sequence, to increase the semantic meaning of the resulting sequence embedding. Sampling from a noisy latent space occupied by noisy sequence embeddings may result in output protein sequences that are more likely to exhibit desirable properties of known protein sequences (e.g., drug-like properties such as binding affinity and specificity, stability, non-immunogenicity, humanity, absence of self-association (or non-aggregation), lack of chemical burden (e.g., aspartate isomerization, oxidation, deamidation)).
[0075] In some exemplary embodiments, encoding a known protein sequence may project the known protein sequence from a sequence space (or discrete space) filled by the protein sequence into a latent space filled by sequence embeddings, each of which is a latent space representation of the corresponding protein sequence. A sequence embedding of a known protein sequence may have a different dimensionality, or amount of features, than the known protein sequence. For example, if the encoding augments the known protein sequence with information in addition to the identity and order of the constituent amino acid residues, the resulting sequence embedding may have a higher dimensionality (or a greater number of features) than the known protein sequence. One example of additional information included in a sequence embedding is structural information indicating the three-dimensional structure (or conformation) adopted by the known protein sequence. For example, in some cases, a sequence embedding of a known protein sequence may include, for each amino acid residue in the known protein sequence, one or more structural tokens that identify one or more neighboring amino acid residues in three-dimensional space (e.g., one or more nearest amino acid residues, one or more amino acid residues within a threshold distance, etc.). In this context, it should be understood that a sequence embedding of a protein sequence may include an array of tokens. In addition to the structural tokens mentioned above, some tokens may encode the identity and, in some cases, the sequential position of each amino acid residue in the protein sequence (e.g., one-hot encoding, etc.).
[0076] In some exemplary embodiments, a protein design computational model may generate output sequences by at least sampling from a latent space, which may be smoothed by adding noise to sequence embeddings therein. Thus, in some cases, a trained energy-based model (EBM) may "walk" the smoothed latent space over one or more iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo, etc.), denoising the noisy sequence embeddings by decoding them, and then extracting the noisy sequence embeddings from there before "jumping" to the true latent space by returning to the sequence space. The sequence (or primary structure) of a protein molecule alone may be insufficient to explain, at least, the presence (or absence) of certain desired properties (e.g., drug-like properties such as binding affinity and specificity) because these properties may also depend on the three-dimensional structure (e.g., secondary structure, tertiary structure, etc.) of the protein molecule. Enriching protein molecule sequences with additional information, such as the aforementioned structural tokens, can increase the semantic meaning of the resulting sequence embeddings by capturing at least some relationships between the protein molecule's sequence, conformation (or three-dimensional structure), and properties. For example, the distance between two or more sequence embeddings in the corresponding latent space can reflect the similarity (or dissimilarity) in protein sequence and conformation (or three-dimensional structure). Thus, the latent space also exhibits greater continuity than a more sparsely filled sequence space. Sampling from the latent space can therefore result in output protein sequences that are diverse and more likely to exhibit the required combination of desirable properties (e.g., drug-like properties).
[0077] In some exemplary embodiments, a protein design computational model may include multiple energy-based models (EBMs) trained in combination to learn different data distributions. For example, in some cases, a protein design computational model may include a first energy-based model (EBM) and a second energy-based model (EBM). In some cases, the first energy-based model may be trained to approximate a first data distribution of protein sequences, and the second energy-based model may be trained to approximate a second data distribution of protein sequences. Further, in some cases, the first energy-based model may be trained to approximate the first data distribution based on the gradient of an energy function associated with the second energy-based model that varies in amount across the second data distribution. For example, if a first data distribution may be associated with a more limited training set (e.g., an insufficient amount of known protein sequences) than a second data distribution, combining training of a first energy-based model (EBM) and a second energy-based model (EBM) in this manner may allow the first energy-based model to learn from a larger training set of the second data distribution while avoiding catastrophic forgetting, which may occur when the first energy-based model is trained on both training sets. For example, in some cases, the second data distribution may be associated with a larger set of known protein sequences (e.g., observed antibody space (OAS)), while the first data distribution may be associated with a smaller subset of known protein sequences that exhibit one or more desirable properties (e.g., antibodies that bind to a specific target molecule). If the subset of known protein sequences includes relatively few known protein sequences, in addition to adjusting parameters of the first energy-based model (EBM) to increase similarity between the protein sequences generated by the first energy-based model (EBM) and the known protein sequences from the first data distribution, the parameters of the first energy-based model may be adjusted based on the gradient of an energy function associated with the second energy-based model.This energy function, which provides a density estimate of the second data distribution of protein sequences, may complement the training of the first energy-based model by providing a surrogate density estimate for at least some of the regions of the first data distribution without proper characterization by known protein sequences.
[0078] In some exemplary embodiments, an energy-based model (EBM) may be trained to generate output sequences with one or more desired properties by applying at least one or more modifications to an input sequence. In some cases, the energy-based model (EBM) may be trained to learn the data distribution of a training set so that modifications made to the input sequence match the pattern of amino acid residues observed in the sample sequences. Examples of modifications that may be made to the input sequence may include changing the identity of one or more amino acid residues in the input sequence and changing the length of the input sequence by inserting and / or deleting one or more amino acid residues. It should be understood that the length of the input sequence may change frequently throughout the generation process because one or more amino acid residues may be inserted and / or deleted during each iteration of gradient-based Markov Monte Carlo (MCMC) sampling. A traditional variable-length representation of the input sequence may require the energy-based model (EBM) to adjust to accommodate each length change, increasing the computational burden of the generation process. Thus, in some cases, the computational complexity resulting from the varying length of the input sequence during the generation process may be reduced by an energy-based model that operates on a fixed-length representation of the input sequence instead of a traditional variable-length representation of the input sequence. For example, the protein design engine may generate a fixed-length representation of the input sequence before generating the corresponding noisy sequence, or optionally before generating the corresponding noisy sequence embedding. In some cases, each amino acid residue in the input sequence corresponds to a structural role of the amino acid residue (e.g., An input sequence may be rendered in a fixed-length representation by applying a structural-role-based numbering scheme in which amino acid residues (selected from a range of integers, such as TIFF2026508122000009.tif5170) are assigned to integer positions of a fixed-length sequence. Gaps at any position in the fixed-length sequence where the input sequence lacks an amino acid residue with a corresponding structural role may be represented by a gap character, such that each position in the fixed-length representation of the input sequence may be occupied by either an amino acid residue (e.g., one of the 20 canonical amino acid residues) or a gap character. Furthermore, amino acid residues may be inserted into the input sequence by replacing tokens encoding gap characters in the fixed-length representation of the input sequence with tokens encoding the identity of the amino acid residue, while amino acid residues may be deleted from the input sequence by replacing tokens encoding the identity of amino acid residues in the fixed-length representation of the input sequence with tokens encoding gap characters.
[0079] FIG. 1 shows a system diagram illustrating an example of a protein design system 100, according to some illustrative embodiments. With reference to FIG. 1, protein design system 100 can include a protein design engine 110, an analysis engine 120, and a client device 130. As shown in FIG. 1, protein design engine 110, analysis engine 120, and client device 130 can be communicatively coupled via a network 140. Client device 130 can be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. Network 140 can be a wired and / or wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.
[0080] In some exemplary embodiments, protein design engine 110 may include encoder 111, noise engine 113, protein design computational model 115, denoising engine 117, and decoder 119. In some cases, protein design engine 110 may apply protein design computational model 115 to generate output sequence 162 based on at least input sequence 152. In the example shown in FIG. 1 , protein design computational model 115 may include one or more energy-based models 170, including, for example, first energy-based model 170a, second energy-based model 170b, etc. In some cases, each of the one or more energy-based models 170 may be trained to approximate a corresponding data distribution. For example, in some cases, first energy-based model 170a and first energy-based model 170a's parameters (e.g., weights, biases, etc.) may approximate a first data distribution of protein sequences. The second energy-based model 170b and the second energy function 175b parameterized by the parameters (e.g., weights, biases, etc.) of the second energy-based model 170b may approximate a second data distribution of protein sequences. If an insufficient amount of known protein sequences characterizing the first data distribution is available to train the first energy-based model 170a, the first energy-based model 170a may be trained to approximate the first data distribution based on a first gradient of the first energy function 175a and a second gradient of the second energy function 175b. For example, as described in more detail below, the first energy-based model 170a may be trained to approximate a first data distribution by applying one or more adjustments to its parameters (e.g., weights, biases, etc.) of the first energy-based model 170a that increase (or maximize) a first similarity between a first output of the first energy-based model 170a and a sample sequence from the first data distribution, and a second similarity between a second output of the second energy-based model 170b and a sample sequence from the second data distribution.
[0081] In some exemplary embodiments, the protein design computational model 115 may generate the output sequence 162 by at least modifying the input sequence 152 and applying the first energy-based model 170a to increase the likelihood that the output sequence 162 is in the first data distribution of protein sequences. In examples in which the first data distribution of protein sequences exhibits one or more desirable properties (e.g., drug-like properties such as affinity, specificity, biological activity, generability, etc.), the output sequence 162 may be generated to also exhibit the one or more desirable properties. As described in more detail below, the protein design computational model 115 may modify the input sequence 152 over one or more successive iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Markov chain Monte Carlo (MCMC) with Langevin dynamics, etc.). For example, in some cases, each iteration of the gradient-based Markov chain Monte Carlo (MCMC) sampling may include drawing samples from the first data distribution that include one or more modifications to the input sequence 152.
[0082] In some cases, sampling from the first data distribution may be guided by the first energy function 175a. For example, samples extracted from low-density regions of the first data distribution populated by protein sequences lacking one or more desirable characteristics may be assigned a high energy value by the first energy function 175a, indicating a low likelihood of their presence in the first data distribution. In contrast, samples extracted from high-density regions of the first data distribution populated by protein sequences with one or more desirable characteristics may be assigned a low energy value by the first energy function 175a, indicating a high likelihood of their presence in the first data distribution. Thus, the slope of the first energy function 175a, which corresponds to changes in energy values, may approximate changes in density across the first data distribution. Sampling from the first data distribution based on the slope of the first energy function 175a may include plotting samples based on changes in the energy values assigned to each sample by the first energy function 175a. For example, in each subsequent iteration of the gradient-based Markov chain Monte Carlo (MCMC), samples may be drawn from increasingly dense regions of the first data distribution populated with protein sequences exhibiting one or more desirable properties. The first energy function 175a may assign lower energy values to those samples to indicate their higher likelihood of being in the first data distribution. It should be understood that in some cases, each iteration of the gradient-based Markov chain Monte Carlo (MCMC) may include further modifying one or more samples from the previous iteration that were determined to have lower energy values than other samples from that previous iteration.
[0083] In some demonstrative embodiments, instead of first energy-based model 170a parameterizing first energy function 175a, first energy-based model 170a may parameterize a score function that outputs, for each sample seen from the first data distribution, a score corresponding to the change in density observed at the sample's location. In some cases, the score function may approximate the gradient of first energy function 175a, which approximates the change in density across the first data distribution, as described above. Thus, in some cases, sampling from the first data distribution may be guided by the score function (instead of first energy function 175a), such that each successive sample is drawn from an incrementally higher density region of the first data distribution. For example, the score function may assign a first score to a first sample that exhibits a more positive local change (e.g., an increase or less decrease) in the density of the first data distribution at a first location of the first sample, and assign a second score to a second sample that exhibits a less positive local change (e.g., a less increase or decrease) in the density of the first data distribution at a second location of the second sample. In some cases, the first energy-based model 170a may derive a third sample from the first data distribution by further modifying the first sample to sample the third sample from a denser region of the first data distribution than the first and second samples.
[0084] In some exemplary embodiments, overfitting of the protein design computational model 115 to sample sequences in a training set may be avoided by the protein design computational model 115 operating on noisy protein sequences. For example, in some cases, the protein design computational model 115 may be trained based on a noisy training set of sample sequences, each of which is a known protein sequence contaminated with noise (e.g., Gaussian noise, such as isotropic Gaussian noise). Furthermore, the protein design computational model 115 may generate the output sequence 162 by applying the first energy-based model 170a to modify the noisy embedding 156 of the input sequence 152. For example, as shown in FIG. 1 , the noise engine 113 may generate the noisy embedding 156 by adding at least noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to the input sequence 152. The protein design computational model 115 may modify the noisy embedding 156 of the input sequence 152, thus applying the first energy-based model 170a to generate the modified noisy embedding 158. In some cases, the first energy-based model 170a may be applied to modify some portions of the noisy embedding 156 but not other portions. For example, in some cases, the modification of the noisy embedding 156 may be limited to one or more adjustable segments of the input sequence 152, thus avoiding altering one or more fixed segments of the input sequence 152. Additionally, the denoising engine 117 may remove noise present in the modified noisy embedding 158 to generate a denoised embedding 160 before the output sequence 162 is generated therefrom. As described in more detail below, the first energy-based model 170a may generate the modified noisy embedding 156 by “walking” the smoothed energy landscape of the noisy protein sequence before the denoising engine 117 denoises the modified noisy embedding 156, thus “jumping” back to the true data distribution.
[0085] In some exemplary embodiments, the noisy embedding 156 of the input sequence 152 may be generated by the noise engine 113 adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to the embedding 154 of the input sequence 152. In some cases, the embedding 154 may include additional information associated with the input sequence 152. For example, in some cases, the embedding 154 may be generated by the encoder 111 by augmenting (or upsampling) the input sequence 152 with structural information in the form of one or more structural tokens, each of which identifies the nearest neighbor amino acid residue in three-dimensional space for each amino acid residue in the input sequence 152. In doing so, the encoder 111 may map the input sequence 152 from a sparsely filled sequence space to a more continuous and semantically meaningful latent space from which the protein design computational model 115 samples the modified noisy embedding 158. For example, the latent space may better capture the relationship between protein sequence, conformation (or three-dimensional structure), and properties, and the distance between two or more sequence embeddings in the latent space reflects the similarity (or dissimilarity) in protein sequences as well as conformation (or three-dimensional structure). Referring again to FIG. 1 , in the example shown, the modified noisy embedding 158 generated by sampling from the latent space may be denoised by the denoising engine 117 before the resulting denoised embedding 160 is decoded by the decoder 119 or mapped from the latent space to the sequence space. The output sequence 162 thus generated may be more likely to exhibit one or more desirable characteristics of the protein sequence in the first data distribution. However, it should be understood that instead of the encoder 111 generating the embedding 154 by augmenting the input sequence 152 with additional information, the encoder 111 may implement an identity function, in which case the embedding 154 may be generated without augmenting the input sequence 152 with any additional information.
[0086] In some exemplary embodiments, the protein design computational model 115 may generate a modified noisy embedding 158 by applying the first energy-based model 175a to at least modify the noisy embedding 156 of the input sequence 152. Examples of modifications include changing the identity of one or more amino acid residues in the input sequence 152 and modifying the length of the input sequence 152 through the insertion and / or deletion (or removal) of one or more amino acid residues. If the first energy-based model 175a operates on a variable-length representation of the input sequence 152, the first energy-based model 175a may require adjustment to accommodate changes in the length of the input sequence 152, which may occur frequently throughout the generation process when one or more amino acid residues are inserted and / or deleted (or removed) between each iteration of gradient-based Markov Monte Carlo (MCMC) sampling. To avoid the computational complexity imposed by the variable-length modification of the input sequence 152, the first energy-based model 175a may operate on a fixed-length representation of the input sequence 152 instead of a variable-length representation of the input sequence 152. For example, if the input sequence 152 is a sequence of fixed length sequences in which each amino acid residue in the input sequence 152 is an integer position (e.g., When rendered in a fixed-length representation by applying a numbering scheme based on structural roles assigned to amino acid residues (e.g., selected from a range of integers, such as TIFF2026508122000010.tif5170), the embedding 154 and noisy embedding 156 of input sequence 152 may have the same length (e.g., the same amount of tokens), regardless of the amount of amino acid residues forming the input sequence 152. As explained in more detail below, if the input sequence 152 corresponds to an immunoglobulin protein (or antibody), the aforementioned structural roles may correspond to amino acid residues occupying a particular complementarity-determining region (CDR) loop or one of the framework regions between a pair of complementarity-determining region (CDR) loops. At any position in the embedding 154 and noisy embedding 156 where the input sequence 152 lacks an amino acid residue with a corresponding structural role, the embedding 154 and noisy embedding 156 of the input sequence 152 may include a gap character to indicate the absence of such amino acid residue.
[0087] Referring again to FIG. 1 , protein design computational model 115 may take a noisy embedding 156 of input sequence 152 and generate a modified noisy embedding 158. As described in more detail below, denoising engine 117 may denoise modified noisy embedding 158 to generate a denoised embedding 160 before decoder 119 decodes denoised embedding 119 to generate output sequence 162. As mentioned above, in some cases, protein design computational model 115 may apply a first energy-based model 175a, which may generate modified noisy embedding 158 by modifying input sequence 152. In some cases, first energy-based model 175a may modify input sequence 152 by changing the identity of one or more amino acid residues in input sequence 152. Alternatively and / or additionally, first energy-based model 175a may modify input sequence 152 by inserting and / or deleting (or removing) one or more amino acid residues from input sequence 152. When input sequence 152 is rendered in the aforementioned fixed-length representation, insertion of a particular type of amino acid residue at a particular position may be achieved by replacing a corresponding gap character in noisy embedding 156 of input sequence 152 with an amino acid residue of that type. Alternatively, deletion (or removal) of an amino acid residue occupying a particular position in input sequence 152 may be achieved by replacing the amino acid residue in noisy embedding 156 of input sequence 152 with a gap character.
[0088] As described above, in some exemplary embodiments, the protein design computational model 115 may operate on noisy protein sequences, which are protein sequences contaminated with noise (e.g., Gaussian noise, etc.). In the example shown in FIG. 1 , the protein design computational model 115 may apply the first energy-based model 170a to modify the noisy embedding 156 of the input sequence 152 and generate a modified noisy embedding 158. In some cases, the noisy embedding 156 may be generated by the noise engine 113 adding noise (e.g., Gaussian noise, etc.) to the embedding 154 of the input sequence 152. It should be understood that the embedding 154 may be enhanced with additional information (e.g., structural information, etc.), although there may also be instances in which the embedding 154 excludes the additional information. Alternatively, the encoder 111 may implement an identity function, meaning that the embedding 154 may capture the same information present in the input sequence 152, including, for example, the identity of each amino acid residue, the sequential position of each amino acid residue, etc. To further explain, Figure 2A depicts a flowchart illustrating an example of a process 200 for computational protein design, according to some exemplary embodiments. With reference to Figures 1 and 2A, process 200 may be performed by protein design engine 110 to train and apply protein design computational model 115 to generate output sequence 162 by modifying at least noisy embedding 156 of input sequence 152. As described in more detail below, noisy embedding 156 of input sequence 152 may be generated based on embedding 154 of input sequence 152, which may or may not be augmented with additional information associated with input sequence 152 (e.g., structural information, etc.).
[0089] At 202, the protein design engine 110 may generate a noisy training set including a plurality of noisy sample sequences. In some exemplary embodiments, to train the protein design computational model 115 on a data distribution populated with specific protein sequences, such as protein sequences exhibiting one or more desirable properties (e.g., drug-like properties), the protein design engine 110 may generate a noisy training set including a plurality of noisy sample sequences. Each noisy sample sequence in this case may be generated by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to a known protein sequence from the data distribution. Furthermore, in some cases, each noisy sample sequence may be a noisy sequence embedding generated by adding noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to an embedding of a known protein sequence, which may or may not be augmented with additional information (e.g., structural information). Training the protein design computational model 115 based on the noisy sample sequences of the noisy training set typically requires the protein design computational model 115 to train a high-dimensional data distribution (e.g., length 100 or 200 kb) based on a disproportionately small number of known protein sequences. Protein sequence of TIFF2026508122000011.tif6170 This may mitigate overfitting and mode collapse that occurs when the image is trained to approximate a given size (dimensions of TIFF2026508122000012.tif6170).
[0090] In the example of the protein design computational model 115 shown in FIG. 1 , the protein design computational model 115 may include a first energy-based model 170a. In this case, training the protein design computational model 115 may include training the first energy-based model 170a to approximate a data distribution based on noisy sample sequences of a noisy training set. In some cases, the first energy-based model 170a may be a machine learning model such as an artificial neural network (ANN), in which case training the first energy-based model 170a may include adjusting one or more parameters (e.g., weights, biases, etc.) of the machine learning model. Doing so may also determine a first energy function 175a parameterized by the parameters of the first energy-based model 170a to output an energy value corresponding to the likelihood of the protein sequence within the first data distribution. In some cases, a noisy training set may be applied to train the first energy-based model 175a, e.g., to avoid overfitting the first energy-based model 175a to some known protein sequences available to characterize the first data distribution. Known protein sequences in TIFF2026508122000013.tif5170 TIFF2026508122000014.tif4170 is It can be converted into a noisy sample array by TIFF2026508122000015.tif5170.
[0091] In some cases, the noise level The noise level may be determined based on the dimensionality and / or sparsity of the first data distribution to increase (or maximize) the quality of the noisy sample sequences in the noisy training set. TIFF2026508122000017.tif3170 may be set to approximately 0.5 (or another value). To further illustrate, consider an entry defined as Matrix with TIFF2026508122000018.tif4170 Consider TIFF2026508122000019.tif4170.
number
[0092] In some exemplary embodiments, the protein design engine 110 may generate each noisy sample sequence of the noisy training set by adding noise to the corresponding embedding of the sample sequence. For example, in some cases, the noise engine 113 may generate the noisy sample sequence based on the known protein sequence by adding at least noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to the embedding of the known protein sequence. It should be understood that the embedding of the known protein sequence may or may not be augmented with additional information (e.g., structural information, etc.), e.g., by the encoder 111. For example, if additional information is present, the embedding of the known protein sequence may include tokens that encode the identity, sequential position, and / or structural information of one or more amino acid residues of the known protein sequence.
[0093] In some cases, each noisy sample sequence in the noisy training set may be fixed in length, meaning that each noisy sample sequence may be the same length (e.g., the same amount of tokens) regardless of the amount of amino acid residue in the corresponding known protein sequence. For example, in some cases, the encoder 111 of the protein design engine 110 may assign to each amino acid residue in the known protein sequence a sequence number corresponding to the structural role of the amino acid residue (e.g., An embedding of a known protein sequence may be generated by applying a structural-role-based numbering scheme, which involves assigning integer positions of a fixed-length sequence (selected from a range of integers, such as TIFF2026508122000027.tif5170) to amino acid residues of the known protein sequence. To keep the length of the embedding the same regardless of the actual amount of amino acid residues in the known protein sequence, encoder 111 may insert gap characters at any positions in the fixed-length sequence where the known protein sequence lacks an amino acid residue with a corresponding structural role. As noted above, in some cases, encoder 111 may further generate the embedding to include encoding the identity of each amino acid residue in the known protein sequence (e.g., one-hot encoding), and in some cases, encoding the consecutive positions of each amino acid residue. To further illustrate, a known protein sequence may be generated by applying a structural-role-based numbering scheme, which involves assigning integer positions of a fixed-length sequence (selected from a range of integers, such as TIFF2026508122000027.tif5170) to amino acid residues of the known protein sequence. TIFF2026508122000028.tif5170, and each token TIFF2026508122000029.tif5170 indicates either the type of amino acid residue or the position gap character TIFF2026508122000030.tif4170. In some cases, noise (e.g., Gaussian noise such as isotropic Gaussian noise) is embedded. TIFF2026508122000031.tif3170 to generate a noisy embedding that is a numeric floating point representation of the original known protein sequence.
[0094] At 204, the protein design engine 110 may train the protein design computational model 115 by applying at least the protein design computational model 115 to generate one or more output sequences and adjusting the protein design computational model 115 to reduce the difference between the one or more generated output sequences and a plurality of noisy sample sequences of the noisy training set. In some exemplary embodiments, training the protein design computational model 115 may include training a first energy-based model 170a based on at least the noisy training set to approximate a data distribution populated by particular protein sequences, such as protein sequences exhibiting one or more desirable properties. In some cases, training the first energy-based model 170a may further include determining a corresponding first energy function 175a, which is parameterized by parameters (e.g., weights, biases, etc.) of the first energy-based model 170a. For example, in some cases, training the first energy-based model 170a may include adjusting one or more parameters (e.g., weights, biases, etc.) of the first energy-based model 170a to reduce (or minimize) the difference between the output sequences generated by the first energy-based model 170a and the noisy sample sequences of the noisy training set. In doing so, the parameters of the first energy function 175a may also be adjusted so that the first energy function 175a outputs lower energy values for a first protein sequence that is within the data distribution than for a second protein sequence that is outside the data distribution.
[0095] In some exemplary embodiments, the protein design engine 110 may train the first energy-based model 170a by performing at least gradient-based Markov chain Monte Carlo (MCMC) sampling, such as Markov chain Monte Carlo (MCMC) sampling with Langevin dynamics, to approximate the gradient of the first energy function 175a. In some cases, the gradient of the first energy function 175a may indicate changes in the density of the data distribution. For example, the gradient of the first energy function 175a may indicate transitions between different density regions of the data distribution, including, for example, transitions between high-density and low-density regions of the data distribution. As described in more detail below, subsequent sampling from the data distribution may be guided by the first energy function 175a, and in particular the gradient of the first energy function 175a, toward high-density regions of the data distribution that are more likely to be filled by protein sequences exhibiting one or more desirable properties. Further, in some cases, gradient-based Markov Chain Monte Carlo (MCMC) sampling to approximate the gradient of first energy function 175a may include adjusting first energy function 170a and parameters of first energy function 175a (e.g., weights, biases, etc.) over successive iterations to increase (or maximize) the similarity between output sequences produced by first energy-based model 170a and noisy sample sequences of the noisy training set while reducing (or minimizing) the energy values determined by first energy function 175a for those sequences.
[0096] To further illustrate, the training of the first energy-based model 170a may be performed using: TIFF2026508122000032.tif5170 input TIFF2026508122000033.tif3170 may include training a first energy function 170a, denoted as mapping the input to a scalar "energy" value. Data distribution associated with TIFF2026508122000034.tif3170 TIFF2026508122000035.tif5170 can be approximated by the following Boltzmann distribution:
number
[0097] In some cases, the first energy-based model 170a is generated by Markov Chain Monte Carlo (MCMC) sampling. It can be trained via contrastive divergence using new sequences (or "samples") drawn from TIFF2026508122000037.tif5170. In the case of gradient-based Markov chain Monte Carlo sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling), each sequence (or "sample") can be initialized from a known protein sequence or a noise sequence before being refined with (discretized) Langevin diffusion.
number
[0098] According to the foregoing formulation, training the first energy-based model 170a may include adjusting the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a to increase (or maximize) the log-likelihood of the noisy sample sequences under the model. That is, the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a may be adjusted to increase the likelihood of the first energy-based model 170a generating output sequences that are similar to the noisy sample sequences of the noisy training set. To this end, the parameters of the first energy-based model 170a may be adjusted to increase the likelihood of the first energy-based model 170a generating output sequences that are similar to the noisy sample sequences of the noisy training set. Noisy training set with increasing energy of noisy data sampled from TIFF2026508122000044.tif5170 That is, when trained, the first energy function 175a of the first energy-based model 170a may output lower energy values for a first protein sequence that is within the data distribution (or sampled from a higher density region of the data distribution) than for a second protein sequence that is outside the data distribution (or sampled from a lower density region of the data distribution). To normalize the energy, an additional TIFF2026508122000046.tif5170 norm penalty may be added.
number
[0099] As mentioned above, the first energy-based model 170a may be trained based on noisy sample sequences of a noisy training set to avoid overfitting the first energy-based model 170a to the few known protein sequences that characterize the data distribution. Protein sequence of TIFF2026508122000048.tif6170 When few known protein sequences are available to characterize a data distribution (dimensions of TIFF2026508122000049.tif6170), training the first energy-based model 170a directly based on the known protein sequences may result in a jagged energy landscape in which there are dramatic changes in energy values between regions populated by known protein sequences. Sampling from the data distribution based on the gradient of the jagged energy landscape may hinder adequate exploration of the data distribution, at least because the steepness of the gradient may limit sampling to regions immediately adjacent to the known protein sequences. In contrast, training the first energy-based model 170a based on noisy sample sequences may result in a smoothed energy landscape in which the gradient of the first energy function 175a is more gradual, allowing for better exploration of the data distribution when sampling from it.
[0100] In some cases, known protein sequences TIFF2026508122000050.tif4170 is converted with additive noise (e.g., Gaussian noise) to create a noisy sample array If TIFF2026508122000051.tif5170 is obtained, the known protein sequence A least squares estimate of TIFF2026508122000052.tif4170 can be obtained by:
number
number
[0101] At 206, the protein design engine 110 may apply the trained protein design computational model 115 to generate an output sequence by modifying at least the input sequence. In some exemplary embodiments, the protein design computational model 115 may apply the first energy-based model 170a to generate the output sequence 162 by modifying at least the input sequence 152 while guided by the first energy function 175a. For example, in some cases, the first energy-based model 170a may modify the noisy embedding 156 of the input sequence 152, which may be generated by the noise engine 113 that adds noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to the embedding 154 of the input sequence 152. Furthermore, in some cases, the embedding 154 may be generated by the encoder 111 to include additional information, such as structural information, associated with the input sequence 152.
[0102] In some cases, first energy-based model 170a may modify input sequence 152 by inserting, deleting (or removing), and / or changing the identity of one or more amino acid residues in input sequence 152. When input sequence 152 is rendered in a fixed-length representation, e.g., by applying a numbering scheme based on structural role, deletion (or removal) of an amino acid residue may be achieved by replacing a token encoding the identity of the amino acid residue with a gap character, whereas insertion of an amino acid residue may be achieved by replacing a gap character with a token encoding the identity of the amino acid residue. The modification of input sequence 152 may be guided by first energy function 175a (e.g., the gradient of first energy function 175a). Notably, in some cases, input sequence 152 may be subjected to successive iterations of modification, each lowering the energy value of input sequence 152.
[0103] For example, in some cases, the input sequence 152 may undergo a first modification and a second modification. Doing so may be equivalent to extracting a first sample and a second sample from the data distribution. In some cases, after extracting the first sample and the second sample from the data distribution, the protein design computational model 115 may apply a first energy function 175a to determine an energy value indicating the likelihood of each sample within the data distribution. A lower energy value in this case may indicate that the sample is drawn from a higher density region of the data distribution, or similarly, that the sample is more likely to be within the data distribution. Thus, in some cases, after extracting the first sample and the second sample, the protein design computational model 115 may apply a first energy-based model 170a to modify the input sequence 152, continuing to extract additional samples from increasingly dense regions of the data distribution, for example, until a sample exhibiting a threshold likelihood of being within the data distribution is extracted. For example, in some cases, the first energy-based model 170a may be applied to further modify the input sequence 152 with the first modification instead of the second modification if the input sequence 152 with the first modification is assigned a lower energy value by the first energy function 175a. Doing so may be similar to "walking" the energy landscape of the data distribution, sampling from incrementally denser regions of the data distribution. In an example where the first energy-based model 170a is modifying a noisy embedding 156 of the input sequence 152, the first energy-based model 170a may operate in a noisy latent space in which the distance between two or more sequence embeddings reflects the similarity (or dissimilarity) in protein sequences as well as the conformation (or three-dimensional structure). The energy landscape of the data distribution may be smoothed by the addition of noise, which reduces abrupt changes in the slope of the first energy function 175a.Because the first energy-based model 170a is trained to approximate the data distribution of protein sequences that exhibit particular desirable properties (e.g., drug-like properties), modifications made to the input sequence 152 may match patterns of amino acid residues observed in known protein sequences, such that the same desirable properties are also present in the output sequence 162 generated therefrom.
[0104] 2B shows a flowchart illustrating another example of a process 250 for protein design, according to some exemplary embodiments. With reference to FIGS. 1 and 2A-2B, process 250 may be performed, for example, by a protein design computational model 115 applied by protein design engine 110 to generate an output sequence based on an input sequence. In some cases, process 250 may perform operation 206 of process 200 shown in FIG. 2A.
[0105] At 252, the protein design engine 110 may encode the input sequence to generate an embedding of the input sequence. In some exemplary embodiments, the encoder 111 may encode the input sequence 152 to generate an embedding 154 of the input sequence 152. The input sequence 152 may correspond to a known protein sequence or a noise sequence (e.g., a sequence of random amino acid residues). In some cases, the encoder 111 may encode the input sequence 152 by generating, for each amino acid residue in the input sequence 152, at least a token that encodes the identity of each amino acid residue. If the input sequence 152 is rendered in a fixed-length representation having the same amount of tokens regardless of the amount of amino acid residues in the input sequence 152, at least some of the tokens in the embedding 154 of the input sequence 152 may identify the type of amino acid residue or gap character occupying the corresponding position in the embedding 154 of the input sequence 152. For example, in some cases, the encoder 111 may generate the fixed-length representation of the input sequence 152 by applying a numbering scheme based on structural roles. Doing so may involve aligning the amino acid residues forming the input sequence 152 to a fixed set of structural roles (e.g., corresponding to various complementarity determining region (CDR) loops or framework regions therebetween) and inserting gap characters where the alignment indicates the absence of amino acid residues with a particular structural role. Thus, If there are possible structural roles for the amount of TIFF2026508122000060.tif4170, the resulting embedding 154 is a sequence of tokens TIFF2026508122000061.tif5170, where each token TIFF2026508122000062.tif5170 indicates the type or position of amino acid residues TIFF2026508122000063.tif4170. In addition, each token TIFF2026508122000064.tif5170 is the position of embedded 154 relative to other tokens It can be generated to include positional encoding to indicate the sequential position of the tokens in TIFF2026508122000065.tif4170.
[0106] In some exemplary embodiments, the encoder 111 may generate an embedding 154 of the input sequence 152, with or without augmenting the embedding 154 with additional information. In some cases, the encoder 111 may implement an identity function, meaning that the embedding 154 may include the same information present in the input sequence 152, including, for example, the identity of each amino acid residue, the sequential position of each amino acid residue, etc. Alternatively, in examples in which the embedding 154 is generated to include additional information, this addition may include, for example, structural information, environmental information, etc. The addition of information may amount to mapping the input sequence 152 from a sequence space (or discrete space) filled by protein sequences to a continuous latent space filled by sequence embeddings, each of which is a latent space representation of a corresponding protein sequence. For example, in some cases, the encoder 111 may generate an embedding 154 of the input sequence 152 to include one or more structural tokens. In some cases, the one or more structural tokens may describe a conformation (or three-dimensional structure) adopted by the input sequence 152. For example, in some cases, a structural token may identify one or more nearest neighboring amino acid residues in three-dimensional space for a corresponding amino acid residue in the input sequence 152 .
[0107] It should be understood that these structural tokens convey a different type of information than positional encoding. That is, instead of amino acid residues adjacent to the primary structure of the input sequence 152, the structural tokens identify amino acid residues adjacent to the folding of the input sequence 152. The existing structural information may increase the semantic meaning of the embedding 154. In cases where the properties of the input sequence 152 depend on the conformation (or three-dimensional structure) adopted by the input sequence 152, incorporating structural information may improve the results of the subsequent generation process, at least because the protein design computational model 115 can take into account at least some of the relationships that exist between the sequence, conformation (or three-dimensional structure), and properties of the input sequence 152.
[0108] At 254, protein design engine 110 may add noise to the embedding of the input sequence to generate a noisy embedding of the input sequence. In some exemplary embodiments, noise engine 113 of protein design engine 110 may generate a noisy embedding 156 for ingestion by protein design computational model 115 (e.g., first energy-based model 170a) based at least on embedding 154. For example, in some cases, noise engine 113 may add noise (e.g., Gaussian noise, such as isotropic Gaussian noise) to embedding 154 to generate noisy embedding 156 of input sequence 152. As described above, in some cases, protein design computational model 115 (e.g., first energy-based model 170a) may be trained to approximate the noisy data distribution of protein sequences having desired properties based on a noisy training set of noisy sample sequences generated from known protein sequences exhibiting particular desired properties. This noisy data distribution may exhibit a smooth energy landscape with gradual gradient changes, facilitating subsequent sampling (e.g., gradient-based Markov Chain Monte Carlo (MCMC) sampling, etc.). In contrast, if the protein design computational model 115 (e.g., first energy-based model 170a) is trained directly based on known protein sequences, the protein design computational model 115 (e.g., first energy-based model 170a) may learn a jagged energy landscape in which dramatic changes in energy values exist between regions populated by known protein sequences. Unlike the gradual gradient of the noisy data distribution, the steep gradient of this jagged energy landscape may hinder proper exploration of the data distribution during the generation process, at least because sampling may be limited to regions immediately surrounding the known protein sequences.As described in more detail below, by operating on the noisy embedding 156 of the input sequence 152, the protein design computational model 115 (e.g., the first energy-based model 170a) can "walk" the smoothed energy landscape of the noisy data distribution to take samples from incrementally denser regions of the noisy data distribution before "jumping" to the true data distribution when a sample is drawn that exhibits a threshold likelihood of being within the noisy data distribution.
[0109] At 256, the protein design computational model 115 may apply an energy-based model (EBM) to generate a modified noisy embedding of the input sequence by at least modifying the noisy embedding of the input sequence based at least on the corresponding energy function. In some exemplary embodiments, the protein design computational model 115 may apply a first energy-based model 170a to modify the noisy embedding 156 of the input sequence and generate a modified noisy embedding 158. In some cases, the first energy-based model 170a may modify the noisy embedding 158 of the input sequence 152 by inserting amino acid residues, deleting (or removing) amino acid residues, and / or changing the identities of amino acid residues in the input sequence 152. As described above, insertion or deletion (or removal) of amino acid residues at specific positions in the input sequence 152 may be achieved without changing the length of the input sequence 152 by swapping tokens representing gap characters.
[0110] In some exemplary embodiments, the protein design computational model 115 may apply the first energy-based model 170a to modify the noisy embedding 158 of the input sequence 152 based on the first energy function 175a of the first energy-based model 170a. In some cases, the noisy embedding 156 may be modified to achieve a lower energy configuration equivalent to a sample drawn from a dense region of the noisy data distribution, as indicated by the energy values output by the first energy function 175a. In some cases, the protein design computational model 115 may perform gradient-based Markov chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov chain Monte Carlo (MCMC) sampling) of the noisy data distribution in which the noisy embedding 156 of the input sequence 152 is modified over multiple successive iterations, with each iteration sampling from an incrementally denser region of the noisy data distribution to increase the likelihood that the resulting modified noisy embedding 158 is in the noisy data distribution. Further, in some cases, modifications made to the noisy embedding 156 of the input sequence 152 may be accumulated over multiple successive iterations. For example, in some cases, the noisy embedding 156 of the input sequence 152 may undergo a first modification and a second modification. The protein design computational model 115 may apply a first energy function 175a to determine a first energy value of the noisy embedding 156 with the first modification and a second energy value of the noisy embedding 156 with the second modification. For subsequent iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling, if the first energy value is lower than the second energy value, the first energy-based model 170a may be applied to further modify the noisy embedding 156 with the first modification, indicating that the noisy embedding 156 with the first modification was sampled from a dense region of the noisy data distribution and is likely within the noisy data distribution.In some cases, one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling may be performed, and the protein design computational model 115 applies the first energy-based model 170a to further modify the noisy embedding 156 of the input sequence 152 until one or more criteria are met. For example, in some cases, the protein design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until a threshold number of iterations have been performed. Alternatively and / or additionally, the protein design computational model 115 may perform one or more additional iterations of gradient-based Markov chain Monte Carlo (MCMC) sampling until the energy value of the modified noisy embedding 156 or the likelihood that the modified noisy embedding 156 is within the noisy data distribution meets one or more thresholds.
[0111] At 258, the protein design engine 110 may denoise the modified noisy embedding of the input sequence to generate a denoised embedding of the input sequence. In some exemplary embodiments, the denoising engine 117 of the protein design computational model 115 may generate a denoised embedding 160 by denoising at least the modified noisy embedding 158 generated by the protein design computational model 115 (e.g., first energy-based model 170a). In some cases, the denoising engine 117 may include one or more machine learning models (e.g., Transformers, etc.) trained to denoise the modified noisy embedding 158 and recover the denoised embedding 160 therefrom. For example, in some cases, the one or more machine learning models may be trained to recover a corresponding known protein sequence for each noisy sample sequence in the noisy training set based on the noisy training set.
[0112] To further illustrate, as noted above, Known protein sequences in TIFF2026508122000066.tif5170 TIFF2026508122000067.tif4170 is converted into a noisy sample array by adding noise. TIFF2026508122000068.tif5170. Thus, in some cases, the denoising engine 117 may apply the modified noisy embedding 158 generated by the protein design computational model 115 (e.g., the first energy-based model 170a) to the sample sequence The denoising may be based at least on a least squares estimator of TIFF2026508122000069.tif4170, which may be given by:
number
number
number
[0113] At 260, the protein design engine 110 may generate an output sequence by at least decoding the denoising embedding of the input sequence. In some exemplary embodiments, the decoder 119 of the protein design engine 110 may generate the output sequence 162 by at least decoding the denoising embedding 160 generated by the denoising engine 117. As mentioned above, in some cases, the protein design computational model 115 (e.g., the first energy-based model 170a) may operate in a noisy latent space to generate a modified noisy embedding 158 if the noisy embedding 156 of the input sequence 152 incorporates additional information (e.g., structural tokens, etc.). Thus, in some cases, in addition to denoising the modified noisy embedding 158, the resulting denoised embedding 160 may be decoded by the decoder 119 to generate the output sequence 162. For example, in some cases, decoding the denoised embedding 160 may include determining the identity and sequential positions of the amino acid residues that form the output sequence 162 based on at least the tokens of the denoised embedding 160.
[0114] 3A depicts a flowchart illustrating an example of a process 300 for training a protein design computational model 115, according to some exemplary embodiments. Referring to FIGS. 1, 2A, and 3A, process 300 may be executed by protein design engine 110 to train a protein design computational model 115, such as, for example, first energy-based model 170a and second energy-based model 170b. As described in more detail below, in some cases, protein design computational model 115 may be trained through gradient-based Markov chain Monte Carlo (MCMC) sampling, including, for example, Markov chain Monte Carlo (MCMC) sampling with Langevin dynamics. Additionally, in some cases, process 300 may implement operation 204 of process 200 shown in FIG. 2A.
[0115] At 302, the protein design engine 110 may apply an energy-based model (EBM) model to generate a first modified sequence. In some exemplary embodiments, the protein design engine 110 may train a protein design computational model 115, including, for example, a first energy-based model 170a, to approximate a data distribution of protein sequences exhibiting one or more desirable properties so that additional protein sequences exhibiting the same desirable properties can be generated by sampling from it. In some cases, the first energy-based model 170a may be trained to approximate such a data distribution based on a training set of sample sequences, each of which is a known protein sequence from the data distribution. In some cases, instead of being trained directly on known protein sequences, the first energy-based model 170a may be trained based on a noisy embedding of known protein sequences. That is, in some cases, the first energy-based model 170a may be trained based on a noisy training set of noisy sample sequences, each of which is an embedding of a known protein sequence contaminated with noise (e.g., Gaussian noise, such as isotropic Gaussian noise).
[0116] In some exemplary embodiments, training the first energy-based model 170a may include applying the first energy-based model 170a to modify an initial sequence (e.g., a known protein sequence or a noise sequence) and adjusting parameters (e.g., weights, biases, etc.) of the first energy-based model 170a to increase the similarity between the resulting modified sequence and the noisy sample sequences in the noisy training set, e.g., incrementally over multiple successive iterations. In some cases, the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a may undergo different adjustments before further adjustments are made that result in protein sequences that are more similar to the noisy sample sequences in the noisy training set. For example, in some cases, a first adjustment may be made to the parameters of the first energy-based model 170a before the first energy-based model 170a with the first adjustment is applied to modify an input sequence and generate at least a first modified sequence. As described in more detail below, the first energy-based model 170a with the second adjustment may be applied to generate at least a second modified sequence before further adjustments are made to either the first or second adjustments to the first energy-based model 170a. In doing so, the first energy-based model 170a may be trained to approximate a noisy data distribution filling a continuous latent space, which may facilitate subsequent sampling (e.g., gradient-based Markov Chain Monte Carlo (MCMC) sampling, etc.), as described above.
[0117] In some demonstrative embodiments, training the first energy-based model 170a may further include determining a first energy function 175a. As described above, in some cases, the first energy function 175a may be parameterized by parameters (e.g., weights, biases, etc.) of the first energy-based model 170a. Thus, in some cases, training the first energy-based model 170a, which includes adjusting the parameters of the first energy-based model 170a, may also include adjusting the parameters of the first energy function 175a. For example, in some cases, the first energy function 175a may be determined by performing gradient-based Markov Chain Monte Carlo (MCMC) sampling (e.g., Langevin Markov Chain Monte Carlo (MCMC) sampling, etc.) to approximate the gradient of the noisy data distribution. Doing so may include adjusting parameters of the first energy function 175a over multiple successive iterations, where the first energy function 175a assigns lower energy values to first sequences that are similar to the noisy sample sequences of the noisy training set than to second sequences that are dissimilar to the noisy sample sequences of the training set. Once the first energy-based model 175a is trained, the first energy function 175a may output energy values that distinguish between protein sequences sampled from high-density regions of the noisy data distribution and protein sequences sampled from low-density regions of the noisy data distribution.
[0118] At 304, the protein design engine 110 may apply the energy-based model with the second adjustment to generate a second modified sequence. In some exemplary embodiments, upon applying the first energy-based model 170a with the first adjustment to generate at least the first modified sequence, the protein design engine 110 may apply the first energy-based model 170a with the second adjustment to generate at least the second modified sequence. It should be understood that the first adjustment and the second adjustment may include different changes to the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a. Thus, applying the first energy-based model 170a with the second adjustment to modify an input sequence may result in a different modified sequence than applying the first energy-based model 170a to modify the same input sequence.
[0119] At 306, the protein design engine 110 may determine that the first modified sequence is more similar to the sample sequences of the training set than the second modified sequence. In some exemplary embodiments, the protein design engine 110 may select the first energy-based model 170a with the first adjustment for further adjustment during subsequent iterations if the modified sequence generated by the first energy-based model 170a with the first adjustment is more similar to the sample sequences of the training set, or possibly to the noisy sample sequences in the noisy training set. The similarity of the first modified sequence to the sample sequences of the training set (or the noisy sample sequences in the noisy training set) than the second modified sequence may indicate that the first energy-based model 170a with the first adjustment more closely matches the data distribution of the sample sequences (or the noisy sample sequences) than the second energy-based model 170a with the second adjustment. In some cases, the similarity between the modified sequence generated by the first energy-based model 170a and the sample sequences of the training set (or the noisy sample sequences of the noisy training set) may be quantified by a similarity metric. Examples of similarity metrics include antibody similarity metrics (e.g., biophysical properties such as molecular weight, length, hydrophobicity, hydrophilicity, etc.), sequence similarity (e.g., edit distance, etc.), naturalness metrics (e.g., likelihood under a pre-trained protein language model), etc. In some cases, the protein design engine 110 may select the first energy-based model 170a with the first adjustment instead of the first energy-based model 170a with the second adjustment to undergo one or more additional iterations of adjustment based on at least the first modified sequence having a higher similarity metric than the second modified sequence.
[0120] At 308, the protein design engine 110 may further adjust the energy-based model with the first adjustment instead of the second adjustment until one or more criteria are met. In some exemplary embodiments, the protein design engine 110 may further adjust the first energy-based model 170a with the first adjustment instead of the first energy-based model 170b with the second adjustment if the first modified sequence generated by the first energy-based model 170a with the first adjustment is more similar to the sample sequence of the training set (or the noisy sample sequence of the noisy training set) than the second modified sequence generated by the first energy-based model 170a with the second adjustment. For example, during subsequent iterations of adjustment, the protein design engine 110 may make further adjustments to the parameters (e.g., weights, biases, etc.) of the first energy-based model 150a with the first adjustment before applying the further adjusted first energy-based model 150a to generate one or more additional modified sequences. In some cases, the first energy-based model 170a may be further adjusted to further increase the similarity between the corrected sequences output by the first energy-based model 170a and the sample sequences of the training set (or the noisy sample sequences of the noisy training set). In some cases, the protein design engine 110 may continue to adjust the first energy-based model 170a until one or more criteria are met. For example, in some cases, the protein design engine 110 may continue to adjust the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a until the protein design engine 110 performs a threshold number of iterations of adjustment. Alternatively and / or additionally, the protein design engine 110 may continue to adjust the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a until the similarity (or similarity metric) between the corrected sequences generated by the first energy-based model 170a and the sample sequences of the training set (or the noisy sample sequences of the noisy training set) meets one or more thresholds.
[0121] 3B depicts a flowchart showing another example of a process 350 for training a protein design computational model 115, according to some exemplary embodiments. Referring to FIGS. 1, 2A, and 3B, process 350 may be performed by protein design engine 110 to train a protein design computational model 115, such as first energy-based model 170a and second energy-based model 170b. In some exemplary embodiments, process 350 may be performed to train first energy-based model 170a to approximate a first data distribution of protein sequences based on the gradient of a second data distribution of protein sequences. For example, in some cases, process 350 may be performed when too few known protein sequences characterizing the first data distribution are available to train first energy-based model 170a. As described in more detail below, in some cases, a second energy-based model 170b may be trained to approximate a second data distribution of protein sequences, such that a second energy function 175b may be applied to provide additional guidance, while the first energy-based model 170a is trained to approximate a first data distribution, for example, via gradient-based Markov Chain Monte Carlo sampling across multiple data distributions (e.g., Langevin Markov Chain Monte Carlo, etc.). In some cases, process 350 may perform operation 204 of process 200 shown in FIG. 2A.
[0122] At 352, the protein design engine 110 may determine a first adjustment to the first energy-based model that reduces a first difference between a first output sequence generated by the first energy-based model and a first plurality of sample sequences from a first data distribution of protein sequences. In some exemplary embodiments, the protein design engine 110 may combine the training of multiple energy-based models, including, for example, a first energy-based model 170a and a second energy-based model 170b. For example, in some cases, training of the first energy-based model 170a to approximate a first data distribution of protein sequences may be combined with training of the second energy-based model 170b to approximate a second data distribution when an insufficient amount of known protein sequences from the first data distribution is available for training the first energy-based model 170a. Thus, in some cases, the protein design engine 110 may determine a first adjustment for the first energy-based model 170a that increases the similarity (or similarity metric) between the output sequence produced by the first energy-based model 170a and a sample sequence from a first data distribution (e.g., a noisy sequence embedding from a noisy data distribution). However, as described in more detail below, instead of directly applying the first adjustment to the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a, the protein design engine 110 may further determine the adjustment made to the parameters (e.g., weights, biases, etc.) of the first energy-based model 170a based on the gradient of a second energy function 175b of a second energy-based model 170b trained to approximate a second data distribution of protein sequences.
[0123] At 354, the protein design engine 110 may determine second adjustments to the second energy-based model that reduce the difference between a second output sequence generated by the second energy-based model and a second plurality of sample sequences from the second data distribution. In some exemplary embodiments, the protein design engine 110 may determine second adjustments to parameters (e.g., weights, biases, etc.) of the second energy-based model 170b to increase the similarity between one or more output sequences generated by the second energy-based model 170b and sample sequences from a second data distribution of protein sequences (e.g., noisy sequence embeddings from a noisy data distribution). As described above, an insufficient amount of known protein sequences from the first data distribution may be available to train the first energy-based model 170a to approximate the first data distribution, while a greater amount of known protein sequences from the second data distribution may be available to train the second energy-based model 170b to approximate the second data distribution. Thus, in some cases, the density of at least some regions in the first data distribution may be uncertain due to the lack of known protein sequences embedded in those regions. In regions of the first data distribution where the density of the first data distribution cannot be determined due to the lack of known protein sequences embedded in those regions, the gradient of the second energy function 175b may provide a surrogate density estimate. Thus, combining the training of the first energy-based model 170a and the second energy-based model 170a may improve the performance of the first energy-based model 170a by at least increasing the precision and accuracy of the approximation of the first data distribution.
[0124] At 356, the protein design engine 110 may train the first energy-based model by applying a third adjustment to the first energy-based model, the third adjustment being determined based on at least the first adjustment and the second adjustment. In some exemplary embodiments, the protein design engine 110 may determine a third adjustment to apply to parameters (e.g., weights, biases, etc.) of the first energy-based model 170a based on at least the first adjustment and the second adjustment. For example, in some cases, the third adjustment may be a sum of the first adjustment and the second adjustment. Alternatively, the third adjustment may be a weighted sum in which the first adjustment and the second adjustment are associated with different weights. Training the first energy-based model 170a may include applying the third adjustment to parameters (e.g., weights, biases, etc.) of the first energy-based model 170a. In some cases, the third adjustment may obtain an estimate of density across the first data distribution of protein sequences and the second data distribution of protein sequences, including regions of the first data distribution where the density of the first data distribution is uncertain due to a lack of known protein sequences that fill those regions. Thus, applying the third adjustment to the first energy-based model 170a may enable the first energy-based model 170a to better approximate the first data distribution despite a lack of known protein sequences that characterize at least some regions of the first data distribution.
[0125] As described above, in some exemplary embodiments, the first energy-based model 170a may be trained to approximate and subsequently sample from a noisy data distribution of noisy protein sequences instead of a true data distribution of protein sequences not perturbed by any noise. Training the first energy-based model 170a to approximate a data distribution of protein sequences, such as a data distribution of protein sequences exhibiting particular desirable properties (e.g., drug-like properties), may include determining the first energy function 170a such that the first energy function 170a assigns lower energy values to protein sequences sampled by higher density regions of the first data distribution than to those sampled from lower density regions of the first data distribution. Furthermore, the slope of the first energy function 170a may approximate the change in density across the first data distribution.
[0126] To further illustrate, Figure 4A depicts a schematic diagram showing an example of sampling from a noisy data distribution, according to some exemplary embodiments. As shown in Figure 4A, TIFF2026508122000087.tif4170 is noise Adding TIFF2026508122000088.tif5170 creates a noisy sequence Can be converted to TIFF2026508122000089.tif4170. Noise Addition of TIFF2026508122000090.tif5170 is a known protein sequence TIFF2026508122000091.tif4170, a noisy sequence TIFF2026508122000092.tif4170, which can be used to generate a noisy data distribution populated by known protein sequences. TIFF2026508122000093.tif4170 exhibits a smoother energy landscape than the data distribution embedded by TIFF2026508122000093.tif4170. In some cases, the first energy-based model 170a may sample from a noisy data distribution, which involves "walking" its energy landscape toward incrementally higher density regions of protein sequences embedded in the noisy data distribution that exhibit desired properties. For example, FIG. 4A shows that "walking" through the energy landscape occurs over a period of sampling iterations. Sample TIFF2026508122000094.tif4170 TIFF2026508122000095.tif4170, sampling iteration Sample in TIFF2026508122000096.tif4170 TIFF2026508122000097.tif4170, also sampling repeat Sample TIFF2026508122000098.tif4170 In some cases, the "walk" across the energy landscape of a noisy data distribution involves extracting the sample Energy values of TIFF2026508122000100.tif4170 are sample The energy value is lower than that of TIFF2026508122000101.tif4170, sample Energy values of TIFF2026508122000102.tif4170 are sample The energy value of the first energy function 175a may be further lower than that of TIFF2026508122000103.tif4170, as guided by the gradient of the first energy function 175a. Furthermore, each sampling iteration may involve further modifying the samples extracted during the previous iteration. Thus, as shown below, the energy value of the first energy function 175a may be further modified by the gradient of the first energy function 175a. Samples drawn from a noisy data distribution between TIFF2026508122000104.tif4170 TIFF2026508122000105.tif4170 is the previous sampling iteration Sample taken between TIFF2026508122000106.tif4170 Noise generated based on TIFF2026508122000107.tif4170 TIFF2026508122000108.tif4170 is normally distributed at each sampling iteration. Extracted from TIFF2026508122000109.tif4170.
number
[0127] Referring again to FIG. 4A, in some cases, the protein sequence TIFF2026508122000111.tif4170 is the corresponding noisy protein sequence drawn from the noisy data distribution TIFF2026508122000112.tif4170 is denoised and denoising engine 117 uses least squares estimator TIFF2026508122000113.tif5170 (e.g. TIFF2026508122000114.tif6170) to project back to the true data distribution. This constitutes the "jump" shown in Figure A. Furthermore, in the example shown in Figure 4, while the first energy-based model 170a "walks" the energy landscape of the noisy data distribution and draws samples from it, a "jump" back to the true data distribution may be performed at each sampling iteration. For example, for a protein sequence TIFF2026508122000115.tif5170 is a sampling repeat Samples drawn from a noisy data distribution between TIFF2026508122000116.tif4170 TIFF2026508122000117.tif4170 can be generated when it is denoised and projected back into the true data distribution, but the protein sequence TIFF2026508122000118.tif5170 is a subsequent sampling iteration Samples drawn from a noisy data distribution between TIFF2026508122000119.tif4170 TIFF2026508122000120.tif4170 may be generated when the data is denoised and projected back into the true data distribution. As described above, the first energy-based model 170a may continue to "walk" the energy landscape of the noisy data distribution and draw samples from it until one or more criteria are met. For example, the first energy-based model 170a may continue to draw samples from a threshold number of sampling iterations if a threshold number of sampling iterations have been performed at that point. Alternatively and / or additionally, the first energy-based model 170a may continue to "walk" the energy landscape of the noisy data distribution up to sample If TIFF2026508122000122.tif4170 exhibits a threshold energy value or a possible threshold in a noisy data distribution, sample One may continue to "walk" the energy landscape of the noisy data distribution until TIFF2026508122000123.tif4170 is elicited.
[0128] Protein sequence Training the first energy-based model 170a to approximate the data distribution of TIFF2026508122000124.tif4170 may cause the first energy-based model 170a to overfit to those particular sequences. This is because the first energy-based model 170a may not be able to accurately predict the distribution of these protein sequences. This means that the first energy-based model 170a can accurately approximate the density of regions of the data distribution that are close to, but not beyond, the data distribution. This phenomenon is illustrated in the top panel (A) of Figure 4B, which shows that the slope (or density estimate) of the data distribution is inaccurate for a large portion of the data distribution. The inability of the first energy-based model 170a to accurately approximate the density of large bands of the data distribution can prevent the first energy-based model 170a from adequately exploring the data distribution during sampling, thus causing mode collapse where the output of the first energy-based model 170a lacks the necessary diversity. In contrast, the bottom panel (B) of Figure 4B shows that the first energy-based model 170a exhibits a significant decrease in the density of regions of the data distribution that are close to, but not beyond, the data distribution. We show that training the first energy-based model 170a based on TIFF2026508122000126.tif4170 can enable the first energy-based model 170a to accurately approximate the density of a larger portion of the data distribution. Thus, training the first energy-based model 170a to approximate a noisy data distribution can prevent overfitting as well as mode collapse.
[0129] 5A depicts a schematic diagram showing an example of sampling from a smoothed discrete space, according to some exemplary embodiments. FIG. 5A shows one variation of a generation process in which the protein design computational model 115 operates within the smoothed discrete space formed when the noise engine 113 adds noise (e.g., Gaussian noise, etc.) to the protein sequence. For example, FIG. 5A shows protein sequences occupying a discrete space (e.g., discrete amino acid space) that is filled by individual (or discrete) protein sequences. TIFF2026508122000127.tif3170 and TIFF2026508122000128.tif4170, each represented by a constituent sequence of amino acid residues. Adding noise (e.g., Gaussian noise) to TIFF2026508122000129.tif3170 creates the first noisy array TIFF2026508122000130.tif4170 can be generated, which contains the protein sequence TIFF2026508122000131.tif3170 into the aforementioned smoothed discrete space, which presents a smoother energy landscape than the initial discrete space. Figure 5A shows how the first energy-based model 170a projects the smoothed discrete space into a noisy array Second noisy sequence from tif2026508122000132.tif4170 TIFF2026508122000133.tif5170 illustrates that the smoothed discrete space may be sampled by "walking" the first energy-based model 170a through multiple successive iterations (e.g., gradient-based Markov Chain Monte Carlo (MCMC) sampling iterations, etc.) of the first noisy sequence. First noisy sequence by modifying TIFF2026508122000134.tif4170 Second noisy sequence from tif2026508122000135.tif4170 TIFF2026508122000136.tif5170. The "walk" through the smoothed discrete space can be guided by the first energy function 175a (e.g., the gradient of the first energy function 175a). Thus, in some cases, the second noisy sequence TIFF2026508122000137.tif5170 is the first noisy sequence Second noisy sequence compared to TIFF2026508122000138.tif4170 This may include a modification to reduce the energy value of TIFF2026508122000139.tif5170, which is the second noisy sequence This means that TIFF2026508122000140.tif5170 is sampled from a dense region in a noisy discrete space.
[0130] To further illustrate, Figure 5B depicts a block diagram illustrating an example of a discrete energy-based model (dEBM) for implementing the first energy-based model 170a, according to some exemplary embodiments. As shown in Figure 5B, the discrete energy-based model (dEBM) The first noisy sequence is taken from TIFF2026508122000141.tif4170 and passed through a multi-layer perceptron (MLP) and a convolutional neural network (CNN). Positionally encode TIFF2026508122000142.tif4170 TIFF2026508122000143.tif4170 (e.g., one-dimensional position encoding TIFF2026508122000144.tif4170) and concatenate the first noisy sequence Embedding TIFF2026508122000145.tif4170 Generate the concatenated output with TIFF2026508122000146.tif4170 and the hidden state TIFF2026508122000147.tif4170 can then be created. TIFF2026508122000148.tif4170 is passed through a multilayer perceptron (MLP) to generate the energy function Returns TIFF2026508122000149.tif5170.
[0131] Referring again to FIG. 5A, in some cases, the "walk" through the smoothed discrete space is a second noisy sequence The first energy-based model 170a may include drawing multiple intermediate samples from the smoothed discrete space, each intermediate sample being an incrementally lower energy configuration drawn from a higher density region of the smoothed discrete space, before arriving at the second noisy array 170a. Further, the first energy-based model 170a may continue to "walk" the smoothed discrete space until one or more criteria are met, at which point the first energy-based model 170a may continue to "walk" the smoothed discrete space until one or more criteria are met, at which point the second noisy array 170a may be drawn. TIFF2026508122000151.tif5170 is denoised by the denoising engine 117 to obtain the protein sequence TIFF2026508122000152.tif4170. The second noisy sequence, as shown in Figure 5A, can be generated. The denoising of TIFF2026508122000153.tif5170 may constitute a "jump" back to discrete space. Thus, the protein sequence TIFF2026508122000154.tif4170 is a distinct protein sequence represented by the constituent sequence of amino acid residues.
[0132] 6 depicts a schematic diagram showing an example of sampling from a smoothed latent space, according to some exemplary embodiments. Figure 6 illustrates another variation of the generative process in which the protein design computational model 115 operates in a smoothed latent space formed when the noise engine 113 adds noise (e.g., Gaussian noise, etc.) to the protein sequence embeddings generated by the encoder 111. In some cases, the protein sequence Before adding noise (e.g., Gaussian noise) to TIFF2026508122000155.tif3170, the encoder 111 calculates at least the protein sequence TIFF2026508122000156.tif3170 by enriching it with additional information (e.g., structural information, environmental information, etc.). Embedding TIFF2026508122000157.tif3170 TIFF2026508122000158.tif3170 can be generated. As shown in Figure 6, the protein sequence Embedding TIFF2026508122000159.tif3170 TIFF2026508122000160.tif3170 may occupy the latent space occupied by sequence embeddings instead of discrete protein sequences found in discrete space (e.g., discrete amino acid space). The noise engine 113 then generates a noise map for the protein sequence. Embedding TIFF2026508122000161.tif3170 The first noisy sequence is generated by adding noise (e.g., Gaussian noise) to TIFF2026508122000162.tif3170. Generate TIFF2026508122000163.tif4170. (As in Figure 5A) Protein sequence Instead of adding noise directly to the protein sequence, Embedding TIFF2026508122000165.tif3170 Adding noise to TIFF2026508122000166.tif3170 results in a smoothed latent space filled by a noisy sequence embedding. TIFF2026508122000167.tif3170. A smoothed latent space may be more continuous and semantically meaningful than its discrete counterpart, at least because the distance between two or more sequence embeddings in the smoothed latent space may reflect the similarity (or dissimilarity) in protein sequences as well as the conformation (or three-dimensional structure) of the protein.
[0133] Referring again to Figure 6, the first energy-based model 170a can "walk" the smoothed latent space while being guided by the first energy function 175a. In a variation of the generation process shown in Figure 6, the first energy-based model 170a can be generated by the noise engine 113 generating the protein sequence, as described above. Embedding TIFF2026508122000168.tif3170 The first noisy sequence is generated by adding noise to "TIFF2026508122000169.tif3170" The first energy-based model 170a may begin the "walk" by drawing one or more samples from the smoothed latent space, each sample including a modification that further reduces its energy value relative to one or more preceding samples. TIFF2026508122000171.tif4170 and the second noisy sequence TIFF2026508122000172.tif5170, where each intermediate sample is an incrementally lower energy configuration drawn from a higher density region of the smoothed latent space. Further, the first energy-based model 170a may continue to "walk" the smoothed latent space until one or more criteria are met, at which point the denoising engine 117 may determine whether the protein sequence TIFF2026508122000173.tif4170 is embedded with noise removed The denoised embedding is generated by decoder 119, which decodes TIFF2026508122000174.tif5170. The second noisy sequence yields TIFF2026508122000175.tif4170 TIFF2026508122000176.tif5170 can be denoised.
[0134] In some exemplary embodiments, training the protein design computational model 115, particularly the first energy-based model 170a, based on a noisy training set containing noisy sample sequences prevents overfitting of the validation loss during maximum likelihood training. As shown in FIG. 7A, the loss of the first energy-based model 170a can converge quickly (e.g., in about 50 training steps) and plateau (e.g., in 100 or more steps) without overfitting. Making the sample sequences noisy provides a strong regularization that prevents overfitting. This effect is significant across a range of noise levels. See TIFF2026508122000177.tif5170. Noise level It should be understood that TIFF2026508122000178.tif6170 (noise-free) is a special case that reflects the reconstruction accuracy of the encoder 111 and decoder 119, or alternatively, the baseline error that may be present in a sequence that undergoes encoding and decoding without adding noise (e.g., TIFF2026508122000179.tif6170), the true protein sequences and protein sequences reconstructed by the decoder 119 from the true protein sequence embeddings generated by the encoder 111 have very few edits (e.g., on average) compared to the clean sample sequences. TIFF2026508122000180.tif6170). These edits tend to occur at higher entropy positions (e.g., positions that are more likely to be occupied by different amino acid residues across different protein sequences), which may reflect the biophysical diversity observed in naturally occurring protein sequences (e.g., antibodies, etc.). However, noise (e.g., Without a standardized data set (TIFF2026508122000181.tif6170), sampling remains challenging because the energy landscape of the protein sequence data distribution lacks the smoothing brought about by the introduction of noise in the sample sequences.
[0135] In some exemplary embodiments, the analysis engine 130 may determine the performance of the protein design computational model 115 based on at least the output sequences 156 across a suite of "antibody-similarity" (ab-similarity) metrics, including, for example, labels derived from amino acid sequences using BioPyton, sequence similarity scores from sequence alignments using DIAMOND, Levenstein edit distances calculated using Edlib, naturalness metrics calculated from the likelihood of masked language models pre-trained on antibody sequences, etc. The sequence feature metrics may be a normalized average Wasserstein distance between the feature distribution of sample sequences in the training set and the validation set, or a normalized average Wasserstein distance between the feature distribution of sample sequences in the training set and the validation set. This can be condensed into a single scalar metric by calculating the average sum edit distance: TIFF2026508122000183.tif6170 summarizes the novelty and diversity of the samples compared to the validation set. The results, summarized in Table 1 below, are based on the variance, which controls the amount of noise added to the sample sequences in the training set. TIFF2026508122000184.tif3170 shows that as the variance increases, a better agreement is reached between the sample feature distribution and the validation set. The DIAMOND similarity metric (Figure 7B) and naturalness metric (Figure 7C) distributions show that the protein design computational model can generate natural sequences with reasonable similarity to the training sequences in the training set while maintaining sequence diversity and sequence novelty.
[0136] [Table 1]
[0137] 8 depicts a schematic diagram showing distributional fitness score-based evaluation of in silico protein designs generated by the protein design computational model 115 against a reference set of validation samples, according to some exemplary embodiments. In some exemplary embodiments, the distributional fitness score may quantify the likelihood of an in silico protein design (e.g., output sequence 162 of FIG. 1 ) relative to a reference distribution while preserving novelty and diversity. In some cases, the distributional fitness score of an in silico protein design may directly correspond to the viability of the in silico protein design as an actual, biophysically valid protein. In some cases, the probability of an in silico protein design fitting a reference distribution may be assessed using a conformal transducer system. For example, TIFF2026508122000187.tif5170, TIFF2026508122000188.tif5170, and TIFF2026508122000189.tif5170, where TIFF2026508122000190.tif3170 shows the characteristics of the sample, TIFF2026508122000191.tif4170 shows the labels. TIFF2026508122000192.tif4170 is the sequence TIFF2026508122000193.tif5170 is a set of real numbers TIFF2026508122000194.tif5170, which can be a measurable function that is equivariant under permutation. New Sample Given TIFF2026508122000195.tif3170, the relevance measure is TIFF2026508122000196.tif4170 is How good is TIFF2026508122000197.tif3170? It is then possible to quantify whether the conformal transducer is similar to TIFF2026508122000198.tif5170. TIFF2026508122000199.tif4170 Each label can be defined as a system of values TIFF2026508122000200.tif5170, reference sequence TIFF2026508122000201.tif5170, and test sample Regarding TIFF2026508122000202.tif4170, There is TIFF2026508122000203.tif7170, and in the formula TIFF2026508122000204.tif7170. Intuitively TIFF2026508122000205.tif5170 is The percentage of in silico protein designs that have a higher fit to the reference distribution than the TIFF2026508122000206.tif5170. In this context, the fitness measure TIFF2026508122000207.tif4170 can be defined as the likelihood under a joint density (e.g., calculated using kernel density estimation) across various properties, such as biophysical and statistical properties (e.g., log-probability under a protein language model). TIFF2026508122000208.tif4170 may contain a set of known protein sequences (e.g., antibodies), labeled TIFF2026508122000209.tif4170 may represent certain desirable properties (e.g., expression, binding affinity, etc.).
[0138] As mentioned above, in some cases, the performance of a protein design computational model 115 may be measured based on a set of "antibody-likeness" (ab-similarity) metrics. The sequence property metrics include a distribution fit score and a normalized average Wasserstein distance between the property distributions of the in silico protein designs and the validation set. This can be condensed into a single scalar metric by calculating the average sum edit distance: TIFF2026508122000211.tif6170 summarizes the novelty and diversity of in silico protein designs, while internal diversity ( TIFF2026508122000212.tif6170) represents the average sum edit distance between in silico protein designs as a group. As shown in Table 2 (below), the protein design computational model 115 can be, for example, Strong antibody similarity (ab-similarity) was achieved when increasing the noise level to TIFF2026508122000213.tif6170. Furthermore, both implementations of the protein design computational model (e.g., energy-based sampling and score-based sampling) achieved faster sampling times and lower memory footprints than conventional methods such as latent sequence diffusion (SeqVDM), score-based models with energy parameterization (DEEN), and pre-trained large-scale language models (GPT3.5).
[0139] [Table 2]
[0140] The performance of the protein design computational model 115 in generating natural, novel, and diverse protein designs was also evaluated in vitro, with the protein design computational model 115 achieving a 97.47% success rate in vitro, with 270 of the 277 in silico antibody designs being successfully expressed and purified in the laboratory. These results are shown in Table 3 below.
[0141] [Table 3]
[0142] Additionally, the performance of the protein design computational model 115 in generating functional protein designs was evaluated in vitro, and the protein design computational model 115 generated a higher percentage of binding antibodies than other methods, such as latent sequence diffusion (SeqVDM), pre-trained large-scale language model (GPT4), Transformer model, and equivariate graph neural network (EGNN). These results are shown in Table 4 below.
[0143] [Table 4]
[0144] The performance of a protein design computational model 115 operating in latent space (lWJS) with different noise levels (sigma) instead of discrete space (dWJS) was also evaluated using the metric Wasserstein distance ( TIFF2026508122000217.tif7170), uniqueness, edit distance ( TIFF2026508122000218.tif6170), and internal diversity ( TIFF2026508122000219.tif6170). Table 5 below summarizes the results of 2000 in silico antibody heavy chain designs generated based on 20 novel seed sequences.
[0145] [Table 5]
[0146] In view of the above-described embodiments of the subject matter, the present application discloses the following list of examples, which, in combination with one feature of a single example or two or more features of said examples, optionally in combination with one or more features of one or more additional examples, are further examples included in the disclosure of the present application.
[0147] Item 1: A computer-implemented method, comprising: generating a first training set including a plurality of noisy sample sequences, wherein each noisy sample sequence in the first training set is generated by at least adding noise to a corresponding sample sequence from a first data distribution; training a protein design computational model by applying the protein design computational model to generate at least one or more output sequences and adjusting the protein design computational model to reduce differences between the one or more generated output sequences and the plurality of noisy sample sequences of the first training set; and applying the trained protein design computational model to generate output sequences by modifying at least input sequences.
[0148] Item 2: The method of item 1, wherein the protein design computational model includes a first energy-based model (EBM).
[0149] Item 3: The method of item 2, wherein training the protein design computational model includes adjusting a plurality of parameters of the first energy-based model that parameterize an energy function of the first energy-based model.
[0150] Item 4: The method described in Item 3, wherein the plurality of parameters are adjusted so that energy values determined by the energy function correspond to likelihoods of one or more generated output sequences within the first data distribution.
[0151] Item 5: The method of item 3 or 4, wherein the plurality of parameters are adjusted such that the energy function outputs a lower energy value for a first generated output sequence that is similar to the plurality of noisy samples of the first training set than for a second generated output sequence that is dissimilar to the plurality of noisy samples of the first training set.
[0152] Item 6: The method of any one of Items 3 to 5, wherein training the protein design computational model includes applying a first energy-based model with a first adjustment to generate a first modified sequence, applying the first energy-based model with a second adjustment to generate a second modified sequence, and upon determining that the first modified sequence is more similar to the plurality of noisy samples of the first training set than the second modified sequence, further modifying the first energy-based model with the first adjustment instead of the second adjustment.
[0153] Item 7: The method of item 6, wherein the first energy-based model is further adjusted until one or more criteria are met, the one or more criteria including at least one of (i) performing a threshold number of iterations of adjustments to the first energy-based model, and (ii) the second modified sequence exhibiting a threshold similarity to a plurality of noisy samples of the first training set.
[0154] Item 8: The method of any one of Items 2 to 7, wherein the protein design computational model further includes a second energy-based model (EBM).
[0155] Item 9: The method of item 8, further comprising: generating a second training set including a plurality of sample sequences from a second data distribution; determining a first adjustment to the first energy-based model that reduces a first difference between a first output sequence generated by the first energy-based model and the plurality of noisy sample sequences of the first training set; determining a second adjustment to the second energy-based model that reduces a second difference between a second output sequence generated by the second energy-based model and the plurality of sample sequences of the second data distribution; and training the first energy-based model by applying at least a third adjustment to the first energy-based model determined based on the first adjustment and the second adjustment.
[0156] Item 10: The method of item 9, wherein the third adjustment corresponds to the sum or weighted sum of the first adjustment and the second adjustment.
[0157] Item 11: The method of any one of items 1 to 10, further comprising: encoding each sample sequence from a first data distribution to generate an embedding for each sample sequence; and generating a plurality of noisy sample sequences of a first training set by at least adding noise to the embedding of each sample sequence.
[0158] Item 12: The method of item 11, wherein each sample sequence from the first data distribution is encoded by augmenting it with additional information.
[0159] Item 13: The method of Item 12, wherein the additional information includes, for each constituent amino acid residue, structural information identifying one or more adjacent amino acid residues in three-dimensional space.
[0160] Item 14: The method of any one of items 1 to 13, wherein the trained protein design computational model generates an output sequence by at least generating a noisy input sequence by adding noise to at least the input sequence, applying an energy-based model to generate a noisy output sequence by at least modifying the noisy input sequence based at least on an energy function of the energy-based model, and generating an output sequence by denoising at least the modified noisy output sequence generated by the energy-based model.
[0161] Item 15: The method of any one of items 1 to 14, wherein the trained protein design computational model generates an output sequence by at least: generating an embedding of the input sequence by at least encoding the input sequence; generating a noisy embedding of the input sequence by at least adding noise to the embedding of the input sequence; applying an energy-based model to at least modify the noisy embedding of the input sequence based at least on an energy function of the energy-based model to generate a modified noisy embedding; denoising the noisy embedding to generate a denoised embedding; and generating an output sequence by at least denoising the noisy embedding.
[0162] Item 16: The method of item 15, wherein the embedding of the input sequence is generated by at least generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
[0163] Item 17: The method of Item 15 or 16, wherein the embedding of the input sequence is generated by at least generating, for at least one amino acid residue of the input sequence, one or more structural tokens that identify one or more adjacent amino acid residues in three-dimensional space.
[0164] Item 18: The method of any one of items 1 to 17, wherein the trained protein design computational model modifies the input sequence by at least one of (i) inserting amino acid residues, (ii) deleting amino acid residues, and (iii) changing the identity of amino acid residues in the input sequence.
[0165] Item 19: The method of any one of items 1 to 18, further comprising generating a fixed-length representation of the input sequence, and applying the trained protein design computational model to generate an output sequence by at least modifying the fixed-length representation of the input sequence.
[0166] Item 20: The method according to Item 19, wherein the fixed-length representation of the input sequence is generated by at least aligning each amino acid residue in the input sequence to a fixed set of structural roles such that each amino acid residue in the input sequence is assigned an integer position corresponding to the structural role of the amino acid residue, and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue having a corresponding structural role.
[0167] Item 21: The method of any one of items 1 to 20, wherein the difference between the one or more generated output sequences and the plurality of noisy sample sequences is quantified by one or more of an antibody similarity metric, an edit distance, and a naturalness metric.
[0168] Item 22: A computer-implemented method including: identifying an input sequence having a plurality of amino acid residues; generating a noisy embedding of the input sequence by at least adding noise to the input sequence; modifying the noisy embedding of the input sequence by applying a protein design computational model trained to at least approximate a data distribution of protein sequences exhibiting one or more desirable properties, wherein the protein design computational model modifies the noisy embedding of the input sequence to increase the likelihood that the resulting modified noisy embedding is present in the data distribution; and generating an output sequence by denoising at least the modified noisy embedding generated by the protein design computational model.
[0169] Item 23: The method of item 22, further including: encoding the input sequence to generate an embedding of the input sequence; generating a noisy embedding of the input sequence by at least adding noise to the embedding of the input sequence; and generating an output sequence by decoding a denoised embedding generated by denoising the modified noisy embedding.
[0170] Item 24: The method of item 23, wherein the input sequence is encoded by generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
[0171] Item 25: The method of items 23 or 24, wherein the input sequence is encoded by at least generating one or more tokens that encode the relative position of each amino acid residue in the input sequence.
[0172] Item 26: The method of items 23 or 24, wherein the input sequence is encoded by generating, for at least one amino acid residue in the input sequence, one or more structural tokens that identify at least one or more adjacent amino acid residues in three-dimensional space.
[0173] Item 27: The method of any one of items 22 to 26, wherein modifying the noisy embedding includes applying an energy-based model (EBM) trained to approximate the data distribution to modify the noisy embedding of the input sequence and generate a first modified noisy embedding; applying an energy function parameterized by the energy-based model (EBM) to modify the noisy embedding of the input sequence and generate a second modified noisy embedding; applying the energy-based model (EBM) to determine a first energy value of the first modified noisy embedding and a second energy value of the second modified noisy embedding; and applying the energy-based model (EBM) to further modify the first modified noisy embedding instead of the second modified noisy embedding based at least on the first energy value and the second energy value.
[0174] Item 28: The method of item 27, in which an energy-based model (EBM) is applied to further modify the first modified noisy embedding until one or more criteria are met.
[0175] Item 29: The method of item 28, wherein the one or more criteria include at least one of (i) performing a threshold number of iterations of modifications to the noisy embedding of the input sequence, and (ii) a first energy value of the first modified noisy embedding that satisfies one or more thresholds.
[0176] Item 30: The method of item 27 or 28, wherein an energy-based model (EBM) is applied to further modify the first modified noisy embedding in place of the second modified noisy embedding based on at least the first energy value and the second energy value indicating that the first modified noisy embedding has a higher likelihood in the data distribution than the second modified noisy embedding.
[0177] Item 31: The method of any one of items 27 to 30, wherein an energy-based model (EBM) is applied to further modify the first modified noisy embedding instead of the second modified noisy embedding based on at least the first energy value and the second energy value indicating that the first modified noisy embedding samples from a higher density region of the data distribution than the second modified noisy embedding.
[0178] Item 32: The method of any one of Items 22 to 31, further comprising generating a fixed-length representation of the input sequence, and generating a noisy embedding of the input sequence based on at least the fixed-length representation of the input sequence.
[0179] Item 33: The method of Item 32, wherein the fixed-length representation of the input sequence is generated by at least aligning each amino acid residue in the input sequence to a fixed set of structural roles such that each amino acid residue in the input sequence is assigned an integer position corresponding to the structural role of the amino acid residue, and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue having a corresponding structural role.
[0180] Item 34: The method of items 32 or 33, wherein the protein design computational model modifies the noisy embedding of the input sequence by at least one of changing the identity of one or more amino acid residues of the input sequence, deleting an amino acid residue occupying a position in the fixed length representation of the input sequence by at least replacing the amino acid residue with a gap character, and inserting an amino acid residue into a position in the fixed length representation of the input sequence by at least replacing the gap residue occupying the position with the amino acid residue.
[0181] Item 35: The method of any one of items 22 to 34, wherein the one or more desirable properties include at least one of expression, affinity, specificity, stability, non-immunogenicity, humanity, absence of self-association, and lack of chemical liability.
[0182] Item 36: The method of any one of Items 22 to 35, wherein the input sequence is a known protein sequence or a noise sequence containing a random sequence of amino acid residues.
[0183] Item 37: A system comprising at least one data processor and at least one memory storing instructions that, when executed by the at least one data processor, result in operations including the method of any one of items 1 to 21 or the method of any one of items 22 to 36.
[0184] Item 38: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including the method of any one of items 1 to 21 or the method of any one of items 22 to 36.
[0185] 9 depicts a block diagram illustrating an example of a computing system 900, according to some exemplary embodiments. Referring to FIGS. 1-9, the computing system 900 may be used to implement the protein design engine 110, the analysis engine 120, the client device 130, and / or any components therein.
[0186] 9, computing system 900 may include a processor 910, a memory 920, a storage device 930, and an input / output device 940. The processor 910, the memory 920, the storage device 930, and the input / output device 940 may be interconnected via a system bus 950. The processor 910 is capable of processing instructions for execution within the computing system 900. Such executed instructions may implement one or more components, such as, for example, the protein design engine 110, the analysis engine 120, the client device 130, etc. In some exemplary embodiments, the processor 910 may be a single-threaded processor. Alternatively, the processor 910 may be a multi-threaded processor. The processor 910 is capable of processing instructions stored in the memory 920 and / or in the storage device 930 to display graphical information for a user interface provided via the input / output device 940.
[0187] Memory 920 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 900. Memory 920 may store, for example, data structures representing a configuration object database. Storage device 930 may provide persistent storage for computing system 900. Storage device 930 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. Input / output device 940 provides input / output operations for computing system 900. In some exemplary embodiments, input / output device 940 includes a keyboard and / or a pointing device. In various implementations, input / output device 940 includes a display device for displaying a graphical user interface.
[0188] According to some demonstrative embodiments, input / output devices 940 may provide input / output operations for network devices. For example, input / output devices 940 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0189] In some exemplary embodiments, computing system 900 can be used to execute various interactive computer software applications that can be used for organizing, analyzing, and / or storing various types of data. Alternatively, computing system 900 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications can include various add-in functions or can be standalone computing products and / or features. When active within an application, functionality can be used to generate a user interface that is provided via input / output devices 940. The user interface can be generated by computing system 900 and presented to a user (e.g., on a computer screen monitor, etc.).
[0190] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0191] These computer programs, sometimes referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.
[0192] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device, such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as, for example, a mouse or trackball, by which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, tactile feedback, etc., and input from the user may be received in any form, including acoustic input, voice input, and tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, etc.
[0193] In the above specification and claims, phrases such as "at least one of" or "one or more of" may appear before a list of consecutive elements or features. The term "and / or" may also be used in listings of two or more elements or features. Unless otherwise implicitly or explicitly stated by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, C," "one or more of A, B, C," and "A, B, and / or C" are intended to mean "A only, B only, C only, A and B together, A and C together, B and C together, or A, B and C together," respectively. Use of the term "based on" above and in the claims means "based at least in part on," and implies that unrecited features or elements are also permitted.
[0194] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the above description do not necessarily represent all implementations of the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. While several variations have been detailed above, other modifications and additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the above-described embodiments may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features disclosed above. Additionally, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order shown or sequential order to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
1. 1. A computer-implemented method comprising: generating a first training set including a plurality of noisy sample sequences, wherein each noisy sample sequence in the first training set is generated by at least adding noise to a corresponding sample sequence from a first data distribution; Protein design computational models, at least applying the protein design computational model to generate one or more output sequences; and adjusting the protein design computational model to reduce differences between the one or more generated output sequences and the plurality of noisy sample sequences of the first training set; and applying the trained protein design computational model to generate an output sequence by modifying at least an input sequence; A method comprising:
2. The method of claim 1 , wherein the protein design computational model comprises a first energy-based model (EBM).
3. 3. The method of claim 2, wherein the training the protein design computational model comprises adjusting a plurality of parameters of the first energy-based model that parameterize an energy function of the first energy-based model.
4. 4. The method of claim 3, wherein the plurality of parameters are adjusted such that an energy value determined by the energy function corresponds to a likelihood of the one or more generated output sequences within the first data distribution.
5. 5. The method of claim 3, wherein the plurality of parameters are adjusted such that the energy function outputs a lower energy value for a first generated output sequence that is similar to the plurality of noisy samples of the first training set than for a second generated output sequence that is dissimilar to the plurality of noisy samples of the first training set.
6. training the protein design computational model, applying the first energy-based model with a first adjustment to generate a first modified sequence; applying the first energy-based model with the second adjustment to generate a second modified sequence; and upon determining that the first modified sequence is more similar to the plurality of noisy samples of the first training set than the second modified sequence, further modifying the first energy-based model with the first adjustment instead of the second adjustment; The method according to any one of claims 3 to 5, comprising:
7. 7. The method of claim 6, wherein the first energy-based model is further adjusted until one or more criteria are met, the one or more criteria including at least one of: (i) performing a threshold number of iterations of adjustments to the first energy-based model; and (ii) the second modified sequence exhibiting a threshold similarity to the plurality of noisy samples of the first training set.
8. The method of any one of claims 2 to 7, wherein the protein design computational model further comprises a second energy-based model (EBM).
9. generating a second training set comprising a plurality of sample sequences from a second data distribution; determining a first adjustment to the first energy-based model that reduces a first difference between a first output sequence produced by the first energy-based model and the plurality of noisy sample sequences of the first training set; determining a second adjustment to the second energy-based model that reduces a second difference between a second output sequence produced by the second energy-based model and the plurality of sample sequences of the second data distribution; and training the first energy-based model by applying to the first energy-based model a third adjustment determined based on the first adjustment and the second adjustment; The method of claim 8 further comprising:
10. The method of claim 9 , wherein the third adjustment corresponds to a sum or a weighted sum of the first adjustment and the second adjustment.
11. encoding each sample sequence from the first data distribution to generate an embedding for each sample sequence; and generating the plurality of noisy sample sequences of the first training set by at least adding noise to the embedding of each sample sequence; The method of any one of claims 1 to 10, further comprising:
12. The method of claim 11 , wherein each sample sequence from the first data distribution is encoded by augmenting it with additional information.
13. 13. The method of claim 12, wherein the additional information comprises structural information that identifies, for each constituent amino acid residue, one or more adjacent amino acid residues in three-dimensional space.
14. The trained computational model for protein design comprises at least generating a noisy input sequence by adding noise to at least said input sequence; applying an energy-based model to generate a noisy output sequence by at least modifying the noisy input sequence based at least on an energy function of the energy-based model; and generating the output array by denoising at least the modified noisy output array generated by the energy-based model; The method of any one of claims 1 to 13, wherein the output array is generated by:
15. The trained computational model for protein design comprises at least generating an embedding of the input sequence by encoding at least the input sequence; generating a noisy embedding of the input sequence by adding at least noise to the embedding of the input sequence; applying an energy-based model to at least modify the noisy embedding of the input sequence based at least on an energy function of the energy-based model to generate a modified noisy embedding; denoising the noisy embedding to generate a denoised embedding; and generating the output array by at least denoising the noisy embedding; The method of any one of claims 1 to 14, wherein the output array is generated by:
16. The embedding of the input sequence includes at least 16. The method of claim 15, wherein the input sequence is generated by generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
17. The embedding of the input sequence includes at least 17. The method of claim 15 or 16, wherein the structural tokens are generated by generating, for at least one amino acid residue of the input sequence, one or more structural tokens that identify one or more adjacent amino acid residues in three-dimensional space.
18. 18. The method of any one of claims 1 to 17, wherein the trained computational protein design model modifies the input sequence by at least one of: (i) inserting amino acid residues, (ii) deleting amino acid residues, and (iii) changing the identity of amino acid residues in the input sequence.
19. generating a fixed length representation of the input array; and applying the trained protein design computational model to generate the output sequence by at least modifying the fixed-length representation of the input sequence; The method of any one of claims 1 to 18, further comprising:
20. The fixed length representation of the input sequence comprises at least: aligning each amino acid residue of the input sequence to a fixed set of structural roles such that each amino acid residue of the input sequence is assigned an integer position corresponding to the structural role of the amino acid residue; and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue with a corresponding structural role; 20. The method of claim 19, wherein the compound is produced by
21. 21. The method of any one of claims 1 to 20, wherein the difference between the one or more generated output sequences and the plurality of noisy sample sequences is quantified by one or more of an antibody similarity metric, an edit distance, and a naturalness metric.
22. 1. A computer-implemented method comprising: identifying an input sequence having a plurality of amino acid residues; generating a noisy embedding of the input sequence by at least adding noise to the input sequence; modifying the noisy embedding of the input sequence by applying a protein design computational model trained to approximate a data distribution of protein sequences exhibiting at least one or more desirable properties, wherein the protein design computational model modifies the noisy embedding of the input sequence to increase the likelihood that the resulting modified noisy embedding will be present in the data distribution; and generating an output sequence by at least denoising the modified noisy embedding generated by the protein design computational model; 11. A computer-implemented method comprising:
23. encoding the input sequence to generate an embedding of the input sequence; generating the noisy embedding of the input sequence by at least adding noise to the embedding of the input sequence; and generating the output array by decoding a denoised embedding produced by the denoising of the modified noisy embedding; 23. The method of claim 22, further comprising:
24. 24. The method of claim 23, wherein the input sequence is encoded by generating, for each amino acid residue in the input sequence, a token that encodes the identity of the amino acid residue.
25. 25. The method of claim 23 or 24, wherein the input sequence is encoded by at least generating one or more tokens that encode the relative position of each amino acid residue in the input sequence.
26. 25. The method of claim 23 or 24, wherein the input sequence is encoded by generating, for at least one amino acid residue in the input sequence, one or more structural tokens that identify at least one or more adjacent amino acid residues in three-dimensional space.
27. modifying the noisy embedding, modifying the noisy embedding of the input sequence and applying an energy-based model (EBM) trained to approximate the data distribution to generate a first modified noisy embedding; modifying the noisy embedding of the input sequence and applying the energy-based model (EBM) to generate a second modified noisy embedding; applying an energy function parameterized by an energy-based model (EBM) to determine a first energy value of the first modified noisy embedding and a second energy value of the second modified noisy embedding; and applying the energy-based model (EBM) to further modify the first modified noisy embedding instead of the second modified noisy embedding based at least on the first energy value and the second energy value; The method of any one of claims 22 to 26, comprising:
28. 28. The method of claim 27, wherein the energy-based model (EBM) is applied to further modify the first modified noisy embedding until one or more criteria are met.
29. 29. The method of claim 28, wherein the one or more criteria include at least one of: (i) performing a threshold number of iterations of modification to the noisy embedding of the input sequence; and (ii) the first energy value of the first modified noisy embedding satisfying one or more thresholds.
30. 29. The method of claim 27 or 28, wherein the energy-based model (EBM) is applied to further modify the first modified noisy embedding instead of the second modified noisy embedding based at least on the first energy value and the second energy value indicating that the first modified noisy embedding has a higher likelihood in the data distribution than the second modified noisy embedding.
31. 31. The method of any one of claims 27 to 30, wherein the energy-based model (EBM) is applied to further modify the first modified noisy embedding instead of the second modified noisy embedding based at least on the first energy values and the second energy values indicating that the first modified noisy embedding samples from a higher density region of the data distribution than the second modified noisy embedding.
32. generating a fixed length representation of the input array; and generating the noisy embedding of the input sequence based on at least the fixed length representation of the input sequence; The method of any one of claims 22 to 31, further comprising:
33. The fixed length representation of the input sequence comprises at least: aligning each amino acid residue of the input sequence to a fixed set of structural roles such that each amino acid residue of the input sequence is assigned an integer position corresponding to the structural role of the amino acid residue; and inserting gap characters at one or more positions where the input sequence does not contain an amino acid residue with a corresponding structural role; 33. The method of claim 32, wherein the compound is produced by
34. the protein design computational model converts the noisy embedding of the input sequence into changing the identity of one or more amino acid residues in the input sequence; deleting amino acid residues occupying positions in said fixed length representation of said input sequence by at least replacing said amino acid residues with gap characters; and inserting an amino acid residue at a position within said fixed length representation of said input sequence by at least replacing a gap residue occupying said position with said amino acid residue; 34. The method of claim 32 or 33, wherein the method is modified by at least one of:
35. 35. The method of any one of claims 22-34, wherein the one or more desirable properties include at least one of expression, affinity, specificity, stability, non-immunogenicity, humanity, absence of self-association, and lack of chemical liability.
36. The method of any one of claims 22 to 35, wherein the input sequence is a known protein sequence or a noise sequence comprising a random sequence of amino acid residues.
37. 1. A system comprising: at least one data processor; at least one memory storing instructions which, when executed by said at least one data processor, result in operations comprising the method of any one of claims 1 to 21 or the method of any one of claims 22 to 36; A system comprising:
38. A non-transitory computer readable medium storing instructions which, when executed by at least one data processor, cause operations comprising the method of any one of claims 1 to 21 or the method of any one of claims 22 to 36.