Hybrid Protein Design
The hybrid protein design approach integrates sequence and structure calculation models to optimize amino acid residues and conformations, addressing the limitations of separate methods by enhancing the likelihood of achieving desired properties in protein design.
Patent Information
- Application Number
- JP2025500211
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-07
- Filing Date
- 2023-07-07
- Publication Date
- 2025-07-30
AI Technical Summary
Existing protein design methods struggle to efficiently generate sequences that exhibit both desired structural properties and functional capabilities, as they often overlook the interdependence of amino acid sequences and three-dimensional structures, leading to sub-optimal results.
A hybrid protein design approach that integrates protein sequence and structure design, using a protein sequence calculation model to inform a protein structure calculation model, narrowing the search space for amino acid residues and conformations to optimize the sequence and structure jointly, thereby reducing computational load while enhancing the likelihood of achieving desired properties.
This method generates protein sequences and structures that are more likely to exhibit binding affinity, stability, and other desired properties by reducing the search space through informed sequence and conformational exploration, improving the efficiency and effectiveness of protein design.
Smart Images

Figure 2025524582000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 359,167, filed on July 7, 2022, the entire content of which is incorporated herein by reference.
[0002] The subject matter described herein generally relates to protein design, and more specifically, to protein design techniques that integrate protein sequence design with protein structure design.
Background Art
[0003] Introduction Proteins are responsible for many essential cellular functions, including, for example, enzymatic reactions, molecular transport, regulation and execution of some biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, and cell - to - cell communication. A protein molecule can comprise one or more polypeptide chains, each of which comprises a plurality of amino acid residues linked to one another by peptide bonds. The primary structure of the molecule, which depends on the sequence of amino acid residues in the polypeptide chain that forms the protein molecule, can refer to the three - dimensional structure (or conformation) of the protein molecule. For example, the secondary structure of a protein molecule can be at least partially defined by the torsional angles (or dihedral angles) of the peptide bonds present in the backbone (or main chain) of the protein molecule, while the tertiary structure of the protein molecule can be defined by the folding of the polypeptide chain.
[0004] The function of a protein molecule can depend on the sequence of amino acids in the polypeptide chain that forms the protein molecule and the three-dimensional structure formed by the polypeptide chain. Therefore, one important objective of protein design is to construct a novel sequence of amino acid residues (such as antibodies, etc.) that exhibits specific desired properties, including the ability to adopt a specific three-dimensional structure. In the case of macromolecular drug discovery, the novel protein sequence can be designed to complement the three-dimensional structure of a target molecule (such as an antigen like a viral antigen, a tumor antigen, etc.) and fold into a three-dimensional structure that is stable enough to allow the corresponding protein molecule to bind to the target molecule.
Summary of the Invention
[0005] A system, method, and product including a computer program product are provided for hybrid protein design. In some exemplary embodiments, a system is provided that includes at least one processor and at least one memory. The at least one memory can include program code that, when executed by the at least one processor, provides an operation. The operation can include identifying a protein sequence calculation model and a protein structure calculation model. Applying the protein sequence calculation model to generate a plurality of proposed protein sequences based at least on an input protein sequence, identifying a set of possible amino acid residues for each position in at least a portion of an output protein sequence based at least on the plurality of proposed protein sequences, using the protein structure calculation model to apply the protein structure calculation model to select a possible amino acid residue from the set of possible amino acid residues for each position in at least a portion of the output protein sequence to generate a first protein structure having the output protein sequence, a method.
[0006] In another aspect, a method for hybrid protein design is provided. The method may include identifying a protein sequence calculation model and a protein structure calculation model. Applying the protein sequence calculation model to generate a plurality of proposed protein sequences based at least on an input protein sequence, identifying a set of possible amino acid residues for each position in at least a portion of an output protein sequence based at least on the plurality of proposed protein sequences, and using the protein structure calculation model to apply the protein structure calculation model to select a possible amino acid residue from the set of possible amino acid residues for inclusion in the output protein sequence for each position in at least a portion of the output protein sequence, thereby generating a first protein structure having the output protein sequence.
[0007] In another aspect, a computer program product including a non-transitory computer-readable medium storing instructions is provided. The instructions may cause an operation to be performed by at least one data processor. The operation may include identifying a protein sequence calculation model and a protein structure calculation model. Applying the protein sequence calculation model to generate a plurality of proposed protein sequences based at least on an input protein sequence, identifying a set of possible amino acid residues for each position in at least a portion of an output protein sequence based at least on the plurality of proposed protein sequences, and using the protein structure calculation model to apply the protein structure calculation model to select a possible amino acid residue from the set of possible amino acid residues for inclusion in the output protein sequence for each position in at least a portion of the output protein sequence, thereby generating a first protein structure having the output protein sequence.
[0008] In some variations, one or more features disclosed herein including the following features can optionally be included in any practicable combination.
[0009] In some variations, multiple proposed protein sequences can be aligned to generate a plurality of aligned protein sequences. For each position in at least a portion of the output protein sequence, a set of possible amino acid residues can be identified based on at least the plurality of aligned protein sequences.
[0010] In some variations, the multiple proposed protein sequences can be aligned by applying one or more of dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, and hidden Markov models.
[0011] In some variations, identifying a set of possible amino acid residues for each position in at least a portion of the output protein sequence can include identifying the first amino acid residue rather than the second amino acid residue for inclusion in the set of possible amino acid residues.
[0012] In some variations, identifying a set of possible amino acid residues for each position in at least a portion of the output protein sequence further includes determining a first frequency at which the first amino acid residue appears at positions across the plurality of proposed protein sequences generated by the protein sequence calculation model, determining a second frequency at which the second amino acid residue appears at positions across the plurality of proposed protein sequences generated by the protein sequence calculation model, and identifying the first amino acid residue rather than the second amino acid residue for inclusion in the set of possible amino acid residues for that position based on at least the first frequency and the second frequency.
[0013] In some variations, the first amino acid residue can be identified for inclusion in the set of amino acid residues based on a first frequency that meets at least one or more thresholds. The second amino acid residue can be identified for exclusion from the set of possible amino acid residues based on a second frequency that does not meet at least one or more thresholds.
[0014] In some variations, the identification of the set of possible amino acid residues for a position in the output protein sequence further comprises determining one or more thresholds based on at least one of the maximum, minimum, median, average, and mode of the frequencies at which each of the plurality of amino acid residues appears at the position across the plurality of proposed protein sequences generated by the protein sequence calculation model.
[0015] In some variations, the set of possible amino acid residues may include only a subset, rather than all, of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, selenocysteine, and pyrrolidine.
[0016] In some variations, the protein structure calculation model may generate the first protein structure by at least determining the identity and conformation of the amino acid residues occupying each position in at least a portion of the output protein sequence based on the energy of the first protein structure having at least the output protein sequence.
[0017] In some variations, the protein structure calculation model may determine the identity and conformation of the amino acid residues occupying each position in at least a portion of the output protein sequence by modifying at least one of the identity and conformation of the amino acid residues to minimize the energy of the first protein structure.
[0018] In some variations, the protein structure calculation model may modify at least one of the identity and conformation of the amino acid residue occupying the position by at least (i) changing the conformation of the amino acid residue occupying the one position or (ii) selecting a different possible amino acid residue for the position from the set of amino acid residues associated with the position.
[0019] In some variations, the protein structure calculation model determines at least the energy of a first protein structure having a first possible amino acid residue from a set of possible amino acid residues, determines the energy of a second protein structure having a second possible amino acid residue from the set of possible amino acid residues, and generates a first protein structure to include the first possible amino acid residue instead of the second possible amino acid residue based on at least the first energy being lower than the second energy, thereby determining the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence.
[0020] In some variations, the protein structure calculation model determines at least the energy of a first protein structure having a first conformation of a first possible amino acid residue, determines the energy of a second protein structure having a second conformation of the first possible amino acid residue, and generates a first protein structure to include the first conformation of the first possible amino acid residue instead of the second conformation of the first possible amino acid residue based on at least the third energy being lower than the fourth energy, thereby further determining the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence.
[0021] In some variations, a property analysis model can be applied to determine the properties of each protein sequence included in a plurality of proposed protein sequences. At least one protein sequence among the plurality of proposed protein sequences can be identified for exclusion based at least on the properties of each protein sequence. At least one protein sequence can be excluded from the plurality of proposed protein sequences before identifying a set of possible amino acid residues for each position in the output protein sequence based at least on the remaining plurality of proposed protein sequences.
[0022] In some variations, the first part of the second protein structure can be identified. The third protein structure can be generated by replacing at least the first part of the second protein structure with at least the second part of the first protein structure.
[0023] In some variations, the third protein sequence of the third protein structure generated to include the second part of the first protein structure and the third part of the second protein structure can be determined. A protein structure calculation model and / or different protein structure calculation models can be applied to determine at least a fourth protein structure having the third protein sequence based at least on the third protein sequence. A similarity metric can be determined to quantify the difference between the third protein structure and the fourth protein structure. The third protein sequence can be identified as a synthesis candidate based on a similarity metric that meets at least one or more thresholds.
[0024] In some variations, the second protein structure can be selected based on a second protein structure that exhibits at least one or more desired characteristics.
[0025] In some variations, the first part of the second protein structure can include the first antigen-binding site of a first antibody having the second protein structure, and the second part of the first protein structure can include the second antigen-binding site of a second antibody having the first protein structure.
[0026] In some variations, the first part of the second protein structure can include the first paratope of a first antibody having the second protein structure, and the second part of the first protein structure can include the second paratope of a second antibody having the first protein structure.
[0027] In some variations, the first portion of the second protein structure can include the first complementarity determining region (CDR) of a first antibody having the second protein structure, and the second portion of the first protein structure can include the second complementarity determining region (CDR) of a second antibody having the first protein structure.
[0028] Implementations of the subject matter can include, but are not limited to, methods according to the descriptions provided herein, and articles of manufacture that include a tangible, machine-readable medium operable to cause one or more machines (such as a computer) to perform one or more of the operations to implement one or more of the features described. Similarly, a computer system can be described that includes one or more processors and one or more memories coupled to the one or more processors. The memory can include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method according to one or more implementations of the subject matter can be implemented by one or more data processors in a single computing system or multiple computing systems. Such multiple computing systems can be connected through one or more connections, including, for example, connections through a network (such as the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of the multiple computing systems, etc., and can exchange data and / or instructions or other commands.
[0029] Details of one or more variations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described in this specification will be apparent from the description and drawings, and from the claims. Specific features of the subject matter of this disclosure are illustrated for purposes of example with respect to deep learning for modeling disease progression, but it should be readily understood that such features are not limiting. The claims that follow this disclosure define the scope of the protected subject matter.
Brief Description of the Drawings
[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate specific aspects of the subject matter disclosed in this specification and, together with the description, serve to explain some of the principles associated with the disclosed implementations.
[0031]
Figure 1
[0032]
Figure 2
[0033]
Figure 3A
[0034]
Figure 3B
[0035]
Figure 4A
[0036]
Figure 4B
[0037]
Figure 5
[0038]
Figure 6A
[0039]
Figure 6B
[0040]
Figure 7A
[0041]
Figure 7B
[0042]
Figure 8A
[0043]
Figure 8B
[0044]
Figure 8C
[0045]
Figure 9
[0046] In practical cases, like reference numerals indicate like structures, features, or elements.
DETAILED DESCRIPTION OF THE INVENTION
[0047] Protein design aims to identify novel protein sequences (e.g., sequences of amino acid residues) that exhibit certain desired properties, such as, for example, expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for the target molecule, lack of nonspecificity, stability (e.g., robustness to various environmental stresses such as conformational stability, thermodynamic stability, protease resistance, and / or the like), non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), lack of chemical liability (e.g., aspartic acid isomerization, oxidation, deamidation), developability, and / or the like. However, protein design is difficult because there are innumerable variations in at least the sequence and structure of proteins, and is particularly a resource-intensive task. For example, the sequence space containing possible substitutions of amino acid residues that can form a protein molecule is enormous (e.g., For the protein sequence of 4,170 amino acid residues in TIFF2025524582000002.tif, approximately Among those possible substitutions in TIFF2025524582000003.tif (4,170), few actually correspond to functional protein sequences. On the other hand, the conformational space occupied by the possible three-dimensional structures of protein molecules is similarly vast, even when limited to several discontinuous (not continuous) structural fluctuations. For example, For the protein sequence of 4,170 amino acid residues in TIFF2025524582000004.tif, assuming each amino acid residue is limited to one of three distinct geometric states (e.g., rotamers), there are approximately possible conformations in TIFF2025524582000005.tif (4,170). As used herein, the term "space" refers to the space of solutions to a given problem. In the context of protein sequence design, the aforementioned sequence space can represent the possible solutions related to constructing protein sequences (of a specified or unspecified length) using a set of amino acid residues such as those forming an antibody. In the case of protein structure design, the aforementioned conformational space can represent the possible solutions associated with determining the geometric states of the amino acid residues forming a protein sequence, including, for example, the geometric states of the backbone atoms and side-chain atoms of each amino acid residue forming the protein sequence.
[0048] Accordingly, a structurally agnostic prediction of a protein sequence can yield many proteins that cannot adopt a three-dimensional structure capable of binding to a target molecule. Predicting only the protein structure can provide many possible protein structures that can bind structurally to a target antigen but lack other desired properties (e.g., expression, specificity for the target molecule, lack of nonspecificity, stability, non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), lack of chemical liability, and / or the like). Since both the sequence space and the conformational space are very large and the predicted protein sequences lack a structure that complements the target molecule and / or other desired properties, many computationally predicted proteins may have little or no therapeutic value. Thus, a rational design approach that includes applying a protein structure calculation model to an input protein sequence only determines the three-dimensional structure adopted by the protein sequence, but the protein sequence itself may remain sub-optimal. For example, separately generated protein sequences (e.g., generated by a language-based protein sequence calculation model) may not be able to exhibit certain desired properties such as binding affinity or stability because they lack at least the recognition of the structural traits by the language-based protein sequence calculation model that contribute to these desired properties.
[0049] The present disclosure describes sub-optimal results brought about by independently determining the sequence and structure of protein molecules. The present system and method recognize that the three-dimensional structure and function of a protein molecule depend on the sequence of amino acid residues that form the protein molecule (e.g., the primary structure of the protein molecule). For example, the binding affinity between a protein molecule and a target molecule may depend on whether the primary structure of the protein molecule can adopt a three-dimensional structure (e.g., secondary and tertiary structures) that complements the three-dimensional structure of the target antigen. In some cases, the practical utility of a protein molecule as a therapeutic agent may further depend on the stability of its three-dimensional structure or a portion thereof (e.g., an antigen-binding fragment (Fab), etc.). Thus, determining a protein sequence that is more likely to exhibit a particular desired property may involve, for example, a joint search over sequence space and conformational space to identify permutations of amino acid residues that can adopt a particular three-dimensional structure. Accordingly, in some exemplary embodiments, the present disclosure provides a protein design engine that can execute a hybrid protein design workflow that integrates protein sequence design and protein structure design such that the generation of the sequence of amino acid residues that form a protein molecule is informed by the three-dimensional structure of the protein molecule. Nevertheless, a naive hybrid protein design approach that combines an exhaustive search of sequence space and conformational space is computationally intractable for protein sequences of significant length. Thus, as described in more detail below, the protein design engine can utilize a protein sequence generated by a protein sequence computational model, not as a single possible therapeutic protein, but as a sequence that can be used to improve or guide the sequence space and conformational space explored by a protein structure computational model when generating the sequence and three-dimensional structure of a protein molecule. By doing so, the computational load associated with the hybrid protein design workflow can be reduced while the resulting protein sequence may be more likely to exhibit one or more desired properties than a sequence generated by a protein sequence model operating without structural awareness.
[0050] In some exemplary embodiments, the protein design engine may execute a hybrid protein design workflow that refines the sequence space and conformational space searched by the protein structure calculation model to generate the sequence and three-dimensional structure of a protein molecule based at least on the protein sequence generated by the protein sequence calculation model. For example, in some cases, the protein design engine can apply a protein sequence calculation model to generate a plurality of proposed protein sequences based at least on the input protein sequence. The protein sequence calculation model may involve, for example, a language-based protein sequence calculation model. In this context, a language-based protein sequence calculation model can refer to a machine learning model (or deep learning model) that applies one or more natural language processing (NLP) techniques to generate protein sequences. Examples of the machine learning architecture of a language-based protein sequence calculation model include autoencoders, transformers, long short-term memory networks, recurrent neural networks, and the like.
[0051] In some cases, the input protein sequence can be selected as a basis for generating a plurality of proposed protein sequences, at least because the input protein sequence exhibits one or more desired characteristics. The protein design engine can narrow the sequence space and conformational space searched by the protein structure calculation model by at least identifying a set of possible amino acid residues for each position in at least a portion of the output protein sequence based at least on the plurality of proposed protein sequences. As described in more detail below, the protein design engine can utilize the set of possible amino acid residues for each position in the output protein sequence by using various methods to exclude at least one possible amino acid residue at one or more positions of the output protein sequence to reduce the sequence space and conformational space searched by the protein structure calculation model. By doing so, the amount of possible substitutions of amino acid residues that form the output protein sequences and their corresponding conformations can be reduced.
[0052] In some cases, the protein design engine can reduce the sequence space and conformational space explored by the protein structure calculation model by applying a property analysis model. The property analysis model of the present disclosure can determine the properties of each protein sequence of a plurality of proposed protein sequences generated by a protein sequence calculation model. In some scenarios, based on those properties, at least one protein sequence can be excluded from the plurality of proposed protein sequences. Then, based on the remaining plurality of proposed protein sequences, by simply identifying the set of possible amino acid residues for each position in the output protein sequence, the amount of possible amino acid residues for one or more positions in the output protein sequence can be reduced. By doing so, the sequence space and conformational space explored by the protein structure calculation model can be further narrowed.
[0053] Furthermore, in some cases, the plurality of proposed protein sequences can be aligned to identify the amino acid residues that appear at each position across the aligned plurality of proposed protein sequences. For example, in some cases, the plurality of proposed protein sequences generated by a protein sequence calculation model can be aligned by applying one or more alignment techniques such as dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, hidden Markov models, etc.
[0054] In some exemplary embodiments, the protein design engine may determine, for each position in at least a portion of the output protein sequence, a set of possible amino acid residues that includes the first amino acid residue rather than the second amino acid residue. For example, in some cases, the set of possible amino acid residues for a particular position in the output protein sequence may be identified as including the first amino acid residue based at least on the first frequency of the first amino acid residue that appears at the position across a plurality of proposed protein sequences generated by the protein sequence calculation model. Alternatively and / or additionally, at least based on the second frequency of the second amino acid residue that appears at the position across a plurality of proposed protein sequences generated by the protein sequence calculation model, a set of possible amino acid residues for the position can be identified to exclude the second amino acid residue. Thus, in some cases, the set of possible amino acid residues for a position may be determined to include only a portion, rather than all, of the possible amino acid residues that can form a protein molecule. For example, if the output protein sequence corresponds to an antibody, the set of amino acid residues for a position may include only a portion, rather than all, of the amino acid residues that can form the antibody, and may include, for example, alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, and the like.
[0055] In some exemplary embodiments, the protein design engine may generate a first protein structure having an output protein sequence by at least applying a protein structure calculation model to select an amino acid residue from a set of possible amino acid residues for inclusion in the output protein sequence for each position in at least a portion of the output protein sequence. For example, in some cases, the protein structure calculation model may generate the first protein structure by at least determining, based on one or more criteria, a plurality of amino acid residues for inclusion in the output protein sequence and the corresponding conformations (or spatial arrangements) of the plurality of amino acid residues (e.g., the backbone and side chain atoms of each amino acid residue). As described, each position in the output protein sequence may, in some cases, be associated with a set of possible amino acid residues that includes a subset of all amino acid residues capable of forming an antibody. Thus, the selection of amino acid residues forming the output protein sequence may be limited to those included in each set of possible amino acid residues. In some cases, the selection of amino acid residues forming the output protein sequence may be limited when one or more segments of the output protein sequence are designated for conservation. In this regard, conservation of a segment of the input protein sequence may involve preventing any insertions, deletions, or substitutions from affecting that segment when generating the output protein sequence. Thus, it should be appreciated that the identity of the amino acid residues in the conserved segment of the output protein sequence remains unchanged from the identity of the amino acid residues occupying the same position in the input protein sequence. That is, the same amino acid residues forming the conserved segment of the input protein sequence remain present in the conserved segment of the output protein sequence. Thus, for the positions of the conserved segments of the output protein sequence, the set of possible amino acid residues for that position may include a single amino acid residue occupying the same position in the input protein sequence.
[0056] In some cases, one or more criteria may include minimizing an energy function of a first protein structure having a plurality of amino acid residues selected for inclusion in the output protein sequence. Further, in some cases, the protein structure calculation model can generate a plurality of protein molecules, each having a different sequence of amino acid residues and / or conformation, before selecting one protein molecule that meets one or more criteria. The narrowing of the sequence space and conformation space explored by the protein structure calculation model can manifest as a dramatic reduction in the amount of possible protein molecules (e.g., having different combinations of amino acid residues and conformations) evaluated by the protein structure calculation model. More specifically, the protein structure calculation model can determine a first energy of a first protein molecule having a first plurality of amino acid residues and a first conformation. Further, the protein structure calculation model can generate a second protein molecule having a second plurality of amino acid residues and a second conformation by modifying at least one of the first plurality of amino acid residues and the first conformation. The protein structure calculation model can determine a second energy of the second protein molecule. If the first energy of the first protein molecule is lower than the second energy of the second protein molecule, the protein structure calculation engine can generate a first protein structure having the first plurality of amino acid residues and the first conformation (instead of the second plurality of amino acid residues and the second conformation). In some cases, the protein structure calculation model can determine the first energy and the second energy by applying an energy function. For example, the protein structure calculation model can apply an energy function based on one or more of initial quantum mechanics, density functional theory (DFT), semi-empirical methods, molecular mechanics force fields, statistical potentials, neural potentials, machine learning models (e.g., trained on structural data), etc.
[0057] In some exemplary embodiments, the protein structure calculation model may further generate a first protein structure by at least determining a first backbone conformation of a first backbone of the first protein structure. The first backbone of the first protein structure can be a chain of consecutive atoms formed by linking backbone atoms of each amino acid residue of the output protein sequence. In some cases, the backbone atoms of each amino acid residue are nitrogen (N) atoms, TIFF2025524582000006.tif3170-carbon( TIFF2025524582000007.tif4170) atoms and an array of atoms including a carboxyl carbon (C) atom. Thus, in some cases, the first backbone conformation of the first backbone may include the spatial arrangement of the backbone atoms of each amino acid residue selected for inclusion in the output protein sequence. Further, in some cases, the spatial arrangement of the backbone atoms is the translation of the first backbone, the rotation of the first backbone, and / or the torsion angle of one or more rotatable bonds formed by the backbone atoms in the first backbone (e.g., TIFF2025524582000008.tif3170-carbon( TIFF2025524582000009.tif4170) the torsion angle ψ of the rotatable bond between the atom and the carbonyl group, TIFF2025524582000010.tif3170-carbon( TIFF2025524582000011.tif4170) the torsion angle φ of the rotatable bond between the atom and the nitrogen (N) atom, the torsion angle ω of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom, etc.) and may be defined by one or more of them.
[0058] In some exemplary embodiments, the protein structure calculation model may determine the first backbone conformation of the first backbone of the first protein structure to have the same conformation as at least a portion of the second backbone of the second protein structure. In some cases, the second protein structure may associate with a protein sequence having one or more desired properties such as expression, binding affinity for a target molecule, specificity for a target molecule, lack of nonspecificity, stability, non-immunogenicity, humanicity, lack of self-association (or non-aggregation). For example, in some cases, the protein sequence may be the input protein sequence on which the protein sequence calculation model was based to generate a plurality of proposed protein sequences or a third protein sequence. Further, in some cases, the second backbone of the second protein structure may represent the second backbone conformation of the second protein structure in the unbound state or the third backbone conformation of the second protein structure bound to the target molecule.
[0059] In some exemplary embodiments, instead of determining that the first backbone conformation of the first backbone has the same conformation as at least a portion of the second backbone of the second protein structure, the protein structure calculation model may determine the first backbone conformation of the first backbone by at least determining the geometric state of the first backbone. In some cases, the geometric state of the first backbone may be defined by the translation and / or rotation of the backbone atoms included in each amino acid residue in the output protein sequence such that the protein structure calculation model determines the geometric state of the first backbone by at least determining the translation and / or rotation of the backbone atoms. Alternatively and / or additionally, the geometric state of the first backbone may be defined by the torsional angle of one or more rotatable bonds formed by the backbone atoms in each amino acid residue in the output protein sequence. Thus, in some cases, the protein structure calculation model can determine the geometric state of the first backbone by at least determining the torsional angle of one or more rotatable bonds formed by the backbone atoms.
[0060] In some exemplary embodiments, upon determining a first backbone conformation of a first backbone of a first protein structure, the protein design computational model may further generate the first protein structure by at least determining, for each position in at least a portion of the first backbone having the first backbone conformation, an amino acid residue selected to be included at the corresponding position in the output protein sequence. Further, the protein structure computational model may generate the first protein structure by at least determining, for each position in at least a portion of the first backbone having the first backbone conformation, a side-chain conformation of the amino acid residue selected to be included at the corresponding position in the output protein sequence. For example, in some cases, the side-chain conformation of the amino acid residue may be determined based at least on the first backbone conformation of the first backbone of the first protein structure. Further, the protein structure computational model may determine the side-chain conformation of the amino acid residue selected to be included in the output protein sequence by at least determining the geometric state of the side chain. For example, in some cases, the geometric state of the side chain of the amino acid residue may be determined by finally selecting a rotamer from a plurality of possible rotamers each including a different combination of torsional angles of rotatable bonds formed by side-chain atoms in the amino acid residue, or the geometric state of the side chain may be determined by determining one or more of the translational, rotational, and / or torsional angles of the side-chain atoms in the amino acid residue.
[0061] In some exemplary embodiments, the protein design engine can verify a first protein structure generated by a protein structure calculation model by at least applying different protein structure calculation models to determine at least a second protein structure having the output protein sequence based at least on the output protein sequence. Further, the protein design engine can verify the first protein structure generated by the protein structure calculation model by determining a similarity metric (e.g., root mean square deviation (RMSD), etc.) that quantifies the difference between the first protein structure and the second protein structure. If the similarity metric meets one or more thresholds, the protein design engine can verify the first protein structure and identify the corresponding output protein sequence as a candidate for synthesis, in vitro measurement, in vivo characterization, etc.
[0062] In some exemplary embodiments, the protein design engine may generate a first protein structure such that one or more portions of the first protein structure can be used as a donor structure for transplantation onto corresponding portions of a second protein structure. For example, in some cases, the second protein structure may be selected such that at least the second protein structure exhibits one or more desired properties such as binding affinity for a target molecule, specificity for a target molecule, lack of nonspecificity, stability, non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), etc. Thus, in some cases, the protein design engine can identify a first portion of the second protein structure before generating a third protein structure by at least substituting the first portion of the second protein structure with a second portion of the first protein structure. In some cases, the protein design engine can verify the third protein structure based at least on a similarity metric (e.g., root mean square deviation (RMSD), etc.) that quantifies the difference between the third protein structure and at least a fourth protein structure generated by a protein structure computational model and / or a different protein structure computational model based on the third protein sequence of the third protein structure. When the first protein structure and the second protein structure are antibodies, the first portion of the second protein structure may include a first antigen-binding site (e.g., a first paratope, a first complementarity-determining region (CDR), etc.), and the second portion of the first protein structure may include a second antigen-binding site (e.g., a second paratope, a second complementarity-determining region (CDR), etc.). Thus, by transplanting the second portion of the first protein structure onto the second protein structure, an antibody can be generated whose sequence and structure combine the sequences and structures of multiple antibodies.
[0063] FIG. 1 shows a system diagram illustrating an example of an analysis system 100 according to some exemplary embodiments. Referring to FIG. 1, the protein design system 100 can include a protein design engine 110, a molecular analysis engine 120, and a client device 130. As shown in FIG. 1, the protein design system 100, the analysis engine 120, and the client device 130 may be communicatively coupled via a network 140. The client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, and the like. The network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and the like.
[0064] In some exemplary embodiments, the protein design engine 110 applies a protein sequence calculation model 113 to generate a plurality of proposed protein sequences 155 including, for example, a first proposed protein sequence 155a, a second proposed protein sequence 155b, a third proposed protein sequence 155c, etc., based on an input protein sequence 150a. In some cases, the protein design engine 110 can receive the input protein sequence 150a including an amino acid residue sequence from the client device 130. Further, in some cases, the input protein sequence 150a can be selected as a basis for generating a plurality of proposed protein sequences 155 because the input protein sequence 150a exhibits one or more desired properties including, for example, expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for the target molecule, lack of non-specificity, stability, non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), etc.
[0065] In some exemplary embodiments, the protein sequence calculation model 113 may include one or more machine learning models trained to generate a plurality of proposed protein sequences 155. For example, in some cases, one or more machine learning models may generate each protein sequence of the plurality of proposed protein sequences 150 by at least sampling a data distribution learned by the one or more machine learning models during training, based on the input protein sequence 155a. That is, each sampling of the data distribution may correspond to a single sampling iteration that generates a single protein sequence of the plurality of proposed protein sequences 155. For example, the first sampling of the data distribution may generate a first proposed protein sequence 155a, and the second sampling of the data distribution may generate a second proposed protein sequence 155b. In some cases, one or more machine learning models may be trained based on various known protein sequences, including protein sequences known to exhibit a particular function as well as protein sequences having no known function. Thus, one or more machine learning models can be trained to learn a data distribution corresponding to a reduced-dimensional representation of the sequence of amino acid residues that form the known protein sequences.
[0066] In some cases, one or more machine learning models can include an autoencoder (e.g., a denoising autoencoder (DAE), etc.), in which case the one or more machine learning models can learn the data distribution by at least learning to generate an encoding of the input protein sequence, and then be decoded to form an output protein sequence that minimally differs from the input protein sequence. During inference, the data distribution associated with the trained autoencoder can be sampled, for example, by encoding the input protein sequence 150a before decoding an intermediate sequence having at least one of damage (e.g., insertion, deletion, and / or modification of amino acid residues) and length change with respect to the input protein sequence 150a. Further, the sampling of the data distribution can be guided by the characteristics of the intermediate sequence. For example, in some cases, the intermediate sequence sampled from the data distribution can be subjected to a characteristic analysis (e.g., by a computational function prediction model), and can be decoded if it is determined that the intermediate sequence exhibits one or more desired characteristics that may also be present in the input protein sequence 150a. If the intermediate sequence does not exhibit one or more desired characteristics, another intermediate sequence can be generated by sampling the data distribution before also subjecting that intermediate sequence to the characteristic analysis. Thus, the intermediate sequences decoded to generate each of the plurality of proposed protein sequences 155 can differ from the input protein sequence 150a, but still exhibit the same (or similar) desired characteristics as the input protein sequence 150a. For example, the first proposed protein sequence 155a, the second proposed protein sequence 155b, and the third proposed protein sequence 155c can differ from the input protein sequence 150a, but retain at least some of the desired characteristics of the input protein sequence 150a.
[0067] In some exemplary embodiments, the protein design engine 110 can sample from a data distribution using various sampling techniques, such as, for example, Markov Chain Monte Carlo (MCMC), importance sampling (IS), rejection sampling, Metropolis-Hastings, Gibbs sampling, slice sampling, exact sampling, etc. Further, as described above, each sampling of the data distribution can correspond to a single sampling iteration that generates one of the plurality of proposed protein sequences 155. In some cases, the protein design engine 110 may continue to apply the protein sequence calculation model 113 to sample the data distribution until one or more conditions are met. For example, in some cases, the protein design engine 110 may continue to apply the protein sequence calculation model 113 to sample the data distribution until the plurality of proposed protein sequences 155 generated by the protein sequence calculation model 113 includes a threshold amount of proposed protein sequences. Alternatively and / or additionally, each of the plurality of proposed protein sequences 155 generated by the protein sequence calculation model 113 can undergo molecular analysis (e.g., by the molecular analysis engine 120 applying the property analysis model 125) to determine one or more properties of each proposed protein sequence. In those cases, the protein design engine 110 can continue to apply the protein sequence calculation model 113 to generate additional proposed protein sequences until the plurality of proposed protein sequences 155 generated by the protein sequence calculation model 113 includes a threshold amount of proposed protein sequences whose properties meet one or more thresholds. Thus, in some cases, the plurality of proposed protein sequences 155 can exclude one or more proposed protein sequences whose properties (e.g., as determined by the property analysis model 125) do not meet one or more thresholds.
[0068] In some exemplary embodiments, the protein design engine 110 can perform variable-length sampling in which a plurality of proposed protein sequences 155 undergoing further analysis have different lengths or amounts of constituent amino acid residues. For example, in some cases, the first proposed sequence 155a may have a different length than the second proposed sequence 155b and / or the third proposed sequence 155c (or may be formed from a different amount of amino acid residues). Alternatively, in some cases, the protein design engine 110 may perform fixed-length sampling in which a plurality of proposed protein sequences 155, including the first proposed sequence 155a, the second proposed sequence 155b, and the third proposed sequence 155c, have the same length (or are formed from the same amount of amino acid residues). In some cases, the protein design engine 110 can perform fixed-length sampling by at least training the protein sequence calculation model 113 based on a training data set that includes protein sequences of the same length. By doing so, the data distribution learned by the protein sequence calculation model 113 can be populated by the encoding of protein sequences that decode to protein sequences of the same length. Accordingly, the plurality of proposed protein sequences 155 generated by applying the protein sequence calculation model 113 can also have the same length (or the same amount of constituent amino acid residues). If the protein sequence calculation model 113 is trained based on protein sequences of different lengths, the protein sequence calculation model 113 can generate protein sequences having different lengths (or different amounts of constituent amino acid residues). In those cases, the protein design engine 110 can perform fixed-length sampling by at least excluding from the plurality of proposed protein sequences 155 any protein sequences generated by the protein sequence calculation model 113 that do not have a particular length (or a particular amount of constituent amino acid residues).
[0069] Referring back to FIG. 1, in some exemplary embodiments, the protein design engine 110 may include an array analyzer 115 that determines a set of possible amino acid residues 160 for each position of the output protein sequence 150b based on a plurality of proposed protein sequences 155 generated at least by the protein sequence calculation model 113. As described above, when the protein design engine 110 performs fixed-length sampling, the plurality of proposed protein sequences 155 may include protein sequences having the same length (or the same number of constituent amino acid residues), but in the case of variable-length sampling, the plurality of proposed protein sequences 155 may have different lengths (or amounts of constituent amino acid residues). In some cases, the array analyzer 115 may align the plurality of proposed protein sequences 157 in order to identify the amino acid residues that appear at each position across the plurality of aligned proposed protein sequences 155. For example, in some cases, the array analyzer 115 may apply at least one or more of dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, hidden Markov models, etc. to the plurality of proposed protein sequences 157 to generate the plurality of aligned proposed protein sequences 155. As described in more detail below, the set of possible amino acid residues 160 for at least some of the positions within the output protein sequence 150b may include a subset of amino acid residues that includes only a portion, and not all, of the possible constituent amino acid residues of the protein molecule. If the output protein sequence 150b corresponds to an antibody, the set of possible amino acid residues 160 may include only a portion, and not all, of the amino acid residues such as alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, etc. that can form an antibody.
[0070] In some exemplary embodiments, the protein design engine 110 can generate a first protein structure 170 by applying a protein structure calculation model 117 to select, for each position in at least a portion of the output protein sequence 150b, an amino acid residue from a corresponding set of possible amino acid residues 160 for inclusion in the output protein sequence 150b. Optionally, the protein structure calculation model 117 can include a physics-based protein structure calculation model, a machine learning-based protein structure calculation model, and the like. As described in more detail below, the protein structure calculation model 117 can generate the first protein structure 170 by at least determining a first backbone conformation of a first backbone of the first protein structure 170. Further, the protein structure calculation model 117 can generate the first protein structure 170 by at least determining a side chain conformation of the amino acid residue selected for inclusion at a corresponding position in the output protein sequence 150b for each position in at least a portion of the first backbone having the first backbone conformation.
[0071] FIG. 2 shows a flowchart illustrating an example of a process 900 for hybrid protein design according to some exemplary embodiments. Referring to FIGS. 1 and 2, the process 900 can be executed by the protein design engine 110 to generate an output protein sequence 150b and a corresponding first protein structure 170 based at least on the input protein sequence 150a.
[0072] In 202, the protein design engine 110 can apply the protein sequence calculation model 113 to generate a plurality of proposed protein sequences based at least on the input protein sequence. In some exemplary embodiments, the protein sequence calculation model 113 can generate a plurality of proposed protein sequences 155 including, for example, a first proposed protein sequence 155a, a second proposed protein sequence 155b, a third proposed protein sequence 155c, etc. based on the input protein sequence 150a. Optionally, the protein sequence calculation model 113 can generate a plurality of proposed protein sequences 155 by at least sampling a data distribution populated by the encoding of various protein sequences. The data distribution can be learned by the protein sequence calculation model 113 by at least training the protein sequence calculation model 113 to encode known protein sequences including protein sequences with known specific functions and protein sequences without known functions such that the known protein sequences can be recovered from the encoding with minimal information loss (e.g., the difference between the decoded protein sequence and the original protein sequence), e.g., by decoding. Each sampling of the data distribution can cause the protein sequence calculation model 113 to generate a single proposed protein sequence different from the input protein sequence 150a by having, for example, at least one of damage (e.g., insertion, deletion, and / or modification of amino acid residues) and length change to the input protein sequence 150a. As described above, the sampling of the data distribution can be guided by the characteristics of the intermediate sequence such that the plurality of proposed sequences 155 generated from the intermediate sequence can exhibit certain desired characteristics that may, in some cases, include the same (or similar) characteristics as the input protein sequence 150a.
[0073] In 204, the protein design engine 110 can identify a set of possible amino acid residues for each position in at least a portion of the output protein sequence based on at least a plurality of proposed protein sequences. In some exemplary embodiments, the sequence analyzer 115 can determine a set of possible amino acid residues 160 for each position of the output protein sequence 150b based on at least a plurality of proposed protein sequences 155. As shown in FIG. 1, in some cases, the sequence analyzer 115 can apply one or more alignment techniques including, for example, dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, hidden Markov models, etc., to align the plurality of proposed protein sequences 155. By doing so, the sequence analyzer 115 can generate the aligned proposed sequences 157 before determining the set of possible amino acid residues 160 for each position of the output protein sequence 150b based on at least the aligned proposed sequences 157.
[0074] For further illustration, FIG. 6A shows a schematic diagram illustrating an example of an aligned proposed array 157 generated by an array analyzer 1115 that aligns at least a first proposed array 155a, a second proposed array 155b, and a third proposed array 155c. In some exemplary embodiments, when a plurality of proposed protein sequences 155 are aligned, the array analyzer 115 can identify the amino acid residues that appear at each position across the aligned proposed array 157. Further, in some cases, the array analyzer 115 can determine the frequency with which each amino acid residue appears at each position across the aligned proposed array 157. In the example shown in FIG. 6A, when the first proposed array 155a, the second proposed array 155b, and the third proposed array 155c are aligned, the array analyzer 115 can determine that the amino acid residue arginine (R) appears at the first position 600a of the first proposed array 155a, the amino acid residue lysine (K) appears at the first position 600a of the second proposed array 155b, and the amino acid residue glutamine (Q) appears at the first position 600a of the third proposed array 155c. The array analyzer 115 can also determine that the amino acid residue alanine (A) appears at the second position 600b of the first proposed array 155a, the second proposed array 155b, and the third proposed array 155c. Further, the array analyzer 115 can determine that the amino acid residue serine (S) appears at the third position 600c of the first proposed array 155a and the third proposed array 155c, and the amino acid residue threonine (T) appears at the third position 600c of the second proposed array 155b.
[0075] In some exemplary embodiments, the alignment analyzer 115 may determine amino acid residues in the output protein sequence 150b whose respective likelihoods of appearing meet one or more thresholds, based at least on the aligned proposed sequences 157. For example, in some cases, the alignment analyzer 115 may identify the amino acid residue that is most likely to occupy each position in at least a portion of the output protein sequence 150b, based at least on the amino acid residues occupying each position across the aligned proposed sequences 157. Thus, in some cases, the alignment analyzer 115 may generate a set of possible amino acid residues 160 for each position, to exclude at least some of the amino acid residues that may form the protein molecule. By doing so, the sequence space of possible permutations of amino acid residues from the protein structure calculation model 117 is reduced, and the output protein sequence 150b and corresponding protein structure 170 are determined during subsequent structure design.
[0076] In some exemplary embodiments, the array analyzer 115 can determine a set of possible amino acid residues 160 for each position in at least a portion of the output protein sequence 150b, based at least on the amino acid residues that appear at each position across the aligned proposed sequences 157. For example, the set of possible amino acid residues 160 for a particular position can include amino acid residues that are more likely to occupy the position within the output protein sequence 150b and can be identified to exclude amino acid residues that are less likely to occupy the position within the output protein sequence 150b. If the output protein sequence 150b is an antibody, the set of possible amino acid residues 160 for that position can include a subset of amino acid residues that can form an antibody. For example, if the output protein sequence 150b is an antibody, the array analyzer 115 can determine, based at least on the amino acid residues that appear at each position across the aligned proposed sequences 157, not all but some of the amino acids alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.
[0077] In some cases, one or more amino acid residues can be identified for inclusion in the set of possible amino acid residues 160 at a particular position of the output protein array 150b, based at least on the frequency with which each amino acid residue appears at that position. For example, if the frequency with which a first amino acid residue appears at that position across the aligned proposed sequence 157 indicates that the first amino acid residue is likely to occupy that position in the output protein array 150b, the first amino acid residue can be identified for inclusion in the set of possible amino acid residues 160 at that position. Conversely, if the frequency with which a second amino acid residue appears at that position across the aligned proposed sequence 157 indicates that the second amino acid residue is unlikely to occupy that position in the output protein array 150b, the second amino acid residue can be excluded from the set of possible amino acid residues at that position. Thus, in some cases, one or more amino acid residues can be identified for inclusion in the set of possible amino acid residues 160 at a position within the output protein array 157b if the frequency with which one or more amino acid residues appear at that position across the aligned proposed sequence 150 meets one or more thresholds. In some cases, the one or more thresholds can be determined based on the maximum, minimum, median, average, and mode frequencies with which each of a plurality of amino acid residues appear at positions across the aligned proposed sequence 157. For example, in some cases, if a first frequency with which a first amino acid residue appears at positions across the aligned proposed sequence 157 exceeds the median frequency with which each of a plurality of amino acid residues appear at positions across the aligned proposed sequence 157, the first amino acid residue can be identified for inclusion in the set of possible amino acid residues 160 at a particular position in the output protein array 150b. In contrast, if a second frequency with which a second amino acid residue appears at positions across the aligned proposed sequence 157 does not exceed the median frequency described above, the second amino acid residue can be excluded from the set of possible amino acid residues 160 at that position.
[0078] For further illustration, FIG. 6B shows a schematic diagram showing an example of a set of possible amino acid residues 160 generated for each position in at least a portion of the output protein sequence 150b. In the example shown in FIG. 6B, the first position 600a of the output protein sequence 150b can be associated with a first set of possible amino acid residues 160a that includes the amino acid residues arginine (R), lysine (K), glycine (Q), and glutamine (Q). The sequence analyzer 115 can select the amino acid residues arginine (R), lysine (K), glycine (Q), and glutamine (Q) for inclusion in the first set of possible amino acid residues 160a because, at least, the frequency at which the amino acid residues appear at the first position 160a across the aligned proposed sequences 157 meets one or more thresholds. On the other hand, the second position 600b of the output protein sequence 150b can be associated with a second set of possible amino acid residues 160b that includes the amino acid residue alanine (A), which can be selected for inclusion based at least on the frequency at which the amino acid residue alanine (A) appears at the second position 600b across the aligned proposed sequences 157 and meets one or more thresholds. Further, the third position 600c of the output protein sequence 150b can be associated with a third set of possible amino acid residues 160c that includes the amino acid residues serine (S) and threonine (T). The amino acid residues serine (S) and threonine (T) can be included in the third set of possible amino acid residues 600c at the third position 600c because, at least, the frequency at which these amino acid residues appear at the third position 160c across the aligned proposed sequences 157 meets one or more thresholds. It should be understood that amino acid residues that are not present in each set of possible amino acid residues 160 can be excluded because, at least, the frequency at which these amino acid residues appear at the corresponding positions across the aligned proposed sequences 157 does not meet one or more thresholds.
[0079] In 206, the protein design engine 110 can generate a protein structure having an output protein sequence by at least applying a protein structure calculation model 117 to select an amino acid residue from a corresponding set of possible amino acid residues for inclusion in the output protein sequence for each position in at least a portion of the output protein sequence. In some exemplary embodiments, the protein design engine 110 can apply the protein structure calculation model 117, and the protein structure calculation model can generate a first protein structure 170 having an output protein sequence 150b by at least selecting an amino acid residue from a corresponding set of possible amino acid residues 160 for inclusion in the output protein sequence 150b for each position in at least a portion of the output protein sequence 150b. The operation of the protein structure calculation model 117 with a set of possible amino acid residues 160 for each position that includes only a part and not all of the possible amino acid residues that can form a protein molecule reduces the computational complexity of the structural design performed by the protein structure calculation model 117. For example, the protein structure calculation model 117 can generate a first protein structure 170 having an output protein sequence 150b by at least evaluating various possible protein molecules each having a different amino acid residue substitution and conformation. Thus, to generate a first protein structure 170 having an output protein sequence 150b, the protein structure calculation model 117 can perform a co-search across the sequence space of possible substitutions of amino acid residues that form a protein molecule and the conformational space of the corresponding three-dimensional structure. This search computational load can be reduced by the protein structure calculation model 117 that utilizes the set of possible amino acid residues 160 for each position to dramatically reduce the amount of possible protein molecules evaluated to generate the output protein sequence 150b.
[0080] In some cases, the protein structure calculation model 117 may perform a joint search over sequence space and conformational space by at least evaluating the three-dimensional structures of different substitutions of amino acid residues formed by the possible amino acid residues in each set 160 of possible amino acid residues. For example, in some cases, the protein structure calculation model 117 can be applied to determine the energy of each three-dimensional protein structure in order to identify the permutation of amino acid residues that folds into a three-dimensional structure with the lowest energy. As described in more detail below, for each permutation of amino acid residues that forms the output protein sequence 150b, the protein structure calculation model 117 can determine the corresponding first protein structure 170 by at least determining the backbone conformation of the first protein structure 170 having that particular permutation of amino acid residues and / or the side-chain conformation of the constituent amino acid residues. For example, in some cases, the protein structure calculation model 117 may determine the geometric state of the backbone and / or side chains of the constituent amino acid residues of the first protein structure 170 with various degrees of freedom. In the case of the backbone conformation, the protein structure calculation model 117 can determine the geometric state of the backbone of the first protein structure 170 based at least in part on the backbone conformations of other protein structures related to one or more desired properties (e.g., expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for a target molecule, lack of non-specificity, stability, non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), etc.). On the other hand, the side-chain conformation of the amino acid residues can be determined by selecting one of a plurality of possible rotamers or by determining the torsional angles of one or more rotatable angles formed by the side-chain atoms of the amino acid residues. FIG. 6B shows an example of a structural design in which the first protein structure 170 generated by the protein structure calculation model 117 is in a bound state binding to the target molecule 650. However, it should be understood that the protein structure calculation model 117 can generate the first protein structure 170 in a bound or unbound state.
[0081] FIG. 3A shows a flowchart illustrating an example of a process 300 for hybrid protein design according to some exemplary embodiments. Referring to FIGS. 1-2 and FIG. 3A, process 300 can be executed by a protein design engine 110 to generate an output protein sequence 150b and a corresponding first protein structure 170 based at least on an input protein sequence 150a. Optionally, process 300 can perform operation 206 of process 200, in which a protein structure calculation model 117 selects constituent amino acid residues of the protein sequence and determines a corresponding protein structure by determining at least the backbone and side-chain conformations of the protein structure.
[0082] In 302, the protein structure calculation model 117 can select a plurality of amino acid residues that meet one or more criteria for inclusion in the protein sequence. In some exemplary embodiments, the protein design engine 110 can apply the protein structure calculation model 117 to select, for each position in at least a portion of the output protein sequence 150b, an amino acid residue from the corresponding set of possible amino acid residues 160 for inclusion in the output protein sequence 150b. Referring again to the example shown in FIG. 6B, the protein structure calculation model 117 can be applied, for example, to select the amino acid residue arginine (R) from the first set of possible amino acid residues 160a, the amino acid residue alanine (A) from the second set of possible amino acid residues 160b, and the amino acid residue serine (S) from the third set of possible amino acid residues 160c for inclusion in the corresponding position of the output protein sequence 150b. As described in more detail below, in some cases, to identify the amino acid residues to include in the output protein sequence 150b, the protein structure calculation model 117 can select a substitution of an amino acid residue that meets one or more criteria from different possible substitutions of amino acid residues formed from the amino acid residues included in the set of possible amino acid residues 160 for each position in at least a portion of the output protein sequence 150b. In some cases, the one or more criteria can include the energy of the first protein structure 170 having the output protein sequence 150b, which can be determined by applying an energy function. The fact that the possible substitutions of amino acid residues are formed from the set of possible amino acid residues 160 for each position in at least a portion of the output protein sequence 150b, instead of all possible amino acid residues that can form an antibody, for example, can dramatically reduce the computational load associated with identifying substitutions of amino acid residues for inclusion in the output protein sequence 150b such that the output protein sequence 150 and the corresponding first protein structure 170 exhibit certain desired properties.
[0083] In 304, the protein structure calculation model 117 can determine the backbone conformation of the protein structure having the protein sequence. In some exemplary embodiments, the protein structure calculation model 117 can further generate the first protein structure 170 by at least determining the backbone conformation of the first protein structure 170. Optionally, the backbone of the first protein structure 170 can include a continuous chain of atoms formed by linking a plurality of backbone atoms from each amino acid residue of the output protein sequence 150b. Further, optionally, the plurality of backbone atoms in each amino acid residue in the output protein sequence can be an array of atoms including a nitrogen (N) atom, an alpha-carbon ( TIFF2025524582000012.tif4170) atom, and a carboxyl carbon (C) atom. It should be understood that the same plurality of backbone atoms can be present in different amino acid residues. In the example of the output protein sequence 150b shown in FIG. 6B, for example, the backbone of the corresponding protein structure 170 can include the backbone atoms of each amino acid residue in the sequence including arginine (R), alanine (A), serine (S), glutamine (Q), aspartic acid (D), valine (V), asparagine (N), threonine (T), alanine (A), valine (V), and alanine (A).
[0084] As described above, the protein structure calculation model 117 can be applied to determine the backbone conformation of the first protein structure 170 with various degrees of freedom. For example, in some exemplary embodiments, the protein structure calculation model 117 is associated with one or more desired properties such as expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for the target molecule, lack of non-specificity, stability, non-immunogenicity, human-likeness, lack of self-association (or non-aggregation), etc. The first protein structure 170 can be determined based on the backbone of another protein structure that can associate with a protein sequence. In some cases, the other protein structure may have the input protein sequence 150a based on which the protein sequence calculation model 113 generated a plurality of proposed sequences 155. Alternatively, the other protein structure may be bound to another protein sequence different from the input protein sequence 150a and the output protein sequence 150b. For example, in some cases, the input protein sequence 150a used to generate a plurality of proposed sequences 155 for determining the output protein sequence 150b may be associated with a first desired property, and the other protein sequence associated with another protein structure that provides the backbone conformation of the first protein structure 170 may be associated with the same first desired property and / or a second desired property. Thus, in some cases, the protein structure calculation model 117 may determine that the first backbone conformation of the first protein structure 170 has the same conformation as at least a portion of the backbone of another protein structure so that the first protein structure 170 can exhibit the same (or similar) desired properties as the other protein structure. As described in more detail below, the protein design engine 110 can perform verification to determine whether the output protein sequence 150b folds into the first protein structure 170 having the backbone conformation of another protein structure.
[0085] In some exemplary embodiments, the protein design engine 110 can apply the protein structure calculation model 117 to determine the backbone conformation of a first protein structure 170 independent of another protein structure. Alternatively, the protein structure calculation model 117 can be applied to determine the geometric state of the backbone of the first protein structure 170 with various degrees of freedom. For example, in some cases, the protein structure calculation model 117 can determine the geometric state of the backbone of the first protein structure 170 by at least determining the translation and / or rotation of at least a portion of the backbone atoms forming the backbone of the first protein structure 170. In some cases, in addition to or instead of the translation and / or rotation of at least a portion of the backbone atoms in the backbone of the first protein structure 170, the protein structure calculation model 117 can determine the geometric state of the backbone of the first protein structure 170 by at least determining the torsion angle of one or more rotatable bonds formed by at least a portion of the backbone atoms. For example, in some cases, the geometric state of the backbone of the first protein structure 170 can be TIFF2025524582000013.tif3170-carbon( TIFF2025524582000014.tif4170) the torsion angle ψ of the rotatable bond between the atom and the carbonyl group, TIFF2025524582000015.tif3170-carbon( TIFF2025524582000016.tif4170) the torsion angle φ of the rotatable bond between the atom and the nitrogen (N) atom, the torsion angle ω of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom, etc., can be determined by determining one or more of them.
[0086] In 306, the protein structure calculation model 117 can generate a protein structure such that each position in at least a portion of the backbone includes an amino acid residue selected for inclusion at the corresponding position in the protein sequence. In some exemplary embodiments, the protein design engine 110 applies the protein structure calculation model 117 such that each position in at least a portion of the backbone of the first protein structure 170 includes at least the side chain atoms of the amino acid residue selected for inclusion at the corresponding position in the output protein sequence 150b, thereby further generating the first protein structure 170. For example, in the example shown in FIG. 6B, the protein structure calculation model 117 may generate the first protein structure 170 by including at least the side chain atoms of each amino acid residue of arginine (R), alanine (A), serine (S), glutamine (Q), aspartic acid (D), valine (V), asparagine (N), threonine (T), alanine (A), valine (V), alanine (A) in the backbone of the first protein structure 170. As described in more detail below, the first protein structure 170 may be further determined by determining the side chain conformation of each amino acid residue selected for inclusion in the output protein sequence 150b by the protein design engine 110 that applies the protein structure calculation model 117.
[0087] At 308, the protein structure calculation model 117 can determine the side-chain conformation of each amino acid residue selected for inclusion in the output protein sequence. In some exemplary embodiments, the protein design engine 110 can further generate a first protein structure 170 having an output protein sequence 150b by applying the protein structure calculation model 117 to at least determine the side-chain conformation of the first protein structure 170. Optionally, the side chains of the first protein structure 170 can include chemical groups attached to the backbone of the first protein structure 170 for each amino acid residue selected for inclusion in the output protein sequence 150b. Optionally, the side chain of an amino acid residue, such as the amino acid residue arginine (R) selected to occupy the first position 600a of the output protein sequence 150b, can include one or more atoms (e.g., side-chain atoms) attached to the backbone atoms of the amino acid residue (e.g., an array of atoms including a nitrogen (N) atom, an alpha-carbon ( TIFF2025524582000017.tif4170) atom, and a carboxyl carbon (C) atom). Optionally, the side chain of an amino acid residue can include one or more atoms (e.g., side-chain atoms) attached to the alpha-carbon ( TIFF2025524582000018.tif4170) atom of the backbone of the amino acid residue.
[0088] In some cases, the protein structure calculation model 117 may determine the side-chain conformation of the first protein structure 170 based on at least the backbone conformation of the first protein structure 170. Further, the protein structure calculation model 117 can be applied to determine the side-chain conformation of the first protein structure 170 with various degrees of freedom. For example, in some cases, the protein structure calculation model 117 may select at least one of a plurality of possible rotamers associated with an amino acid residue, each corresponding to a single possible conformation of the constituent side-chain atoms, to determine the side-chain conformation of each amino acid residue in at least a portion of the output protein sequence 150b. By determining the side-chain conformation of the first protein structure 170 in this way, each amino acid residue in the first protein structure 170 is restricted to one of several individual structural variations, which may be more computationally efficient than determining the side-chain conformation over continuous structural variations. Alternatively, the protein structure calculation model 117 may determine the side-chain conformation of each amino acid residue in at least a portion of the output protein sequence 150b by determining one or more of the translation of the side-chain, the rotation of the side-chain, and / or the torsion angle of one or more rotatable bonds formed by the side-chain atoms.
[0089] FIG. 3B shows a flowchart illustrating an example of a process 350 for hybrid protein design according to some exemplary embodiments. Referring to FIGS. 1-2 and FIGS. 3A-3B, process 350 can be executed by protein design engine 110 to generate an output protein sequence 150b and a corresponding first protein structure 170 based at least on output protein sequence 150a. In some cases, process 350 can perform operation 206 of process 200, where protein structure calculation model 117 selects constituent amino acid residues of the protein sequence and determines a corresponding protein structure by determining at least the backbone and side-chain conformations of the protein structure. Further, in some cases, process 350 can perform operation 302 of process 300, where protein structure calculation model 117 selects a plurality of amino acid residues that meet one or more criteria, such as minimization of an energy function of first protein structure 170 having output protein sequence 150b, for inclusion in output protein sequence 150b.
[0090] In 352, the protein structure calculation model 117 can determine the energy of a first protein molecule having a first plurality of amino acid residues and a first conformation. In some exemplary embodiments, the protein structure calculation model 117 can generate a first protein structure 170 by at least determining a permutation of amino acid residues for inclusion in the output protein sequence 150b and a corresponding conformation (or spatial arrangement) of constituent atoms having a minimum energy. Optionally, the protein structure calculation model 117 can determine a permutation of amino acid residues for inclusion in the output protein sequence 150b by at least selecting, for each position in the output protein sequence 150b, an amino acid residue from a corresponding set of possible amino acid residues 160 for inclusion in the output protein sequence 150b. Further, in addition to determining the identity of the amino acid residues at each position in at least a portion of the output protein sequence 150b, the protein structure calculation model 117 can determine the conformation of the corresponding first protein structure 170, for example, by determining the backbone and side-chain conformations of the first protein structure 170 having the output protein sequence 150b.
[0091] In some cases, the protein structure calculation model 117 can determine the output protein sequence 150b and the first protein structure 170 formed by the output protein sequence 150b by at least generating a plurality of protein molecules having different amino acid residue sequences and / or different conformations in order to identify a specific protein sequence and conformation that meet one or more criteria. That is, the output protein sequence 150b and the corresponding protein structure 170 generated by the protein structure calculation model 117 can meet one or more criteria such as minimization of the energy of the first protein structure 170 having the output protein sequence 150b. Therefore, in some cases, when generating a first protein molecule having a first plurality of amino acid residues and a first conformation, the protein structure calculation model 117 can determine the first energy of the first protein molecule, for example, by applying an energy function. In some cases, the protein structure calculation model 117 can apply an energy function based on one or more of initial quantum mechanics, density functional theory (DFT), semi-empirical methods, molecular mechanics force fields, statistical potentials, neural potentials, machine learning models (e.g., trained on structural data), etc. For example, in some cases, the energy function for determining the first energy of the first protein molecule may be an energy function based on physics that determines the total energy of the first protein molecule, including one or more of electrostatic energy, covalent bond energy, van der Waals energy, etc. It should be understood that different energy functions can be associated with different accuracies and computational complexities. For example, a more accurate energy function such as an initial quantum mechanics-based energy function may impose a greater computational overhead than a less accurate energy function such as a molecular mechanics force field-based energy function. Therefore, in some cases, the protein structure calculation model 117 may apply the first energy function instead of the second energy function to determine the first energy of the first protein molecule based at least on the respective accuracies and / or computational complexities of the first energy function and the second energy function.As will be described in more detail below, the protein structure calculation model 117 can generate additional protein molecules having different sequences and / or different conformations of amino acid residues before identifying a protein molecule having the lowest energy (e.g., total energy).
[0092] At 354, the protein structure calculation model 117 can generate a second protein molecule having a second plurality of amino acid residues and a second conformation by modifying at least one of the first plurality of amino acid residues and the first conformation. In some exemplary embodiments, the protein structure calculation model 117 can generate one or more additional protein molecules having a sequence of amino acid residues and / or a different conformation that is different from the first protein molecule generated in operation 352. For example, in some cases, the protein structure calculation model 117 may generate a second protein molecule having at least one of a sequence of amino acid residues and a different conformation that is different from the first protein molecule generated in operation 352. In some cases, the second protein molecule can be generated by modifying the sequence of amino acid residues forming the first protein molecule, e.g., by inserting, deleting, and / or modifying one or more of the amino acid residues in the first protein molecule. Alternatively and / or additionally, the second protein molecule can be generated by modifying the conformation of the first protein molecule, e.g., by modifying one or more of the backbone conformation and the side chain conformation of the first protein molecule.
[0093] In 356, the protein structure calculation model 117 can determine the second energy of the second protein molecule. In some exemplary embodiments, when generating a second protein molecule having a second plurality of amino acid residues and a second conformation, the protein structure calculation model 117 can determine the second energy of the second protein molecule, for example, by applying an energy function. In some cases, the energy function for determining the second energy of the second protein molecule may also be a physics-based energy function that determines the total energy of the second protein molecule, including one or more of electrostatic energy, covalent bond energy, van der Waals energy, etc. for the second protein molecule. For example, in some cases, the protein structure calculation model 117 can determine the second energy of the second protein molecule by applying an energy function based on one or more of first-principles quantum mechanics, density functional theory (DFT), semi-empirical methods, molecular mechanics force fields, statistical potentials, neural potentials, machine learning models (e.g., trained on structural data), etc.
[0094] In 358, the protein structure calculation model 117 can generate a protein structure having a first plurality of amino acid residues and a first conformation based on at least the first energy being lower than the second energy. In some exemplary embodiments, the protein structure calculation model 117 can generate an output protein sequence 170b of the first protein structure 150 to have a first plurality of amino acid residues forming the first protein molecule when the first energy of the first protein molecule is lower than the second energy of the second protein molecule. Further, when the first energy of the first protein molecule is lower than the second energy of the second protein molecule, the protein structure calculation model 117 can generate the first protein structure 170 to have the first conformation of the first protein molecule. Further, in some cases, the protein structure calculation model 117 can generate a third protein molecule having a third plurality of amino acid residues and a third conformation by modifying the first protein molecule (e.g., the first plurality of molecules and / or the first conformation) or the second protein molecule (e.g., the second plurality of molecules and / or the second conformation). When the first energy of the first protein molecule is lower than the third energy of the third protein molecule, the protein structure calculation model 117 can generate the first protein structure 170 to have the first conformation of the first protein molecule.
[0095] FIG. 4A shows a flowchart illustrating another example of a process 400 for hybrid protein design according to some exemplary embodiments. Referring to FIGS. 1-3A, 3B, and 4A, process 400 can be executed by protein design engine 110 to verify a protein structure generated by an in-silico workflow. For example, in some cases, protein design engine 110 can execute process 400 to verify a protein structure generated by a protein structure calculation model, such as protein structure calculation model 117 shown in FIG. 1. Alternatively and / or additionally, protein design engine 110 can execute process 500 to verify a protein structure generated by at least grafting a portion of a protein structure (e.g., a donor protein structure) onto a corresponding portion of another protein structure (e.g., a recipient or template protein structure).
[0096] At 402, protein design engine 110 can apply a protein structure calculation model to determine at least a second protein structure having the same protein sequence as the first protein structure, based on the protein sequence of the at least first protein structure. In some exemplary embodiments, when protein design engine 110 applies protein design calculation model 117 to determine a first protein structure 170 having output protein sequence 150b, protein design engine 110 can apply a different protein structure calculation model to verify that output protein sequence 150b folds into the first protein structure 170 (e.g., having the backbone and side-chain conformations of the first protein structure 170 determined by protein design calculation model 117). For example, in the example shown in FIG. 4B, protein design engine 110 can apply protein structure calculation model 117, which can be a different protein structure calculation model than protein structure calculation model 450, to determine at least a second protein structure 455 formed by output protein sequence 150b.
[0097] In some cases, the protein design engine 110 can apply the protein structure calculation model 450 or a plurality of different protein structure calculation models to generate a second protein structure 455 or a plurality of protein structures formed by the output protein array 150b. Further, in some cases, the protein structure calculation model 450 may be different from the protein structure calculation model 117 because it includes at least a machine learning model different from the protein structure calculation model 117 and / or implements a structure design protocol different from that of the protein structure calculation model 450. For example, if the protein structure calculation model 117 determines the first protein structure 170 by at least determining the amino acid residues included in the output protein array 150b and the conformation (or spatial arrangement) of the corresponding amino acid residues (e.g., the constituent atoms in each amino acid residue), the protein structure calculation model 450 may determine the second protein structure 465 based at least on the output protein array 150b.
[0098] At 404, the protein design engine 110 can determine / calculate a similarity metric that quantifies the difference between the first protein structure and the second protein structure. In some exemplary embodiments, the protein design engine 110 may include a structure analyzer 460 that determines a similarity metric 465 that quantifies the difference between a first protein structure 170 generated by the protein structure calculation model 117 and a second protein structure 455 generated by the protein structure calculation model 450. It should be understood that when the protein design engine 110 generates not only the second protein structure 465 but also a plurality of protein structures associated with the output protein array 150b, the structure analyzer 460 may determine the similarity metric 465 for each protein structure. In some cases, the similarity metric 465 may include a root mean square deviation (RMSD) calculated based on the best superimposed atomic coordinates of the first protein structure 170 and the second protein structure 455. However, it should be understood that the similarity metric 465 may include other values that quantify the structural differences between the first protein structure 170 and the second protein structure 455. For example, in some cases, the structure analyzer 460 may perform principal component analysis (PCA), in which case the similarity metric 465 may include the correlation between the respective symmetry interaction matrices of the first protein structure 170 and the second protein structure 455. The symmetry interaction matrix of a protein structure can be constructed to include relationship parameters between secondary elements such as distances, orientations, and / or other relevant structure invariants.
[0099] In 406, the protein design engine 110 may identify the protein sequence of the first protein structure as a synthesis candidate based on a similarity metric that meets at least one other (predetermined) threshold. In some exemplary embodiments where the similarity metric 465 that quantifies the structural difference between the first protein structure 170 and the second protein structure 150b meets one or more thresholds, the protein design engine 110 may verify that the output protein sequence 150b folds into the first protein structure 170 determined by the protein structure calculation model 117. That is, when a different protein structure calculation model, such as the protein structure calculation model 450, determines the same (or sufficiently similar) protein structure for the output protein sequence 150b, the protein design engine 110 can verify that the first protein structure 170 has the actual conformation of the output protein sequence 117b. When the protein design engine 110 generates not only the output protein structure 465 but also a plurality of protein structures related to the output protein sequence 150b, the protein design engine 110 may verify that the first protein structure 170 has the actual conformation of the output protein sequence 150b when the similarity metric 465 of a threshold amount of protein structures generated by a different protein structure calculation model, such as the second protein structure 455 generated by the protein structure calculation model 450, meets one or more thresholds.
[0100] If a first protein structure 170 (e.g., a particular conformation of the first protein structure 170) is associated with one or more desired properties, the protein design engine 110 may identify the output protein sequence 150b for further analysis, such as synthesis, in vitro measurement, in vivo property evaluation, etc., if the protein design engine 110 can verify that the output protein sequence 150b folds into the conformation of the first protein structure 170. For example, the first protein structure 170 may exhibit one or more desired properties such as expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for the target molecule, lack of non-specificity, stability, non-immunogenicity, humanicity, lack of self-association (or non-aggregation), etc. Thus, if the protein design engine 110 can verify that the output protein sequence 150b folds into the first protein structure 170, at least the protein design engine 110 may determine that the physical protein structure synthesized from the output protein sequence 150b is likely to exhibit the same desired properties as the in silico generated first protein structure 170, and thus the output protein sequence 150b is identified for further analysis.
[0101] FIG. 5 shows a flowchart depicting another example of a process 500 for hybrid protein design, according to some exemplary embodiments. Referring to FIGS. 1 - 5, the process 500 may be performed by the protein design engine 110 to generate a protein structure by grafting at least a protein structure onto another protein structure. Optionally, the process 200 described in FIG. 2 and / or the process 300 described in FIG. 3A may be performed to generate the donor protein structure and / or the recipient (or template) protein structure used in the process 500. As described in more detail below, the donor protein structure may be a portion of a protein molecule (e.g., an antigen-binding fragment (Fab), paratope, complementarity-determining region (CDR), variable region (Fv), etc.) that is transplanted onto a corresponding portion of the recipient (or template) protein structure to form a complete protein molecule.
[0102] In 502, the protein design engine 110 may generate a first protein structure that includes at least a first portion of a protein molecule. As described above, in some exemplary embodiments, the protein design engine 110 may apply a protein sequence calculation model 113 and a protein structure calculation model 117 to generate a first protein structure 170 having an output protein sequence 150b. In some examples, the first protein structure 170 generated by the protein sequence calculation model 113 and the protein structure calculation model 117 may be a donor structure that includes a portion of the entire protein molecule, such as a paratope, variable region (Fv), antigen-binding fragment (Fab), or complementarity-determining region (CDR) of an antibody. For example, in some cases, the protein design engine 110 may apply the protein sequence calculation model 113 and the protein structure calculation model 117 to generate a specific portion of the protein molecule (not the entire protein molecule). In some cases, the aforementioned portions of the protein molecule may be one or more domains, each of which is an independent folded portion of the protein molecule. The first protein structure 170 that functions as a donor structure may include all portions of the protein molecule generated by the protein sequence calculation model 113 and the protein structure calculation model 117, or, in some cases, further sub-portions of that portion of the protein molecule. For example, when the protein design engine 110 applies the protein sequence calculation model 113 and the protein structure calculation model 117 to generate an antigen-binding fragment (Fab) of an antibody molecule, the first protein structure 170 may be one or more of the variable region (Fv) or complementarity-determining region (CDR) of the antigen-binding fragment (Fab). Alternatively, the first protein structure 170 may be a portion of the entire protein molecule (e.g., an antibody, etc.) generated by the protein sequence calculation model 113 and the protein structure calculation model 117.When the protein sequence calculation model 113 and the protein structure calculation model 117 generate an entire protein molecule such as an entire antibody, the first protein structure 170 can be a protein molecule such as a paratope, variable region (Fv), antigen-binding fragment (Fab), or complementarity-determining region (CDR) of the antibody, which is identified by the protein design engine 110 to function as a donor structure. In the example shown in FIGS. 7A-7B, a donor structure including at least a part of the first protein structure 170 generated by the protein sequence calculation model 113 and the protein structure calculation model 117 can be transplanted into a second protein structure 700 (e.g., a recipient (or template) protein structure) to form a third protein structure 750 shown in the embodiment shown in FIG. 7A.
[0103] In 504, the protein design engine 110 may replace a second part of the second protein structure with the first protein structure to generate a third protein structure. In some exemplary embodiments, the protein design engine 110 may replace a second part (e.g., a recipient (or template) protein structure) of the second protein structure 700 with the first protein structure 170 to generate the third protein structure 750. For example, in some cases, the first protein structure 170 may be a part of a protein molecule such as a paratope, variable region (Fv), antigen-binding fragment (Fab), or complementarity-determining region (CDR) of an antibody. Thus, the protein design engine 110 may transplant the first protein structure 170 into the second protein structure 700 by replacing at least the corresponding part of the second protein structure 700 with the first protein structure 170. For example, if the first protein structure 170 is the first paratope, the first variable region (Fv), the first antigen-binding fragment (Fab), or the first complementarity-determining region (CDR) of the first antibody, the protein design engine 110 may replace the second paratope, the second variable region (Fv), the second antigen-binding fragment (Fab), or the second complementarity-determining region (CDR) of the second antibody with the first protein structure 170.
[0104] In some cases, the protein design engine 110 may transplant the first protein structure 170 into the second protein structure 700 to generate a third protein structure 750. In some cases, the third protein structure 750 may be a complete protein molecule (e.g., an antibody, etc.) that combines portions of a plurality of protein structures including the first protein structure 170 and the second protein structure 700. For example, in some cases, the third protein structure 750 may include a variable region (Fv) of the first protein structure 170 and a constant region (Fc) of the second protein structure 700. Alternatively and / or additionally, the third protein structure 750 may include an antigen-binding fragment (Fab) of the first protein structure 170 and a crystallizable fragment (Fc) of the second protein structure 700. In some cases, the third protein structure 750 may include one or more complementarity-determining regions (CDRs) (or hypervariable regions) of the first protein structure 170 and one or more framework regions of the second protein structure 700.
[0105] In 506, the protein design engine 110 can determine the protein sequence of a third protein structure generated to include at least a first portion of a first protein structure and a third portion of a second protein structure. In some exemplary embodiments, the protein design engine 110 may determine the protein sequence 760 of the third protein structure 750, which is generated to include portions of the first protein structure 170 and the second protein structure 700 and is associated with different underlying protein sequences. (An exemplary sequence 760 is shown in FIG. 7B.) Thus, the protein sequence 760 of the third protein structure 750 may also include at least a portion of the output protein sequence 150a of the first protein structure 170 and the protein sequence of the second protein structure 700. For example, in some cases, the protein sequence 760, the third protein structure 750 may include a first sequence of amino acid residues that form the variable region (Fv) of the first protein structure 170 and a second sequence of amino acid residues that form the constant region (Fc) of the second protein structure 700. Alternatively and / or additionally, the protein sequence 760 of the third protein structure 750 may include a first sequence of amino acid residues that form the antigen-binding fragment (Fab) of the first protein structure 170 and a second sequence of amino acid residues that form the crystallizable fragment (Fc) of the second protein structure 700. In some cases, the protein sequence 760 of the third protein structure 750 may include a first sequence of amino acid residues that form one or more complementarity-determining regions (CDRs) (or hypervariable regions) of the first protein structure 170 and a second sequence of amino acid residues that form one or more framework regions of the second protein structure 700. As described in more detail below, in some cases, the protein design engine 110 applies a protein structure computational model to verify that the protein sequence 760 of the third protein structure 750 assumes the conformation of the third protein structure 750 formed by grafting the first protein structure 170 (e.g., donor structure) onto the second protein structure 700 (e.g., acceptor or template structure).
[0106] At 508, the protein design engine 110 can apply a protein structure calculation model to determine at least a fourth protein structure having the same protein sequence as the third protein structure based on the protein sequence of the at least third protein structure. As described above, in some exemplary embodiments, the protein design engine 110 can verify that the protein sequence 760 of the third protein structure 750, which includes the sequences of amino acid residues from the first protein structure 170 and the second protein structure 700, adopts the conformation of the third protein structure 750 by at least verifying that the first protein structure 170 (e.g., donor structure) grafts to the second protein structure 750 (e.g., acceptor or template structure) to form the third protein structure 700. In some cases, the protein design engine 110 can verify the third protein structure 750 by applying one or more protein structure calculation models 450 to determine one or more additional protein structures 455 for evaluation of the third protein structure 750 based at least on the protein sequence 760 of the third protein structure 750. For example, in some cases, the protein design engine 110 can apply the protein structure calculation model 117 used to generate the first protein structure 170 to generate one or more additional protein structures 455 for verifying the third protein structure 750. Alternatively and / or additionally, the protein design engine 110 can apply the protein structure calculation model used to generate the second protein structure 700 and / or another protein structure calculation model to generate one or more additional protein structures 455 for verifying the third protein structure 750. As will be described in more detail below, verification of the third protein structure 750 can include evaluating the third protein structure 750 against one or more additional protein structures 455 generated based on the protein sequence 760 of the third protein structure 750.
[0107] At 510, the protein design engine 110 can determine a similarity metric that quantifies the difference between a third protein structure and a fourth protein structure. In some illustrative embodiments, the structure analyzer 460 of the protein design engine 110 can determine a similarity metric 465 that quantifies the difference between a third protein structure 750 and one or more additional protein structures 455 generated based on the protein sequence 460 of the third protein structure 750. In some cases, the similarity metric 465 can include a root mean square deviation (RMSD) that quantifies the structural differences between the third protein structure 750 and each of the additional protein structures 455 generated based on the protein sequence 760 of the third protein structure 750. Further, in some cases, the similarity metric 465 can indicate the likelihood that the protein sequence 760 of the third protein structure 750 assumes the conformation of the third protein structure 750 generated by grafting the first protein structure 170 onto the second protein structure 700.
[0108] At 512, the protein design engine 110 can identify the protein sequence of a third protein structure as a candidate for synthesis based on a similarity metric that meets at least one or more thresholds. In some exemplary embodiments, if a similarity metric 465 that quantifies the structural difference between the third protein structure 750 and one or more additional protein structures 455 meets one or more thresholds, the protein design engine 110 can verify that the protein sequence 760 folds into the third protein structure 700 generated by grafting the first protein structure 170 onto the second protein structure 750. In some cases, the protein design engine 110 can verify that the third protein structure 750 represents the actual conformation of the protein sequence 700 if one or more protein structure calculation models 450 determine a protein structure of the same (or sufficiently similar) protein sequence 760 as the first protein structure 170 grafted onto the second protein structure 760. If the protein design engine 110 generates a plurality of protein structures 455 having the protein sequence 760 of the third protein structure 750, the protein design engine 110 can verify that the third protein structure 750 has the actual conformation of the protein sequence 760 if the similarity metric 465 of a threshold amount of additional protein structures 455 meets one or more thresholds.
[0109] In some exemplary embodiments, upon verifying the third protein structure 750, the protein design engine 110 can select the corresponding protein sequence 760 for further analysis, including, for example, synthesis, in vitro measurement, in vivo characterization, etc. For example, in some cases, the third protein structure 760, or one or more of its components such as the first protein structure 170 and / or the second protein structure 700, may exhibit one or more desired properties such as expression, binding affinity for a target molecule (e.g., an antigen such as a viral antigen or a tumor antigen), specificity for the target molecule, lack of nonspecificity, stability, non-immunogenicity, humanicity, lack of self-association (or non-aggregation), etc. Thus, if the protein design engine 110 can verify that the protein sequence 760 folds into the third protein structure 750, the protein design engine 110 can determine that it is highly likely that the physical protein structure synthesized from the protein sequence 760 by the protein design engine 110 will exhibit the same desired properties as the third protein structure 760, and thus can identify the protein sequence 760 of the third protein structure 750 for further analysis. For example, if the first protein structure 170 is a paratope, antigen-binding fragment (Fab), variable region (Fv), or complementarity-determining region (CDR) that exhibits binding affinity for a target molecule, and the third protein structure 750 is generated to include the first protein structure 170 grafted onto the second protein structure 700 as a corresponding constant region, crystallizable fragment (Fc), or framework region of the third protein structure 750, the protein design engine 110 may determine that it is highly likely that the protein structure synthesized from the protein sequence 760 of the third protein structure 750 will exhibit the same binding affinity for the target molecule.
[0110] FIG. 8A is a schematic diagram showing an example of a process for hybrid protein design in which alternative protein sequences are generated for an antigen-binding fragment of the antibody trastuzumab, according to some exemplary embodiments. Referring to FIG. 8A, a protein design engine 110 can determine one or more interface residues based at least on a co-crystal structure 800 of a first protein structure 825 corresponding to the antibody trastuzumab bound to a second protein structure 850 corresponding to the protein human epidermal growth factor receptor 2 (HER2). In this regard, an interface residue can refer to an amino acid residue in one protein molecule, such as the first protein structure 825 of the antibody molecule trastuzumab, that interacts and contacts another protein molecule, such as the second protein structure 850 of the human epidermal growth factor receptor 2 (HER2) protein molecule. As shown in FIG. 8A, the interface residues of the antibody trastuzumab can include the constituent amino acid residues of the antigen-binding fragment (Fab) of the antibody trastuzumab. In some cases, the interface residues of the antibody trastuzumab can include the amino acid residues forming the variable domains of the light chain ( TIFF2025524582000019.tif5170) and heavy chain ( TIFF2025524582000020.tif5170) of the antibody trastuzumab.
[0111] Referring back to FIG. 8A, in some exemplary embodiments, one or more alternative protein sequences for the antibody trastuzumab (or a particular portion thereof) can be generated by applying at least the protein sequence calculation model 113. In the example shown in FIG. 8A, the protein sequence calculation model 113 can be an autoencoder (e.g., a variational autoencoder, etc.) that generates each alternative protein sequence of the antibody trastuzumab by sampling the data distribution occupied by the encoding of known protein sequences. Thus, for each sampling iteration of the data distribution, the protein sequence calculation model 113 can encode at least a portion of the original protein sequence of the antibody trastuzumab before decoding an intermediate sequence having at least one of a disruption (e.g., insertion, deletion, and / or modification of amino acid residues) and a length change to the protein sequence of trastuzumab.
[0112] In some cases, the protein design engine 110 can apply the protein sequence calculation model 113 to generate alternative sequences for a portion of the antibody trastuzumab (e.g., paratope, antigen-binding fragment (Fab), variable region (Fv), complementarity-determining region (CDR), etc.). Alternatively, the protein design engine 110 can apply the protein sequence calculation model 113 to generate alternative sequences for the entire trastuzumab molecule. For example, in some cases, the protein design engine 110 can apply the protein sequence calculation model 113 to generate a plurality of proposed protein sequences based on the original protein sequence (or a portion thereof) of the antibody trastuzumab, and then, based on the proposed protein sequences, determine the possible amino acid residues set for each position of the alternative trastuzumab protein sequence. In some cases, the possible amino acid residues set at a position within the alternative trastuzumab protein sequence can include one or more amino acid residues that appear at that position with sufficient frequency across the proposed protein sequences generated by the protein sequence calculation model 113.
[0113] In some exemplary embodiments, alternative protein sequences of the antibody trastuzumab (or portions thereof) can be generated by applying a protein structure calculation model 117, which determines alternative trastuzumab sequences based on the possible amino acid residues set for each position and the corresponding conformations of protein molecules having alternative protein sequences. If the conformation assumed by the alternative protein sequence is sufficiently similar to the conformation of the original trastuzumab molecule (e.g., the first protein structure 825 within the co-crystal structure 800), the protein design engine 110 can identify the alternative protein sequence for further analysis, including, for example, synthesis, in vitro measurement, in vivo characterization, etc. For example, the protein design engine 110 can apply a protein structure calculation model, which may or may not be the protein structure calculation model 117 applied to generate alternative sequences and structures of trastuzumab, to determine the conformation of a protein molecule having an alternative protein sequence. If the conformation of the protein molecule having the alternative protein sequence shows sufficient similarity to the original structure of trastuzumab, e.g., when a similarity metric (e.g., root mean square deviation (RMSD), etc.) that quantifies the difference between the two molecules meets one or more thresholds, the protein design engine 110 can identify an alternative sequence of trastuzumab for further analysis (e.g., synthesis, in vitro measurement, in vivo characterization, etc.).
[0114] Figure 8B shows a violin graph 860 showing the predicted binding affinities of alternative trastuzumab arrays generated by different hybrid protein design workflows, and a violin graph 865 showing the measured affinities, according to some exemplary embodiments. Referring to Figure 8B, the protein structures generated by each hybrid protein design workflow are at least quantified by the difference between the first energy of the bound trastuzumab human epidermal growth factor receptor 2 (HER2) protein complex and the second energy of the separated protein molecules, (for human epidermal growth factor receptor 2 (HER2)) The predicted binding affinity of each alternative trastuzumab protein sequence can be evaluated by calculating the interfacial energy (dG separation) that indicates. The binding affinity of each alternative trastuzumab protein sequence can also be empirically determined, for example, by measuring the dissociation constant ( TIFF2025524582000021.tif5170). Violin graphs 860 and 865 show the distributions of interfacial energy (dG separation) and measured binding affinity (pkd) across a first population of unbound alternative trastuzumab protein sequences (labeled negative) and a second population of bound alternative trastuzumab protein sequences, respectively. These two populations of alternative trastuzumab protein sequences were generated by a hybrid protein design workflow that included a protein sequence calculation model trained based on a training data set that included a first plurality of labeled unbound protein sequences and a second plurality of labeled bound protein sequences. As shown by graphs 850 and 860, most of the bound alternative trastuzumab protein sequences generated by the hybrid protein design workflow exhibit high binding affinity, as indicated by higher dissociation constants and lower interfacial energies.
[0115] FIG. 8C shows graph 870 depicting the stability of each alternative trastuzumab protein structure as indicated by the total energy of the molecule. Graph 870 of FIG. 8C shows that alternative trastuzumab protein sequences generated by the various hybrid protein design workflows described herein fold into a three-dimensional structure (e.g., at a lower total energy) that is more stable than protein sequence design techniques that do not rely on a conventional structure.
[0116] In view of the above-described embodiments of the subject matter, the present application discloses the following list of examples that are further examples of the present application, by combining one feature of a single example or two or more features of said examples, optionally in combination with one or more features of one or more further examples included in the disclosure of the present application.
[0117] Item 1: A computer-implemented method comprising identifying a protein sequence calculation model and a protein structure calculation model, applying the protein sequence calculation model to generate a plurality of proposed protein sequences based at least on an input protein sequence, identifying, for each position in at least a portion of an output protein sequence, a set of possible amino acid residues based at least on the plurality of proposed protein sequences, and using the protein structure calculation model to generate a first protein structure having the output protein sequence by selecting a possible amino acid residue from the set of possible amino acid residues for each position in at least a portion of the output protein sequence.
[0118] Item 2: The method of item 1, further comprising aligning the plurality of proposed protein sequences to generate a plurality of aligned protein sequences, and identifying, for each position in at least a portion of the output protein sequence, a set of possible amino acid residues based at least on the plurality of aligned protein sequences.
[0119] Item 3: The method according to item 1 or 2, wherein a plurality of proposed protein sequences are aligned by applying one or more of dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, and hidden Markov models.
[0120] Item 4: The method according to any one of items 1 to 4, which includes identifying a first amino acid residue for inclusion in a set of possible amino acid residues for each position in at least a portion of the output protein sequence, but does not include identifying a second amino acid residue.
[0121] Item 5: The method according to item 4, which further includes determining a first frequency at which a first amino acid residue appears at a position across a plurality of proposed protein sequences generated by a protein sequence calculation model, determining a second frequency at which a second amino acid residue appears at a position across the plurality of proposed protein sequences generated by the protein sequence calculation model, and identifying the first amino acid residue for inclusion in the set of possible amino acid residues for the position based on at least the first frequency and the second frequency, but not identifying the second amino acid residue.
[0122] Item 6: The method according to item 5, wherein the first amino acid residue is identified for inclusion in the set of amino acid residues based on a first frequency that meets at least one or more thresholds, and the second amino acid residue is identified for exclusion from the set of possible amino acid residues based on a second frequency that does not meet at least one or more thresholds.
[0123] Item 7: Identifying a set of possible amino acid residues for positions in an output protein sequence, and determining one or more thresholds based on at least one of a maximum value, a minimum value, a median value, an average value, and a mode value of the frequency of occurrence of each of a plurality of amino acid residues at positions across a plurality of proposed protein sequences generated by a protein sequence calculation model. The method according to item 6 further includes this.
[0124] Item 8: The method according to any one of items 1 to 7, wherein the set of possible amino acid residues includes some but not all of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, selenocysteine, and a part of pyrrolidine.
[0125] Item 9: The protein structure calculation model generates a first protein structure by at least determining the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence based on the energy of the first protein structure having at least the output protein sequence. The method according to any one of items 1 to 8.
[0126] Item 10: The method according to item 9, wherein the protein structure calculation model determines the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence by at least modifying at least one of the identity and conformation of the amino acid residue to minimize the energy of the first protein structure.
[0127] Item 11: The method according to item 10, wherein the protein structure calculation model modifies at least one of the identity and conformation of the amino acid residue occupying the position by at least (i) changing the conformation of the amino acid residue occupying one position, or (ii) selecting a different possible amino acid residue for the position from the set of amino acid residues associated with the position.
[0128] Item 12: The method according to any one of Items 9 to 11, wherein the protein structure calculation model determines at least the energy of a first protein structure having a first possible amino acid residue from a set of possible amino acid residues, determines the energy of a second protein structure having a second possible amino acid residue from the set of possible amino acid residues, and generates a first protein structure that includes the first possible amino acid residue instead of the second possible amino acid residue based on at least the first energy being lower than the second energy, thereby determining the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence.
[0129] Item 13: The method according to Item 12, wherein the protein structure calculation model determines at least the energy of a first protein structure having a first conformation of a first possible amino acid residue, determines the energy of a first protein structure having a second conformation of the first possible amino acid residue, and generates a first protein structure that includes the first conformation of the first possible amino acid residue instead of the second conformation of the first possible amino acid residue based on at least the third energy being lower than the fourth energy, thereby further determining the identity and conformation of the amino acid residue occupying each position in at least a portion of the output protein sequence.
[0130] Item 14: The method according to any one of Items 1 to 14, wherein the protein structure calculation model generates a first protein structure by at least determining a first backbone conformation of a first backbone of the first protein structure having an output protein sequence.
[0131] Item 15: The method according to Item 14, wherein the first backbone of the first protein structure is a continuous chain of atoms formed by linking a plurality of backbone atoms from each amino acid residue of the output protein sequence.
[0132] Item 16: The plurality of backbone atoms are nitrogen (N) atoms, TIFF2025524582000022.tif3170 - carbon( TIFF2025524582000023.tif4170) atoms and carboxyl carbon (C) atoms, the method according to item 15.
[0133] Item 17: Determining the first backbone conformation of the first backbone of the first protein structure includes determining that the first backbone of the first protein structure has the same conformation as at least a part of the second backbone of the second protein structure, the method according to any one of items 14 to 16.
[0134] Item 18: The method according to item 17, wherein the second protein structure is associated with a protein sequence determined to have one or more desired properties.
[0135] Item 19: The method according to item 18, wherein the protein sequence having one or more desired properties is an input protein sequence or another protein sequence.
[0136] Item 20: The second backbone of the second protein structure represents the second backbone conformation of the unbound second protein structure or the third backbone conformation of the second protein structure bound to a target molecule, the method according to any one of items 17 to 19.
[0137] Item 21: Determining the first backbone conformation of the first backbone of the first protein structure includes determining the translation of the plurality of backbone atoms contained in each amino acid residue of at least a part of the output protein sequence, the method according to any one of items 14 to 20.
[0138] Item 22: Determining the first backbone conformation of the first backbone of the first protein structure includes determining the rotation of the plurality of backbone atoms contained in each amino acid residue of at least a part of the output protein sequence, the method according to any one of items 14 to 21.
[0139] Item 23: The method according to any one of Items 14 to 22, wherein determining the first backbone conformation of the first protein structure includes determining the torsional angles of one or more rotatable bonds formed by a plurality of backbone atoms included in each amino acid residue of the output protein sequence.
[0140] Item 24: The method according to any one of Items 14 to 23, wherein the protein structure calculation model further generates a first protein structure including a plurality of side chain atoms of an amino acid residue selected to be included at a corresponding position of the output protein sequence at each position of at least a portion of the first backbone having the first backbone conformation.
[0141] Item 25: The method according to Item 24, further including applying a different protein structure calculation model to determine at least a second protein structure having the output protein sequence based at least on the output protein sequence, and determining a similarity metric that quantifies the difference between the second protein structure and the first protein structure generated to have the first backbone conformation.
[0142] Item 26: The method according to Item 25, further including identifying the output protein sequence as a candidate for synthesis based at least on a similarity metric that satisfies one or more thresholds.
[0143] Item 27: The method according to any one of Items 14 to 26, wherein the protein structure calculation model further generates the first protein structure by at least determining the side chain conformation of an amino acid residue selected to be included at a corresponding position of the output protein sequence for each position of at least a portion of the first backbone having the first backbone conformation.
[0144] Item 28: The method according to Item 27, wherein the side chain conformation of the amino acid residue is determined based at least on the first backbone conformation of the first backbone of the first protein structure.
[0145] Item 29: The method according to any one of Items 27 to 28, wherein determining the side-chain conformation of an amino acid residue selected for inclusion in the output protein sequence includes selecting a rotamer from a plurality of possible rotamers.
[0146] Item 30: The method according to Item 29, wherein each rotamer of the plurality of possible rotamers includes a different combination of torsional angles of one or more rotatable bonds formed by a plurality of side-chain atoms of the amino acid residue.
[0147] Item 31: The method according to any one of Items 27 to 30, wherein determining the side-chain conformation of an amino acid residue selected for inclusion in the output protein sequence includes determining the torsional angle of one or more rotatable bonds formed by a plurality of side-chain atoms of the amino acid residue.
[0148] Item 32: The method according to any one of Items 27 to 31, wherein determining the side-chain conformation of an amino acid residue selected for inclusion in the output protein sequence includes determining the translation of a plurality of side-chain atoms of the amino acid residue.
[0149] Item 33: The method according to any one of Items 27 to 32, wherein determining the side-chain conformation of an amino acid residue selected for inclusion in the output protein sequence includes determining the rotation of a plurality of side-chain atoms of the amino acid residue.
[0150] Item 34: Applying a characterization model to determine the characteristics of each protein sequence included in a plurality of proposed protein sequences, identifying at least one protein sequence among the plurality of proposed protein sequences for exclusion based at least on the characteristics of each protein sequence, and excluding at least one protein sequence before identifying a set of possible amino acid residues for each position in the output protein sequence based at least on the remaining plurality of proposed protein sequences. The method according to any one of Items 1 to 33, further comprising.
[0151] Item 35: The method according to any one of Items 1 to 34, further comprising generating a third protein structure by identifying a first portion of a second protein structure and at least replacing the first portion of the second protein structure with at least a second portion of a first protein structure.
[0152] Item 36: Determining the third protein sequence of a third protein structure generated to include a second portion of a first protein structure and a third portion of a second protein structure, applying a protein structure calculation model and / or different protein structure calculation models to determine at least a fourth protein structure having the third protein sequence based on at least the third protein sequence, determining a similarity metric that quantifies the difference between the third protein structure and the fourth protein structure, and identifying the third protein sequence as a synthesis candidate based on the similarity metric that satisfies at least one or more thresholds. The method according to Item 35.
[0153] Item 37: The method according to Item 35 or 36, wherein the second protein structure is selected based on a second protein structure that exhibits at least one or more desired characteristics.
[0154] Item 38: The method according to any one of Items 35 to 37, wherein the first portion of the second protein structure includes the first antigen-binding site of a first antibody having the second protein structure, and the second portion of the first protein structure includes the second antigen-binding site of a second antibody having the first protein structure.
[0155] Item 39: The method according to any one of Claims 35 to 38, wherein the first portion of the second protein structure includes the first paratope of a first antibody having the second protein structure, and the second portion of the first protein structure includes the second paratope of a second antibody having the first protein structure.
[0156] Item 40: The method according to any one of items 35 to 39, wherein the first part of the second protein structure includes the first complementarity-determining region (CDR) of the first antibody having the second protein structure, and the second part of the first protein structure includes the second complementarity-determining region (CDR) of the second antibody having the first protein structure.
[0157] Item 41: The method according to any one of items 1 to 40, wherein the protein structure calculation model generates the first protein structure by at least determining a plurality of amino acid residues for inclusion in the output protein sequence and the corresponding conformation of the plurality of amino acid residues based on one or more criteria.
[0158] Item 42: The method according to item 41, wherein the one or more criteria include minimizing the energy function of the first protein structure having the plurality of amino acid residues selected for inclusion in the output protein sequence.
[0159] Item 43: The method according to item 41 or 42, wherein the protein structure calculation model determines the first energy of the first protein molecule having at least the first plurality of amino acid residues and the first conformation, generates a second protein molecule having the second plurality of amino acid residues and the second conformation by modifying at least one of the first plurality of amino acid residues and the first conformation, determines the second energy of the second protein molecule, and generates the first protein structure having the first plurality of amino acid residues and the first conformation based on the fact that at least the first energy is lower than the second energy.
[0160] Item 44: The method according to any one of items 1 to 43, wherein the protein sequence calculation model includes one or more machine learning models trained to generate a plurality of proposed protein sequences based on the input protein sequence.
[0161] Item 45: The method according to any one of items 1 to 44, wherein a protein sequence calculation model generates a plurality of proposed protein sequences by at least sampling a data distribution populated by a plurality of encoded protein sequences, and each sampling of the data distribution generates an encoding of a protein sequence having at least one of a disruption and a length change to an input protein sequence.
[0162] Item 46: A system comprising at least one data processor and at least one memory storing instructions, which, when executed by the at least one data processor, result in operations including the method according to any one of items 1 to 45.
[0163] Item 47: A non-transitory computer-readable medium storing instructions which, when executed by at least one data processor, result in operations including the method according to any one of embodiments 1 to 45.
[0164] FIG. 9 is a block diagram showing an example of a computing system 900 according to some exemplary embodiments. Referring to FIGS. 1 to 9, the computing system 900 can be used to implement a protein design engine 110, a molecular analysis engine 120, a client device 130, and / or any component thereof.
[0165] As shown in FIG. 9, computing system 900 can include a processor 910, a memory 920, a storage device 930, and an input / output device 940. The processor 910, the memory 920, the storage device 930, and the input / output device 940 can be interconnected via a system bus 950. The processor 910 is capable of processing instructions for execution within the computing system 900. Such executed instructions can implement, for example, a protein design engine 110, an analysis engine 120, a client device 130, and the like. In some exemplary embodiments, the processor 910 can be a single-threaded processor. Alternatively, the processor 910 can be a multi-threaded processor. The processor 910 is capable of processing instructions stored in the memory 920 and / or the storage device 930 to display graphical information for a user interface provided via the input / output device 940.
[0166] The memory 920 is a computer-readable medium, such as volatile or non-volatile, that stores information within the computing system 900. The memory 920 can store, for example, a data structure representing a configuration object database. The storage device 930 is capable of providing persistent storage for the computing system 900. The storage device 930 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 940 provides input / output operations for the computing system 900. In some exemplary embodiments, the input / output device 940 includes a keyboard and / or a pointing device. In various embodiments, the input / output device 940 includes a display device for displaying a graphical user interface.
[0167] According to some exemplary embodiments, the input / output device 940 can perform input / output operations for network devices. For example, the input / output device 940 can include an Ethernet port or other networking ports to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0168] In some exemplary embodiments, the computing system 900 can be used to execute various interactive computer software applications that can be used for the compilation, analysis, and / or storage of various forms of data. Alternatively, the computing system 900 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generation, management, editing of spreadsheet documents, word processing documents, and / or other objects), computing functions, communication functions, etc. The applications can include various add-in functions or can be stand-alone computing products and / or functions. When activated within an application, the functionality can be used to generate a user interface provided via the input / output device 940. The user interface can be generated by the computing system 900 and presented to the user (e.g., on a computer screen monitor, etc.).
[0169] One or more aspects or features of the subject matter described in this specification can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can be included in an implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device. The programmable system or computing system can include clients and servers. Clients and servers are generally located at remote locations and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other.
[0170] These computer programs, sometimes referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), and includes a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium can store such machine instructions non-transitorily, for example, in non-transitory solid-state memory, magnetic hard drive, or any equivalent storage medium. A machine-readable medium can alternatively or additionally store such machine instructions transiently, for example, in a processor cache or other random access memory associated with one or more physical processor cores.
[0171] To provide interaction with a user, one or more aspects or features of the subject matter described in this specification can be implemented on a computer having, for example, a display device such as a cathode ray tube (CRT), liquid crystal display (LCD), or light emitting diode (LED) monitor for displaying information to the user, a keyboard, and a pointing device such as a mouse or trackball by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback such as, for example, visual feedback, auditory feedback, tactile feedback, and the input from the user can be received in any form including acoustic input, voice input, tactile input. Other possible input devices include touch screens, or other touch sensor-based devices such as single or multi-point resistive or capacitive track pads, speech recognition hardware and software, optical scanners, optical pointers, digital image capture devices, and associated interpretation software, and the like.
[0172] In the above specification and claims, phrases such as "at least one of ~" or "one or more of ~" may appear before a list of consecutive elements or features. The term "and / or" may also be used in the enumeration of two or more elements or features. Unless there is an implicit or explicit contradiction in the context in which it is used, such phrases are intended to mean any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B", "one or more of A and B", "A and / or B" are each intended to mean "only A, only B, or A and B together". A similar interpretation is intended for lists containing three or more items. For example, the phrases "at least one of A, B, C", "one or more of A, B, C", "A, B, and / or C" are each intended to mean "only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together". The use of the term "based on" in the above and the claims means "based at least in part on" and means that features or elements not recited are also permitted.
[0173] The subject matter described in this specification can be embodied in a system, apparatus, method, and / or article, depending on the desired configuration. The implementations described in the above description do not necessarily represent all implementations according to the subject matter described in this specification. Rather, these are merely some examples that are consistent with aspects related to the described subject matter. Although several variations have been detailed above, other changes and additions are possible. In particular, additional features and / or variations can be provided in addition to those described in this specification. For example, the above-described embodiments can be directed to various combinations and sub-combinations of the disclosed features, and / or combinations and sub-combinations of some of the additional features disclosed above. Further, the logical flows depicted in the accompanying figures and / or described in this specification do not necessarily require the particular order, or sequential order, shown to obtain a desired result. Other embodiments can also be within the scope of the following claims.
Claims
1. Identifying a protein sequence calculation model and a protein structure calculation model; Applying the protein sequence calculation model to generate a plurality of proposed protein sequences based at least on an input protein sequence; Identifying, for each position in at least a portion of the output protein sequence, a set of possible amino acid residues based at least on the plurality of proposed protein sequences; and Applying the protein structure calculation model to select, for each position in at least the portion of the output protein sequence, a possible amino acid residue from the set of possible amino acid residues for inclusion in the output protein sequence, thereby generating, using the protein structure calculation model, a first protein structure having the output protein sequence A computer-implemented method comprising.
2. Aligning the plurality of proposed protein sequences to generate a plurality of aligned protein sequences; and Identifying, for each position in at least the portion of the output protein sequence, the set of possible amino acid residues based at least on the plurality of aligned protein sequences The method according to claim 1, further comprising.
3. The method according to claim 1 or 2, wherein the plurality of proposed protein sequences are aligned by applying one or more of dynamic programming, progressive alignment, hierarchical alignment, iterative alignment, motif discovery, deep learning models, and hidden Markov models.
4. The method according to any one of claims 1 to 3, wherein identifying the set of possible amino acid residues for each position in at least the portion of the output protein sequence comprises identifying a first amino acid residue for inclusion in the set of possible amino acid residues, but does not include identifying a second amino acid residue.
5. Identifying the set of possible amino acid residues for each position in at least the portion of the output protein sequence comprises Determining a first frequency at which a first amino acid residue appears at the position across the plurality of proposed protein sequences generated by the protein sequence calculation model Determining a second frequency of occurrence of a second amino acid residue at the position across the plurality of proposed protein sequences generated by the protein sequence calculation model, and identifying the first amino acid residue for inclusion in the set of possible amino acid residues for the position based on at least the first frequency and the second frequency, but not identifying the second amino acid residue The method according to claim 4, further comprising.
6. The first amino acid residue is identified for inclusion in the set of amino acid residues based on the first frequency satisfying at least one or more thresholds, and the second amino acid residue is based on the second frequency not satisfying at least the one or more thresholds. The method according to claim 5, identified for exclusion from the set of possible amino acid residues.
7. Identifying the set of possible amino acid residues for the position in the output protein sequence is Determining the one or more thresholds based on at least one of a maximum value, a minimum value, a median value, an average value, and a mode value of the frequency of occurrence of each of a plurality of amino acid residues at the position across the plurality of proposed protein sequences generated by the protein sequence calculation model The method according to claim 6, further comprising.
8. The set of possible amino acid residues includes some but not all of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, selenocysteine, and pyrrolidine. The method according to any one of claims 1 to 7.
9. The protein structure calculation model generates the first protein structure by at least determining the identity and conformation of amino acid residues occupying each position in at least the portion of the output protein sequence based on the energy of the first protein structure having at least the output protein sequence. The method according to any one of claims 1 to 8.
10. The method according to claim 9, wherein the protein structure calculation model determines the identity and conformation of the amino acid residues occupying each position in at least a portion of the output protein sequence by at least modifying at least one of the identity and the conformation of the amino acid residues so as to minimize the energy of the first protein structure.
11. The method according to claim 10, wherein the protein structure calculation model modifies at least one of the identity and the conformation of the amino acid residue occupying the position by at least (i) changing the conformation of the amino acid residue occupying one position, or (ii) selecting a different possible amino acid residue for the position from the set of amino acid residues associated with the position.
12. The protein structure calculation model at least determines a first energy of the first protein structure having a first possible amino acid residue from the set of possible amino acid residues, determines a second energy of the first protein structure having a second possible amino acid residue from the set of possible amino acid residues, and generates the first protein structure to include the first possible amino acid residue instead of the second possible amino acid residue based on at least that the first energy is lower than the second energy to determine the identity and the conformation of the amino acid residues occupying each position in at least a portion of the output protein sequence, according to any one of claims 9 to 11.
13. The protein structure calculation model at least determines a third energy of the first protein structure having a first conformation of the first possible amino acid residue, determines a fourth energy of the first protein structure having a second conformation of the first possible amino acid residue, and generates the first protein structure to include the first conformation of the first possible amino acid residue instead of the second conformation of the first possible amino acid residue based on at least that the third energy is lower than the fourth energy The method of claim 12, further determining the identity and conformation of the amino acid residues occupying each position in at least said portion of said output protein sequence. **Claim 14** Applying a property analysis model to determine the properties of each protein sequence included in said plurality of proposed protein sequences; Identifying at least one protein sequence among said plurality of proposed protein sequences for exclusion, based at least on said properties of each protein sequence; and Excluding said at least one protein sequence from said plurality of proposed protein sequences, prior to identifying the set of possible amino acid residues for each position in said output protein sequence, based at least on the remaining plurality of proposed protein sequences. The method according to any one of claims 1 to 13, further comprising. **Claim 15** Identifying a first portion of a second protein structure; and Generating a third protein structure by at least substituting said first portion of said second protein structure with at least a second portion of said first protein structure. The method according to any one of claims 1 to 14, further comprising. **Claim 16** Determining a third protein sequence of said third protein structure generated to include said second portion of said first protein structure and a third portion of said second protein structure; Applying said protein structure calculation model and / or a different protein structure calculation model to determine at least a fourth protein structure having said third protein sequence, based at least on said third protein sequence; Determining a similarity metric that quantifies the difference between said third protein structure and said fourth protein structure; and Identifying said third protein sequence as a candidate for synthesis, based at least on said similarity metric satisfying one or more thresholds. The method of claim 15, further comprising. **Claim 17** The method according to claim 15 or 16, wherein said second protein structure is selected based on said second protein structure exhibiting at least one or more desired properties. **Claim 18** The method according to any one of claims 15 to 17, wherein the first portion of the second protein structure comprises a first antigen-binding site of a first antibody having the second protein structure, and the second portion of the first protein structure comprises a second antigen-binding site of a second antibody having the first protein structure.
19. The method according to any one of claims 15 to 18, wherein the first portion of the second protein structure comprises a first paratope of a first antibody having the second protein structure, and the second portion of the first protein structure comprises a second paratope of a second antibody having the first protein structure.
20. The method according to any one of claims 15 to 19, wherein the first portion of the second protein structure comprises a first complementarity-determining region (CDR) of a first antibody having the second protein structure, and the second portion of the first protein structure comprises a second complementarity-determining region (CDR) of a second antibody having the first protein structure.
21. At least one data processor, and At least one memory storing instructions that, when executed by the at least one data processor, cause an operation comprising the method according to any one of claims 1 to 20 A system comprising.
22. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause an operation comprising the method according to any one of claims 1 to 20.