Fragment-based quantum mechanical calculation of protein properties

By separating the polypeptide sequence into data units and combining a mixed method of classical molecular mechanics and quantum mechanics, the efficiency and accuracy problems of calculating the properties of complex proteins in the prior art are solved, and rapid and accurate prediction of protein properties are achieved.

CN120283283APending Publication Date: 2025-07-08MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102139.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing calculation methods When predicting the molecular properties of proteins, classical molecular mechanics methods cannot capture quantum effects caused by electron motion, while density functional theory is accurate but computationally intensive, making it difficult to quickly and accurately simulate complex protein systems on conventional processors.

Method used

Using a fragment-based quantum mechanics computing system, the force and energy of each data unit is calculated using density functional theory and machine learning models by separating the polypeptide sequence into multiple data units by combining classical molecular mechanics and quantum mechanics, and the total energy and energy of the polypeptide sequence is calculated by mixing strategies.

Benefits of technology

Fast and accurate prediction of protein molecular properties is achieved, reducing calculation time while maintaining high accuracy, suitable for drug design and protein research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283283A_ABST
    Figure CN120283283A_ABST
Patent Text Reader

Abstract

A computing system for fragment-based quantum mechanical calculation of protein properties is provided. The processor implements a protein fragmentation module that separates the computer-readable polypeptide sequence into a plurality of data units. For each sub-sequence of three adjacent amino acids in the polypeptide sequence, a first amino acid, a second amino acid, and a third amino acid are identified, each amino acid having a respective backbone comprising an amino group, a carbon, and a carboxyl group, and a side chain attached to an alpha carbon. The protein fragmentation module generates data units representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon, and stores the generated data units in a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In the field of computational chemistry, computer-based techniques have been developed to predict molecular properties through computer simulations. These molecular properties can have a wide range of effects on the appearance and function of molecules or materials, and thus have attracted attention in various fields. For example, in the field of drug design, changes in molecular properties can affect the efficacy of drugs. In the field of drug discovery, molecular properties can affect the likelihood that a naturally discovered material can be used for therapeutic purposes. In the field of quantum chemistry, quantum mechanical calculations of the contributions of electrons to the physical and chemical properties of molecules and materials are a fundamental area of research. As described below, there are still opportunities for improvement in computational methods for predicting molecular properties, and the improvements will have applications outside the field of computational chemistry. Summary of the Invention

[0002] To address the issues discussed herein, a computerized system and method for fragment-based quantum mechanical calculations of protein properties are provided. In one aspect, the computerized system includes a processor that executes instructions using a portion of an associative memory to implement a protein fragmentation module. The protein fragmentation module separates a computer-readable polypeptide sequence representing multiple amino acids into multiple data units. For each subsequence of three adjacent amino acids in the polypeptide sequence, the protein fragmentation module is configured to identify a first amino acid, identify a second amino acid, identify a third amino acid, generate a data unit, and store the generated data unit. The first amino acid has a first backbone that includes a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon. The second amino acid has a second backbone that includes a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon. The third amino acid has a third backbone that includes a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon. The data unit includes data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon.

[0003] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that will be further described in the Detailed Implementation below. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of this disclosure. Brief Description of the Drawings

[0004] Figure 1 A schematic diagram of a computational system for fragment-based quantum mechanical calculations of protein properties according to an embodiment of the present disclosure is shown.

[0005] Figure 2 shows a computing system using Figure 1 to generate data units from a truncated polypeptide sequence.

[0006] Figure 3 shows a computing system using Figure 1 to generate data units from a tetrapeptide.

[0007] Figure 4 shows a subset of atoms coexisting in a data unit generated from a Figure 3 tetrapeptide.

[0008] Figure 5 shows the atoms coexisting in a data unit generated from a Figure 3 tetrapeptide, and an indication of redundant units.

[0009] Figure 6 shows the pairs of coexisting atoms to be included in the force calculation of a polypeptide for each atom in a Figure 3 tetrapeptide.

[0010] Figure 7 shows Figure 3 the atoms in a tetrapeptide for which additional interactions need to be calculated by molecular mechanics.

[0011] Figure 8 shows Figure 3 the atoms in a tetrapeptide for which additional interactions need to be calculated by a combination of quantum mechanics and molecular mechanics.

[0012] Figure 9 shows a flowchart of a method for fragment-based quantum mechanical calculation of protein properties according to an example implementation of the present disclosure.

[0013] Figure 10 shows an example computing environment in which embodiments of the present disclosure may be implemented. DETAILED IMPLEMENTATION

[0014] Computer-based technologies have been developed to predict molecular properties through computer simulations. For example, molecular dynamics (MD) simulation is a widely used computational tool for simulating atomic motion. As atoms change their physical positions over the simulation time period, the MD model calculates the potential energy and the net force on each atom of the molecular system, thereby describing the dynamics and thermodynamic properties of the molecular system. MD is widely used in the fields of physics, chemistry, biology, and pharmacy because understanding the mechanisms of protein molecules enables progress in drug design, protein design, enzyme engineering, etc.

[0015] MD simulations can be performed using classical molecular mechanics (MM) or quantum mechanics (QM). Classical MM is based on Newtonian mechanics and has been widely used for proteins. Classical MM simulations using empirical force fields can achieve fast simulation results for large systems, but have the drawback of being unable to capture quantum effects caused by electron motion. Additionally, the parameters of the force fields calculated within such simulations are usually not transferable.

[0016] In contrast, QM provides highly accurate calculations for atoms and molecules and can thus be used to study biological processes with electronic transitions. Density functional theory (DFT) is the most widely used method in quantum simulations. DFT is a powerful quantum physical calculation technique that can accurately predict various molecular properties, such as the energy and forces of molecules, the shape of molecules, etc., in many cases. Although MD simulations driven by DFT can accurately calculate energy and forces, DFT is time-consuming and computationally intensive. A single model for a simple molecule on a conventional processor typically takes several hours, while simulations for proteins with 1000 or more atoms usually take several months. Therefore, for complex molecular systems, calculating accurate DFT solutions is not practical on current hardware. These factors pose obstacles to accurately and effectively predicting the molecular properties of proteins.

[0017] To address these problems, a computational system for fragment-based quantum mechanics calculations of protein properties is provided. Although directly running QM on biomolecules is computationally infeasible, applying a hybrid strategy using QM and classical MM enables more effectively and accurately determining the forces on each atom in a polypeptide sequence (i.e., a protein). The embodiments discussed herein describe a novel method using polypeptide fragments (i.e., data units) to calculate the molecular properties of proteins using a combination of QM and classical MM.

[0018] First refer to Figure 1, computing system 10 includes at least one computing device. Computing system 10 is shown as including a first computing device 14 and a second computing device 16. The first computing device 14 includes a processor 18 and a memory 22, and the second computing device 16 includes a processor 20 and a memory 24. The illustrated implementation is exemplary in nature, and other configurations are possible. In the following description, the first computing device will be described as server 14, and the second computing device will be described as client computing device 16, and the corresponding functions performed at each device will be described. It should be understood that in other configurations, computing system 12 may include a single computing device that performs the significant functions of both server 14 and client computing device 16, and the first computing device may be a computing device other than a server. In other alternative configurations, the functions described as being performed at server 14 may alternatively be performed at client computing device 16, and vice versa.

[0019] Continue Figure 1 , processor 18 is configured to implement a protein fragmentation module 26 hosted at server 14. Protein fragmentation module 26 separates a computer-readable polypeptide sequence 28 representing a plurality of amino acids into a plurality of data units. Polypeptide sequence 28 may be stored at a protein sequence database 30 (such as UniProt, Swiss-Prot, Protein Research Foundation (PRF), etc.), and is sent to protein fragmentation module 26 upon receipt of user input via user interface 32 at client computing device 16. It should be understood that a polypeptide chain of one hundred or more amino acids linked together by covalent peptide bonds is generally considered a protein. In the embodiments described herein, computing system 10 is configured to determine forces and energies for proteins as well as for polypeptides consisting of fewer than 100 amino acids.

[0020] The amino acids in polypeptide sequence 28 may be represented in single-letter code (e.g., ALGY for alanine, leucine, glycine, and tyrosine) or three-letter code (e.g., AlaLeuGlyTyr for alanine, leucine, glycine, and tyrosine). As discussed in detail below with reference to Figure 2 and 3 , the data unit generator 34 included in protein fragmentation module 26 is configured to separate polypeptide sequence 28 into a plurality of data units 36, each data unit 36 representing the atomic structure of an amino acid subsequence in polypeptide sequence 28. The plurality of data units 36 may be stored in data unit database 38 on server 14. It should be understood that data unit database 38 may include a plurality of containers 40, each container storing data units 36 derived from a corresponding polypeptide sequence 28.

[0021] The processor 18 is also configured to implement a data unit property calculation module 42 that calculates the force F of each atom in each of the plurality of data units 36 and calculates the energy E of each data unit 36. As Figure 1 shown and discussed in detail below with reference to Figures 4 to 6 , the data unit property calculation 40 may include a quantum simulation program 44 that applies DFT to calculate the force of each atom in each data unit 36 and calculates the energy of each data unit 36. Alternatively, in some embodiments, the force of each atom in each data unit 36 and the energy of each data unit 36 may be determined via a machine learning (ML) model 46, such as a vector-scalar interaction graph neural network (ViSNet).

[0022] Continuing to refer to Figure 1 , the processor 18 is also configured to implement a polypeptide property calculation module 48 that calculates the force of the polypeptide sequence based on the calculated force of each atom in each of the plurality of data units 26 and calculates the energy of the polypeptide sequence 28 based on the calculated energy for each of the plurality of data units. As described in detail below with reference to Figure 8 and 9 , the energy of the polypeptide sequence 28 is calculated by summing the calculated energies of each of the plurality of data units 26 and subtracting the energy of the repeating regions shared by adjacent data units of the polypeptide sequence. Similarly, the force of the polypeptide sequence 28 is calculated by summing the calculated forces of each of the plurality of data units 26 and subtracting the force of the repeating regions shared by adjacent data units of the polypeptide sequence. The polypeptide property calculation module 48 includes a classical MM simulation program 50 to calculate the interaction between the backbone atoms of the data unit 26 and the side chain atoms of non-adjacent data units. The polypeptide property calculation module 48 also includes a hybrid QM-MM simulation program 52. This program enables the interaction between the side chain atoms of data units separated by a distance less than or equal to a distance threshold to be calculated via equilibrium QM (counterpoise QM) applying DFT, while the interaction between the side chain atoms of data units separated by a distance greater than the distance threshold is calculated via MM.

[0023] Once determined, the calculated energies and forces 54 for each polypeptide sequence 28 can be stored in the Protein Energy and Force Database 56 on the server 14. In response to user input at the client computing device 16, the energies and forces for the polypeptide sequence can be displayed as a graph 60 on the display 58 in the user interface 32. In any implementation described herein, it should be understood that the server 14 communicates with the client computing device 16 via the network 62, which allows a user of the client computing device to access the data and programs stored on the server 14, including the data stored in the Protein Sequence Database 30, the Data Unit Database 38, and the Protein Energy and Force Database 56.

[0024] Protein fragmentation

[0025] Each amino acid includes a backbone having an amino group (NH2), an alpha carbon (Cα), and a carboxyl group (COOH), and a side chain R attached to the alpha carbon. There are generally considered to be 21 amino acid side chains, and each amino acid side chain determines the identity of the amino acid. When forming a polypeptide chain, the amino group of the downstream amino acid forms a peptide bond with the carboxyl group of the upstream amino acid in a biochemical reaction that releases a water molecule. Reading the amino acid sequence from left to right, the first amino group forms the N-terminus at the start of the sequence, and the last carboxyl group forms the C-terminus at the end of the sequence.

[0026] For each subsequence of three adjacent amino acids in the polypeptide sequence 28, the data unit generator 34 is configured to identify a first amino acid, a second amino acid, and a third amino acid. The first amino acid has a first backbone that includes a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon. The second amino acid has a second backbone that includes a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon. The third amino acid has a third backbone that includes a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon. The data unit generator 34 is configured to generate a data unit 36 that includes data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon.

[0027] Figure 2Shows examples of two generated data units 26A, 26B. As shown, the truncated polypeptide sequence 28 is separated into a first data unit 36A indicated by a dashed line and a second data unit 36 indicated by a dotted line. The first alpha carbon and the first carboxyl group of each data unit 36 include the N-terminal acetyl group (ACE) of the data unit, and the third amino group and the third alpha carbon include the C-terminal N-methylamino group (NME) of the data unit. Each data unit 36 also includes data representing a first peptide bond P1 and a second peptide bond P2. The first peptide bond is formed between the N-terminal ACE and the second amino group, and the second peptide bond is formed between the second carboxyl group and the C-terminal NME. For both peptide bonds, each data unit 36 can be considered a novel dipeptide (DIP).

[0028] The overlapping region between the first data unit 36A and the second data unit 36B in the truncated polypeptide sequence 28 is indicated by parentheses. The overlapping region includes the N-terminal ACE from the second data unit 36B and the C-terminal NME from the first data unit 36A. As discussed in detail below, when calculating the forces and energies for the polypeptide sequence, the forces and energies of each overlapping region (i.e., redundant ACE-NME unit 64) must be subtracted from the equation.

[0029] For each data unit 36, data representing one or more additional hydrogen atoms is added to the first alpha carbon in each data unit according to the first bond length and the first direction of the previous bond between the first alpha carbon and the first side chain. Additionally, data representing one or more additional hydrogen atoms is added to the third alpha carbon in each data unit according to the third bond length and the third direction of the previous bond between the third alpha carbon and the third side chain. The Limited Memory Broyden-Fletcher-Goldfarb-Shanno quasi-Newton (LBFGS) algorithm is applied to optimize the positions of the one or more additional hydrogen atoms.

[0030] Figure 3 Illustrates a tetrapeptide separated into four data units 36A, 36B, 36C, and 36D. The four data units are shown separately in boxes. Each data unit includes a backbone having an amino group (N X H), an alpha carbon (CA X ), and a hydroxyl group (C X O X ), and a side chain R X attached to the alpha carbon. The alpha carbon and the hydroxyl group from the upstream amino acid in the polypeptide sequence include an ACE cap at the N-terminus, and the amino group and the alpha carbon from the downstream amino acid include an NME cap at the C-terminus. For example, in data unit 26A, the backbone includes N 1 H, CA 1 and C 1 O 1 . The side chain R 1is attached to the α-carbon CA 1 The upstream α-carbon CA 0 and the hydroxyl group C 0 O 0 form an N-terminal ACE cap, and the downstream amino group N 2 H and the α-carbon CA 2 form a C-terminal NME cap. Separating the tetrapeptide into four data units generates three redundant ACE-NME units, as shown by the dashed, dotted, and double-dotted lines in Figure 3 as shown.

[0031] For each amino acid represented in polypeptide sequence 28, a data unit 36 is generated and stored in the data unit database 38. Then, for each data unit 36, the energy and forces required to determine the force field of the polypeptide sequence are calculated. The generalized force field consists of two parts: the energy and force calculations for each data unit 36, and the two-body interaction calculations between neighboring data units 36. From these two aspects, the total energy and forces for the entire polypeptide (i.e., protein) can be accurately determined.

[0032] Data Unit Property Calculation

[0033] The following paragraphs provide additional descriptions of the implementation for calculating the molecular properties of each data unit 36. As described above, there are two different ways to calculate the molecular properties of data units, including quantum mechanics (QM) based on ORCA and deep learning (DL) models.

[0034] As described above, the data unit property calculation module 42 includes a quantum simulation program 44. In the quantum mechanics (QM) mode, the quantum simulation program 44 applies density functional theory (DFT) to calculate the forces on each atom in each generated data unit 36 and calculates the energy of each data unit. An example implementation of such a QM program is ORCA, a general quantum chemistry program package that includes modern electronic structure methods such as DFT. Using ORCA, a DFT such as the M06-2X density functional is applied to a basis set such as the 6-31G(d) basis set to calculate the forces on each atom and the energy of the data unit 36. The M06-2X functional is a high-nonlocal functional with twice the nonlocal exchange term (2X), and it is only parameterized for non-metals. In the 6-31G basis set, each inner shell (1s orbital) STO is a linear combination of 6 primitive functions, and each valence shell STO is divided into inner and outer (double zeta) using 3 and 1 primitive Gaussian functions, respectively.

[0035] As described above, when combining data units 36 to determine the total energy of a polypeptide sequence, the redundant ACE-NME units 64 between adjacent amino acids must be taken into account. Thus, the total energy of the entire protein can be approximately calculated by summing the energies of data units 36 and subtracting the energies of all redundant ACE-NME units 64, as shown in Equation 1, where n is the number of amino acids or data units.

[0036] The forces on the atoms in the same data units 36 and ACE-NME 64 are calculated according to Equation 2.

[0037] In Equation 2, i represents the atom for which the force is calculated, m represents all the data units to which atom i belongs, n represents all the ACE-NME units 64 to which atom i belongs, and j represents any other atom coexisting with atom i in the same data unit 36 or ACE-NME unit 64.

[0038] Figure 4 and Figure 5 shows example atoms coexisting in the tetrapeptide introduced in Figure 3 and discussed above. The tetrapeptide is illustrated again in Figure 4 for reference. The atoms included in data unit 36A, i.e., dipeptide 1 (DIP1), are in italics; the atoms included in data unit 36B, i.e., dipeptide 2 (DIP2), are underlined; the atoms included in data unit 36C, i.e., dipeptide 3 (DIP3), are in bold; the atoms included in data unit 36D, i.e., dipeptide 4 (DIP4), are in italics and underlined.

[0039] See Figure 4 the first row, showing the neighboring atoms coexisting with the alpha carbon CA 1 CA 1 coexists with the atoms in the first data unit 36A and the second data unit 36B, as well as the atoms in the first ACE-NME. The forces for each pair of atoms of CA 1 and the coexisting atoms are calculated and summed. For the atoms in data units and ACE-NMEs that do not coexist with this atom, they are not represented in atomic form. For example, CA 1 does not coexist with the atoms in data unit 36C (denoted as DIP3), data unit 36D (DIP4), ACE-NME 2 or ACE-NME 3 in the atoms.

[0040] Atoms included in the repeating regions are indicated by boxes, and atoms from all repeating regions except one are removed from the calculation, as shown by the crossed-out boxes. For example, including C 1 O 1 、N 2 H 2 and CA 2 The regions are included in the first data unit 36A, the second data unit 36B, and the first ACE–NME unit. Thus, for the calculation of the force on the atom pair with CA 1 the atoms in the second data unit 36B and the first ACE–NME unit are excluded from the calculation. Figure 4 shows a subset of the atoms included in the tetrapeptide, while Figure 5 shows the coexisting atoms for each atom in the data units 36A, 36B, 36C, 36D, and ACE–NME in the tetrapeptide 1 、ACE–NME 2 and ACE–NME 3 Each atom in the and ACE–NME

[0041] Figure 6 shows the coexisting atom pairs for each atom in the tetrapeptide that will be included in the calculation of the force of the polypeptide. After each row of coexisting atoms is a summary of the interactions. For example, in Figure 6 the first row of, CA 0 H3 coexists with C 0 O 0 、N 1 H、CA 1 H、C 1 O 1 、N 2 H、CA 2 H、R 1 which can be summarized as six heavy atoms from CA i to CA i+2 and one side chain R i+1 .

[0042] Alternatively, the data unit property calculation module 42 can run the ML model 46 to calculate the force on each atom in each generated data unit 36 and calculate the energy of each data unit. For example, the ML model 46 can be implemented as a vector–scalar interaction graph neural network (ViSNet). Using this method, the coordinates and atom types for each data unit 36 or ACE-NME unit 64 are inputs to the ViSNet model, and the model produces the force on each atom and the energy for the data unit 36.

[0043] Polypeptide property calculation

[0044] Using the above quantum simulation program 44 or ML model 46 makes it possible to calculate all the energies and forces in the same data units 36 and ACE-NME unit 64. However, the additional interactions between different units have not been calculated. The following paragraphs provide an additional description of the implementation of the molecular properties for calculating the additional interactions between the data unit 36 and the ACE-NME unit 64 to determine the forces and energies for the polypeptide sequence. As described above, there are two different methods for calculating the molecular properties of the polypeptide sequence 28, including the classical MM program 50 and the QM-MM simulation program 52.

[0045] Figure 7 Shows the atoms in the tetrapeptide for the additional interactions that need to be calculated (see Figure 3 and 4 ). Figure 7 The top diagram A of 1 shows the interactions that have not been calculated between the atoms in the first data unit 36A and the atoms in the third data unit 36C. Specifically, the boxed area in the top diagram A) indicates the interactions between the CA 1 O 1 N 2 H atoms in the first data unit 36A and the R 3 C 3 O 3 N 4 H, C 4 H 3 atoms in the third data unit 36C that need to be calculated.

[0046] Figure 7 The bottom diagram B) of 0 illustrates the interactions that have not been calculated between the atoms in the first data unit 36A and the atoms in the second data unit 36B and the third data unit 36C. Specifically, the boxed area in the bottom diagram B) indicates the interactions between the C 3 H 0 O 0 N 1 H and R 1 atoms in the first data unit 36A and the R 2 C 2 O 2 N 3 H, CA 3 R 3 C 3 O 3 N 4 H, C 4 H 3The interactions between atoms that need to be calculated. As described above and will be described in detail below, there are two methods available for estimating these interactions.

[0047] In the first method, additional interactions are calculated by MM. The MM method includes two types of interactions: Coulomb interaction and van der Waals interaction. Then, the corresponding parameters from the molecular dynamics (MD) force field (FF) simulation program and the distances between atoms are used to calculate the energy and force, as shown in Equations 3 and 4 below.

[0048] The MD FF simulation program can be, for example, Assisted Model Building with Energy Refinement (AMBER), which uses the FF19SB force field that uses amino acid-specific backbone parameters and improves the modeling of amino acid-dependent properties (such as helix propensity).

[0049] The energy E and force F with the subscript "units" represent the values obtained from the combination of data unit 36 and ACE-NME unit 64 (Equations 1, Equation 2), and A indicates the set of atoms in each data unit 36. As shown in Equation 4, the summations in the second and third terms iterate over all atoms j whose indices are after the current atom i and do not coexist with atom i in any data unit 36.

[0050] Using the second method, additional interactions are calculated by a combination of QM and MM. In the QM-MM method, the interactions between neighboring side chains (i.e., side chains within a distance threshold λ of each other) are calculated via counterpoise correction for quantum simulation at the DFT level, while the interactions between the remaining atom pairs are calculated using the MM method.

[0051] Figure 8 Shows the atoms in the tetrapeptide where additional interactions need to be calculated (see Figure 3 and 4 ). Figure 8 The top diagram A) of 1 shows the interactions between the atoms of the first side chain R 2 , the second side chain R 3 , and the third side chain R 1 to R 2The distance between is within the threshold distance λ, while R 1 to R 3 and R 2 to R 3 The distance between is greater than the threshold distance λ. Thus, as described below, the interaction between the side chain R 1 and R 2 The interaction between the atoms in will be calculated by quantum simulation and DFT via equilibrium correction, while R 1 and R 3 and between R 2 and R 3 The interaction between will be calculated using the MM method described above.

[0052] Figure 8 The middle graph B) of shows the additional interactions calculated using the MM method. The interactions between the atoms in the side chain are not included in this calculation because those interactions are determined based on the threshold distance λ. Figure 8 The bottom graph C of shows the additional interactions between the atoms calculated by MM. When using the MM method as described in Figure 7 These interactions include the atoms in the R 1 side chain. However, when using the QM-MM method, depending on the threshold distance λ, the interactions between the atoms in the R 1 side chain are calculated using QM or MM and are therefore not included in the additional interactions calculated by default using MM.

[0053] To construct an equilibrium QM system, the coordinates of the two side chains included in the calculation are extracted, and hydrogen atoms are added to the β-carbon of the side chain in the α-carbon direction according to the C-H bond length. If the side chain is glycine, hydrogen atoms are added according to the H-H bond length. If the side chain is proline, two hydrogen atoms are added: one to the β-carbon in the α-carbon direction and one to the delta-carbon in the N-terminal direction. Then, three systems are constructed using the two side chains. The first system has the two side chains and their basis functions, the second system has the first side chain and the basis functions of the two side chains, and the third system has the second side chain and the basis functions of the two side chains. The algorithm is shown in equations 5 and 6 shown below, where the first side chain is A, the second side chain is B, and λ defines the distance between the side chains in the polypeptide sequence.

[0054] In the energy E calculation of equation 5, Indicates the interaction of the first system with side chains A and B with the basis functions of A and B, while Indicates the interaction of the second system with side chain A with the basis functions of A and B. The sum reflects the counterpoise energy. The other subscripts in the calculation of the force F shown in Equation 6 have similar meanings.

[0055] Figure 9 FIG. 5 shows a flowchart of a method 900 for fragment-based quantum mechanical calculations of protein properties according to an example implementation of the present disclosure. The method 900 can be implemented by the hardware and software of the above-described computing system 10, or by other suitable hardware and software.

[0056] It should be understood that steps 902 to 910 of the method 900 are performed for each subsequence of three adjacent amino acids in the polypeptide sequence. In step 902, the method 900 includes identifying a first amino acid. As described above, the first amino acid has a first backbone that includes a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon.

[0057] Continuing from step 902 to step 904, the method 900 includes identifying a second amino acid. As described above, the second amino acid has a second backbone that includes a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon.

[0058] Proceeding from step 904 to step 906, the method 900 includes identifying a third amino acid. As described above, the third amino acid has a third backbone that includes a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon.

[0059] Advancing from step 906 to step 908, the method 900 includes generating a data unit. As described above, the data unit includes data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon. The first alpha carbon and the first carboxyl group include the N-terminal acetyl group (ACE) of the data unit, and the third amino group and the third alpha carbon include the C-terminal N-methylamino group (NME) of the data unit. The data unit also includes data representing a first peptide bond and a second peptide bond, where the first peptide bond is formed between the N-terminal ACE and the second amino group, and the second peptide bond is formed between the second carboxyl group and the C-terminal NME. The data units generated for the polypeptide sequence together represent the atomic structure of the amino acids in the polypeptide sequence.

[0060] The method may further include adding data representing one or more additional hydrogen atoms to the first alpha carbon in each data unit according to a first bond length and a first direction of a previous bond between the first alpha carbon and the first side chain, and adding data representing one or more additional hydrogen atoms to the third alpha carbon in each data unit according to a third bond length and a third direction of a previous bond between the third alpha carbon and the third side chain.

[0061] Continuing from step 908 to step 910, method 900 includes storing the generated data units in a database. Multiple data units for a polypeptide sequence can be stored in a container in the database, and the database can include multiple containers, each storing data units derived from a corresponding polypeptide sequence.

[0062] Advancing from step 910 to step 912, method 900 includes calculating the force on each atom in the data unit. Advancing from step 912 to step 914, method 900 includes calculating the energy of the data unit. In the quantum mechanics mode, density functional theory is applied to calculate the force on each atom in the generated data unit and to calculate the energy of the data unit. In the machine learning mode, the coordinates and atom types for each data unit are input into a machine learning model to calculate the force on each atom in the generated data unit and to calculate the energy of the data unit.

[0063] Continuing from step 914 to step 916, method 900 includes calculating the force of the polypeptide sequence based on the calculated force on each atom in each of the multiple data units. The energy of the polypeptide sequence is calculated by summing the calculated energies of each of the multiple data units and subtracting the energy of the repeating regions shared by adjacent data units of the polypeptide sequence.

[0064] Proceeding from step 916 to step 918, method 900 includes calculating the energy of the polypeptide sequence based on the calculated energy for each of the multiple data units. The force of the polypeptide sequence is calculated by summing the calculated forces of each of the multiple data units and subtracting the force of the repeating regions shared by adjacent data units of the polypeptide sequence.

[0065] As described in detail above, the interactions between the backbone atoms of a data unit and the side chain atoms of non-adjacent data units are calculated via molecular mechanics, the interactions between the side chain atoms of data units separated by a distance less than or equal to a distance threshold are calculated via equilibrium quantum mechanics applying DFT, and the interactions between the side chain atoms of data units separated by a distance greater than the distance threshold are calculated via molecular mechanics.

[0066] Figure 10 A non-limiting embodiment of a computing system 1000 that can perform one or more of the above methods and processes is schematically shown. Computing system 1000 is shown in a simplified form. Computing system 1000 can embody the above and Figure 1The computer system 10 as shown. The computing system 1000 can take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).

[0067] The computing system 1000 includes a logical processor 1002, volatile memory 1004, and a non-volatile storage device 1006. The computing system 1000 can optionally include a display subsystem 1008, an input subsystem 1010, a communication subsystem 1012, and / or Figure 10 other components not shown herein.

[0068] The logical processor 1002 includes one or more physical devices configured to execute instructions. For example, the logical processor can be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions can be implemented to perform tasks, implement data types, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.

[0069] The logical processor can include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, the logical processor can include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 1002 can be single-core or multi-core, and the instructions executed thereon can be configured for sequential, parallel, and / or distributed processing. The components of the logical processor can optionally be distributed across two or more separate devices, which can be located remotely and / or configured to coordinate processing. Aspects of the logical processor can be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In such a case, it should be understood that these virtualized aspects run on different physical logical processors of various different machines.

[0070] The non-volatile storage device 1006 includes one or more physical devices configured to hold instructions executable by the logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 1006 can be transformed, for example, to hold different data.

[0071] The non-volatile storage device 1006 may include removable and / or built-in physical devices. The non-volatile storage device 1006 may include optical memories (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memories (e.g., ROM, EPROM, EEPROM, flash memory, etc.), and / or magnetic memories (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.) or other mass storage device technologies. The non-volatile storage device 1006 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that the non-volatile storage device 1006 is configured to hold instructions even when the non-volatile storage device 1006 is powered off.

[0072] The volatile memory 1004 may include a physical device that includes random access memory. The volatile memory 1004 is typically used by the logic processor 1002 to temporarily store information during the processing of software instructions. It should be understood that when the volatile memory 1004 is powered off, the volatile memory 1004 generally does not continue to store instructions.

[0073] Aspects of the logic processor 1002, the volatile memory 1004, and the non-volatile storage device 1006 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field programmable gate arrays (FPGA), programmable and application specific integrated circuits (PASIC / ASIC), programmable and application specific standard products (PSS / ASSP), system on a chip (SOC), and complex programmable logic devices (CPLD).

[0074] The terms "module", "program", and "engine" may be used to describe aspects of the computing system 1000, which are typically implemented in software by a processor to use a portion of the volatile memory to perform a specific function that involves a transformational process of specifically configuring the processor to perform the function. Thus, a portion of the volatile memory 1004 may be used, via the logic processor 1002 executing instructions held by the non-volatile storage device 1006, to instantiate a module, program, or engine. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may encompass individuals or groups of executable files, data files, libraries, drivers, scripts, database records, etc.

[0075] When included, the display subsystem 1008 can be used to present a visual representation of data held by the non-volatile storage device 1006. The visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data held by the non-volatile storage device and thus transform the state of the non-volatile storage device, the state of the display subsystem 1008 can likewise be transformed to visually represent the changes in the underlying data. The display subsystem 1008 can include one or more display devices utilizing almost any type of technology. Such display devices can be combined with the logic processor 1002, volatile memory 1004, and / or non-volatile storage device 1006 in a shared enclosure, or such display devices can be peripheral display devices.

[0076] When included, the input subsystem 1010 can include one or more user input devices such as a keyboard, mouse, touch screen, or game controller, or interface with such devices. In some embodiments, the input subsystem can include or interface with selected natural user input (NUI) components. Such components can be integrated or peripheral, and the conversion and / or processing of input actions can be handled on-vehicle or off-vehicle. Example NUI components can include a microphone for voice and / or sound recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition; and electric field sensing components for evaluating brain activity; and / or any other suitable sensors. When included, the communication subsystem 1012 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 1012 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network or a wired or wireless local or wide area network (such as HDMI over a Wi-Fi connection). In some embodiments, the communication subsystem can allow the computing system 600 to send messages to and / or receive messages from other devices via a network such as the Internet.

[0077] The following paragraphs provide additional descriptions of aspects of the present disclosure. In one aspect, a computing system for fragment-based quantum mechanical calculations of protein properties is provided. The computing system may include a processor that executes instructions using a portion of an associative memory to implement a protein fragmentation module that separates a computer-readable polypeptide sequence representing a plurality of amino acids into a plurality of data units. The protein fragmentation module may be configured for each subsequence of three adjacent amino acids in the polypeptide sequence: to identify a first amino acid having a first backbone including a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon; to identify a second amino acid having a second backbone including a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon; to identify a third amino acid having a third backbone including a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon; to generate a data unit including data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon; and to store the generated data unit in a database.

[0078] In this aspect, additionally or alternatively, the first alpha carbon and the first carboxyl group may include an N-terminal acetyl group (ACE) of the data unit, the third amino group and the third alpha carbon may include a C-terminal N-methylamino group (NME) of the data unit, and the data unit further includes data representing a first peptide bond and a second peptide bond, the first peptide bond being formed between the N-terminal ACE and the second amino group, and the second peptide bond being formed between the second carboxyl group and the C-terminal NME.

[0079] In this regard, additionally or alternatively, data representing one or more additional hydrogen atoms may be added to the first alpha carbon in each data unit according to a first bond length and a first direction of a previous bond between the first alpha carbon and the first side chain, and data representing one or more additional hydrogen atoms may be added to the third alpha carbon in each data unit according to a third bond length and a third direction of a previous bond between the third alpha carbon and the third side chain.

[0080] In this aspect, additionally or alternatively, the Limited Memory Broyden-Fletcher-Goldfarb-Shanno quasi-Newton (LBFGS) algorithm may be applied to optimize the positions of one or more additional hydrogen atoms.

[0081] In this aspect, additionally or alternatively, the processor may be further configured to execute instructions to implement a data unit property calculation module that calculates the force of each atom in the data unit and calculates the energy of the data unit.

[0082] In this regard, additionally or alternatively, in the quantum mechanics (QM) mode, the data unit property calculation module may apply density functional theory (DFT) to calculate the forces on each atom in the generated data unit and calculate the energy of the data unit.

[0083] In this regard, additionally or alternatively, in the machine learning mode, the data unit property calculation module may input the coordinates and atom types for each data unit into a machine learning model to calculate the forces on each atom in the generated data unit and calculate the energy of the data unit.

[0084] In this regard, additionally or alternatively, the processor may be further configured to execute instructions to implement a polypeptide property calculation module. The polypeptide property calculation module calculates the force of a polypeptide sequence based on the forces on each atom calculated in each of the plurality of data units, and calculates the energy of the polypeptide sequence based on the energies calculated for each of the plurality of data units. The energy of the polypeptide sequence may be calculated by summing the energies calculated for each of the plurality of data units and subtracting the energy of the repetitive regions shared by adjacent data units of the polypeptide sequence. The force of the polypeptide sequence may be calculated by summing the forces calculated for each of the plurality of data units and subtracting the forces of the repetitive regions shared by adjacent data units of the polypeptide sequence.

[0085] In this regard, additionally or alternatively, the interaction between the backbone atoms of a data unit and the side chain atoms of a non-adjacent data unit may be calculated via molecular mechanics.

[0086] In this regard, additionally or alternatively, the interaction between the side chain atoms of data units separated by a distance less than or equal to a distance threshold may be calculated via equilibrium quantum mechanics applying DFT, and the interaction between the side chain atoms of data units separated by a distance greater than the distance threshold may be calculated via molecular mechanics.

[0087] On the other hand, a method for fragment-based quantum mechanical calculation of protein properties is provided. The method may include: for each subsequence of three adjacent amino acids in a polypeptide sequence, identifying a first amino acid having a first backbone including a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon; identifying a second amino acid having a second backbone including a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon; identifying a third amino acid having a third backbone including a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon; generating a data unit including data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon; and storing the generated data unit in a database.

[0088] In this aspect, additionally or alternatively, the first alpha carbon and the first carboxyl group include an N-terminal acetyl group (ACE) of the data unit, the third amino group and the third alpha carbon include a C-terminal N-methylamino group (NME) of the data unit, and the data unit further includes data representing a first peptide bond and a second peptide bond, the first peptide bond being formed between the N-terminal ACE and the second amino group, and the second peptide bond being formed between the second carboxyl group and the C-terminal NME.

[0089] In this aspect, additionally or alternatively, the method may further include adding data representing one or more additional hydrogen atoms to the first alpha carbon in each data unit according to a first bond length and a first direction of a previous bond between the first alpha carbon and the first side chain, and adding data representing one or more additional hydrogen atoms to the third alpha carbon in each data unit according to a third bond length and a third direction of a previous bond between the third alpha carbon and the third side chain.

[0090] In this aspect, additionally or alternatively, the method may further include calculating the force of each atom in the data unit and calculating the energy of the data unit.

[0091] In this regard, additionally or alternatively, the method may further include: in a quantum mechanics mode, applying density functional theory to calculate the force of each atom in the generated data unit and calculating the energy of the data unit.

[0092] In this aspect, additionally or alternatively, the method may further include, in a machine learning mode, a data unit property calculation module inputting the coordinates and atom types of each data unit into a machine learning model to calculate the force of each atom in the generated data unit and calculating the energy of the data unit.

[0093] In this regard, additionally or alternatively, the method may further include calculating the force of the polypeptide sequence based on the force of each atom calculated in each of the plurality of data units, and calculating the energy of the polypeptide sequence based on the energy calculated for each of the plurality of data units. The energy of the polypeptide sequence may be calculated by summing the energies calculated for each of the plurality of data units and subtracting the energy of the repetitive regions shared by adjacent data units of the polypeptide sequence. The force of the polypeptide sequence may be calculated by summing the forces calculated for each of the plurality of data units and subtracting the force of the repetitive regions shared by adjacent data units of the polypeptide sequence.

[0094] In this regard, additionally or alternatively, the method may further include calculating the interaction between the backbone atoms of a data unit and the side chain atoms of a non-adjacent data unit via molecular mechanics.

[0095] In this regard, additionally or alternatively, the method may further include: calculating the interaction between the side chain atoms of data units separated by a distance less than or equal to a distance threshold via equilibrium quantum mechanics applying DFT, and calculating the interaction between the side chain atoms of data units separated by a distance greater than the distance threshold via molecular mechanics.

[0096] On the other hand, a computing system for fragment-based quantum mechanics calculations of protein properties is provided. The computing system may include a processor that uses a portion of an associative memory to execute instructions to implement a protein fragmentation module that separates a computer-readable polypeptide sequence representing a plurality of amino acids into a plurality of data units. The protein fragmentation module may be configured to generate, for each subsequence of three adjacent amino acids in the polypeptide sequence, a data unit that includes data representing a first alpha carbon and a first carboxyl group from a first amino acid, a second amino group, a second alpha carbon, a second carboxyl group, and a second side chain from a second amino acid, and a third amino group and a third alpha carbon from a third amino acid. Using a quantum simulation program, a data unit property calculation module may apply density functional theory (DFT) to calculate the force of each atom in the generated data unit and calculate the energy of the data unit. A polypeptide property calculation module may calculate the force of the polypeptide sequence based on the force of each atom calculated in each of the plurality of data units, and may calculate the energy of the polypeptide sequence based on the energy calculated for each of the plurality of data units.

[0097] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered restrictive as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various acts shown and / or described may be performed in the order shown and / or described, in other orders, in parallel, or omitted. Similarly, the order of the processes described above may be changed.

[0098] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein, and any and all equivalents thereof.

Claims

1. A computing system for fragment-based quantum mechanics calculations of protein properties, comprising: A processor that uses a portion of an associative memory to execute instructions to implement a protein fragmentation module that separates a computer-readable polypeptide sequence representing a plurality of amino acids into a plurality of data units, where The protein fragmentation module is configured to, for each subsequence of three adjacent amino acids in the polypeptide sequence: Identify a first amino acid having a first backbone that includes a first amino group, a first α-carbon, and a first carboxyl group, and a first side chain attached to the first α-carbon, Identify a second amino acid having a second backbone that includes a second amino group, a second α-carbon, and a second carboxyl group, and a second side chain attached to the second α-carbon, Identify a third amino acid having a third backbone that includes a third amino group, a third α-carbon, and a third carboxyl group, and a third side chain attached to the third α-carbon, Generate a data unit that includes data representing the first α-carbon, the first carboxyl group, the second amino group, the second α-carbon, the second carboxyl group, the second side chain, the third amino group, and the third α-carbon, and Store the generated data unit in a database.

2. The computing system according to claim 1, where The first α-carbon and the first carboxyl group include an N-terminal acetyl group (ACE) of the data unit, The third amino group and the third α-carbon include a C-terminal N-methylamino group (NME) of the data unit, and The data unit further includes data representing a first peptide bond and a second peptide bond, where the first peptide bond is formed between the N-terminal ACE and the second amino group, and the second peptide bond is formed between the second carboxyl group and the C-terminal NME.

3. The computing system according to claim 1, where Data representing one or more additional hydrogen atoms is added to the first α-carbon in each data unit according to a first bond length and a first direction of a previous bond between the first α-carbon and the first side chain, and Data representing one or more additional hydrogen atoms is added to the third α-carbon in each data unit according to a third bond length and a third direction of a previous bond between the third α-carbon and the third side chain.

4. The computing system according to claim 1, where the processor is further configured to execute instructions to implement: A data unit property calculation module that calculates the force of each atom in the data unit and calculates the energy of the data unit.

5. The computing system according to claim 4, where In a quantum mechanics (QM) mode, the data unit property calculation module applies density functional theory (DFT) to calculate the force of each atom in the generated data unit and calculates the energy of the data unit.

6. The computing system according to claim 4, where In the machine learning mode, the data unit property calculation module inputs the coordinates and atom types of each data unit into a machine learning model to calculate the forces of each atom in the generated data unit and calculate the energy of the data unit.

7. The computing system according to claim 4, wherein the processor is further configured to execute instructions to implement: a polypeptide property calculation module that calculates the force of the polypeptide sequence based on the forces of each atom calculated in each data unit of the plurality of data units and calculates the energy of the polypeptide sequence based on the energy calculated for each data unit of the plurality of data units, wherein the energy of the polypeptide sequence is calculated by summing the energies calculated for each data unit of the plurality of data units and subtracting the energy of the repeating regions shared by adjacent data units of the polypeptide sequence, and the force of the polypeptide sequence is calculated by summing the forces calculated for each data unit of the plurality of data units and subtracting the forces of the repeating regions shared by adjacent data units of the polypeptide sequence.

8. The computing system according to claim 7, wherein the interaction between the backbone atoms of a data unit and the side chain atoms of a non-adjacent data unit is calculated via molecular mechanics.

9. The computing system according to claim 7, wherein the interaction between the side chain atoms of data units separated by a distance less than or equal to a distance threshold is calculated via equilibrium quantum mechanics applying DFT, and the interaction between the side chain atoms of data units separated by a distance greater than the distance threshold is calculated via molecular mechanics.

10. A method for fragment-based quantum mechanics calculation of protein properties, the method comprising: for each subsequence of three adjacent amino acids in a polypeptide sequence: identifying a first amino acid having a first backbone including a first amino group, a first alpha carbon, and a first carboxyl group, and a first side chain attached to the first alpha carbon; identifying a second amino acid having a second backbone including a second amino group, a second alpha carbon, and a second carboxyl group, and a second side chain attached to the second alpha carbon; identifying a third amino acid having a third backbone including a third amino group, a third alpha carbon, and a third carboxyl group, and a third side chain attached to the third alpha carbon; generating a data unit including data representing the first alpha carbon, the first carboxyl group, the second amino group, the second alpha carbon, the second carboxyl group, the second side chain, the third amino group, and the third alpha carbon; and storing the generated data unit in a database.

11. The method according to claim 10, wherein the first alpha carbon and the first carboxyl group include the N-terminal acetyl group (ACE) of the data unit, the third amino group and the third alpha carbon include the C-terminal N-methylamino group (NME) of the data unit, and The data unit further includes data representing a first peptide bond and a second peptide bond, the first peptide bond being formed between the N-terminal ACE and the second amino group, and the second peptide bond being formed between the second carboxyl group and the C-terminal NME.

12. The method according to claim 10, the method further comprising: calculating a force for each atom in the data unit; and calculating an energy of the data unit.

13. The method according to claim 12, the method further comprising: in a quantum mechanics mode, applying density functional theory to calculate a force for each atom in the generated data unit and to calculate the energy of the data unit.

14. The method according to claim 12, the method further comprising: calculating a force of the polypeptide sequence based on the calculated force for each atom in each data unit of the plurality of data units; and calculating an energy of the polypeptide sequence based on the calculated energy for each data unit of the plurality of data units, wherein the energy of the polypeptide sequence is calculated by summing the calculated energy for each data unit of the plurality of data units and subtracting the energy of the repeating regions shared by adjacent data units of the polypeptide sequence, and the force of the polypeptide sequence is calculated by summing the calculated force for each data unit of the plurality of data units and subtracting the force of the repeating regions shared by adjacent data units of the polypeptide sequence.

15. The method according to claim 14, the method further comprising: calculating interactions between side chain atoms of data units separated by a distance less than or equal to a distance threshold via equilibrium quantum mechanics applying DFT; and calculating interactions between side chain atoms of data units separated by a distance greater than the distance threshold via molecular mechanics.