Methods for de novo protein design without templates
Through the combination of SCUBA energy function and ABACUS method, the problem of main chain structure generation in protein design is solved, high-precision and diverse protein design is achieved, and spontaneously foldable amino acid sequences and structures are generated to meet specific functional needs.
Patent Information
- Application Number
- CN202111197820.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-10-14
AI Technical Summary
The prior art is difficult to generate a main chain spatial structure with high designability in protein design, resulting in too single design results, lacking the rich and diverse natural protein structure, and being unable to fully search for the optimal structure that meets the design needs.
SCUBA, a statistical energy function focusing on the main chain, was used for continuous optimization sampling, and through random dynamics model and simulated annealing optimization technology, a stable main chain structure with extremely small energy in the initial main chain structure was found, and amino acid sequence selection and main chain structure correction iteration were combined with the ABACUS method to generate amino acid sequences that can be spontaneously stable folded.
The high-precision, spontaneous and stable folding protein structure is achieved in de novo design, which can meet specific functional requirements and generate proteins with new topological structures. The amino acid sequence can be successfully expressed and consistent with the design model at the atomic level.
Smart Images

Figure CN113870943B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of protein design, and specifically relates to a method for de novo design of part or all of a protein. This method uses pre-specified features as constraints to generate the main-chain spatial structure of the protein to be designed, and then determines the amino acid sequence of the protein to be designed, so that the protein to be designed possesses the pre-specified features. When generating the main-chain spatial structure of the protein to be designed, this method does not require using known specific protein fragments as structural templates to splice together to produce the protein structure to be designed. Instead, it uses computer optimization to obtain a mathematical model (statistical energy function model) learned from a large number of natural protein structures. Background Art
[0002] There are many protein design methods in the prior art, but each has its own shortcomings.
[0003] Automatically design amino acid sequences given a main chain structure
[0004] This technology takes the spatial structure of the polypeptide main chain given by the user as the target, and automatically selects the amino acid sequence so that the protein molecule with the amino acid sequence can be spontaneously and stably folded into the target spatial structure. The earliest document reporting the successful automatic design of amino acid sequences is Dahiyat et al., Science 278,82-87 (1997). This work and the Rosetta Design (Kuhlman et al., Science 302,1364-1368 (2003)) developed by Baker and collaborators at the University of Washington, USA, both achieve automatic sequence design by optimizing the energy function based mainly on the physical model. One of the main inventors of the present invention and collaborators proposed, verified and published the ABACUS method based on statistical energy function (Xiong et al., Nature Communications 5,5330 (2014) and Bioinformatics 36,136-144 (2020) for automatically designing amino acid sequences for given main chain structures.
[0005] To automatically design amino acid sequences using the above technology, the user is required to first provide a main chain spatial structure with high designability. The so-called "high designability" of the main chain spatial structure refers to the real existence of an amino acid sequence that can spontaneously and stably fold into this structure. Theoretical and experimental studies have shown that among all the possible main chain spatial structures that are chemically reasonable (such as reasonable internal coordinates such as bond lengths, bond angles and dihedral angles, and no spatial steric conflicts between atoms), only a very small proportion of structures have high designability. The above background technology only provides solutions to the problem of how to design amino acid sequences in protein de novo design, and does not consider how to generate a main chain spatial structure with high designability.
[0006] Generate new main chain structure by splicing existing protein structure fragments
[0007] Currently available and experimentally verified methods for generating highly designable main-chain structures from scratch all require using main-chain fragments of known protein structures as local structural templates, and utilizing assembly and combination of different fragments to generate main-chain structures that may be highly designable.
[0008] Since fragment structures can only discretely represent the possible local structural space of polypeptides, the design results are limited by the resolution, accuracy, and completeness of the existing fragment set, making the de novo protein structure too simplistic, overly biased towards idealized local structures, and lacking the rich diversity of natural protein structures. In practical applications, it is impossible to fully search the designable main chain structure space to obtain the optimal structure that meets the design requirements.
[0009] A theoretical method for designing main chain structures based on the statistical energy function of continuous optimization of the main chain structure
[0010] This was proposed by the main inventors and collaborators of the present invention, and the theoretical model and preliminary computational verification were published in the preprint paper bioRxiv 2019,673897,DOI:10.1101 / 673897. In this paper, we disclose a theoretical method for learning high-dimensional, continuous, and analytical neural network statistical energy functions from natural protein structure data. Preliminary verification shows that the statistical energy function SCUBA obtained by this method can be used to drive continuous optimization sampling of the main chain spatial structure. The paper verifies that the natural main chain structure is stable in SCUBA-driven molecular dynamics simulations. In addition, the artificially constructed initial main chain structure is optimized by simulated annealing using SCUBA, and the resulting main chain structure can have a secondary structure stacking structure that is highly similar to that of the natural protein. Summary of the Invention
[0011] The present invention solves the problem of how to design part or all of a protein from scratch, that is, how to specify part or all of the amino acid sequence of the protein to be designed so that the protein can spontaneously fold into a stable spatial structure in which the amino acid residues have the expected spatial arrangement, thereby enabling the protein to be used to meet certain functional requirements.
[0012] The present invention provides a complete de novo computational protein design tool chain (including a design workflow, a computer-programmed model used in the process, and model parameters that make the computational results physically meaningful). The key steps of the tool chain include de novo generation of a physically designable main-chain spatial structure under undetermined sequence conditions (the meaning of physically designable main-chain structure is that the amino acid sequence that spontaneously folds into the structure actually exists), followed by an automatic sequence design algorithm to preliminarily select an amino acid sequence for the structure, further optimizing the target main-chain structure based on the sequence, and then reselecting the sequence. After several rounds of iteration, the final amino acid sequence that can spontaneously and stably fold into the target structure is obtained. During the above structure and sequence optimization process, the algorithm can make the automatically generated spatial structure and sequence conform to specific constraints proposed by the user, so that the designed protein meets certain user-specified functional requirements.
[0013] The protein de novo design method of the present invention divides the de novo design of all or part of a protein into two stages. The first stage is to generate a high-precision main chain structure with high designability, and the second stage is the iterative stage of sequence selection and main chain correction.
[0014] In the high-precision main chain structure design stage, the backbone-centered statistical energy function SCUBA is used as the objective function. When the sequence is undetermined, the main chain structure is optimized by computational methods to find a stable main chain structure with extremely small total SCUBA energy.
[0015] The energy term in SCUBA is a continuous analytical function derived from training non-redundant native protein structure data using the Nearest Neighbor Counting-Neural Network (NC-NN) method. Therefore, optimization techniques such as stochastic kinetic models and simulated annealing can be used to find the energy-minimizing backbone on the SCUBA energy surface within a certain range of initial backbone structures within a limited computational time.
[0016] During the main chain structure optimization phase, the search range for conformations can be modified by controlling the optimization algorithm parameters (e.g., simulated annealing temperature range, simulation time, etc.). User-defined constraint energy terms can be added during the optimization process to tailor the resulting main chain structure to specific user requirements.
[0017] Due to the complexity of the energy surface, loop regions (regions connecting secondary structures) in the main chain structure obtained only through continuous simulation optimization may still have locally high energy, reducing the designability of the main chain structure. This can be achieved by regenerating a large number of closed loop structures through computational methods, optimizing the SCUBA energy using each of these as the initial structure, and selecting the optimized main chain with low energy.
[0018] During the aforementioned main-chain optimization process, the amino acid sequence is unknown. Since the main-chain conformation obtained by optimizing the SCUBA energy is insensitive to the side-chain type of the amino acid residues, when optimizing the main-chain SCUBA energy, the model can include residues with simple side-chain structures, such as leucine and valine, to act as space-occupying side chains.
[0019] During the iterative sequence selection and main-chain structure modification phase, the present invention first uses the main-chain structure generated in the first phase to determine the amino acid sequence using methods described in the prior art for designing amino acid sequences given a given main-chain structure (e.g., the ABACUS method). After replacing the side chains with the designed sequence, the SCUBA energy of the overall structure is optimized to generate a main-chain structure modified according to the specific sequence. This modified main-chain structure is then used to redesign the sequence. Multiple rounds of sequence design and main-chain structure modification iterations can be performed, and the optimal sequence generated in this process is selected according to one or more meaningful criteria (e.g., the lowest ABACUS sequence energy) as the final design result for subsequent experimental analysis.
[0020] The present invention provides the following technical solutions:
[0021] In one aspect, the present invention provides a method for de novo protein design, characterized in that the method comprises:
[0022] a. Produce a main chain structure with high designability;
[0023] b. Iteration of amino acid sequence selection and backbone modification.
[0024] In some embodiments, generating a main chain structure with high designability includes: using a main chain focused statistical energy function SCUBA as an objective function, optimizing the main chain structure using a computational method when the sequence is to be determined, and finding a stable main chain structure with extremely small total SCUBA energy.
[0025] In some embodiments, the energy term in SCUBA is a continuous analytical function obtained by training non-redundant natural protein structure data using a neighbor counting-neural network method, and then using a stochastic kinetic model and simulated annealing, after a limited computing time, the optimized main chain represented by the minimum point on the SCUBA energy surface is found starting from the initial main chain structure.
[0026] In some embodiments, the scope of the conformation search is varied by controlling the parameters of the optimization algorithm (eg, the temperature range of simulated annealing, the length of simulation time).
[0027] In some embodiments, during the optimization process, a user-defined constrained energy term is added so that the generated main chain structure meets specific user requirements.
[0028] In some embodiments, local resampling and optimization are performed, and a large number of closed loop region structures are regenerated by computational methods for the loop regions in the main chain structure obtained by continuous simulation optimization, the SCUBA energy is optimized, and the optimized main chain with low energy is selected.
[0029] In some embodiments, the model used to optimize the main chain structure SCUBA energy includes residues with simple side chain structures (eg, leucine, valine) that play a role in side chain space occupation.
[0030] In some embodiments, the main chain structure generated in step a is used to design an amino acid sequence, and the amino acid sequence is preferably designed using the ABACUS method.
[0031] In some embodiments, after replacing the side chains with the designed amino acid sequence, the SCUBA energy of the overall structure is optimized to produce a main-chain structure modified according to the specific sequence.
[0032] In some embodiments, the amino acid sequence is redesigned using the corrected main chain structure, and after multiple rounds of amino acid sequence design-main chain structure correction iterations, the optimal amino acid sequence is screened, preferably selecting the optimal amino acid sequence based on the lowest ABACUS sequence energy.
[0033] In another aspect, the present invention provides a device for de novo protein design, characterized in that the device comprises:
[0034] Main chain structure generation module, used to generate a main chain structure with high designability;
[0035] The amino acid sequence selection and main chain correction iterative module is used to select the amino acid sequence and perform iterative correction on the main chain.
[0036] In some embodiments, the main chain structure generation module includes a calculation module, which uses the statistical energy function SCUBA of the focused main chain as the objective function. When the sequence is to be determined, the main chain structure is optimized by a calculation method to find a stable main chain structure with extremely small total SCUBA energy.
[0037] In some embodiments, the computing module further includes a first storage module, which stores energy terms in SCUBA. The energy terms in SCUBA are continuous analytical functions obtained by training non-redundant natural protein structure data using a neighbor counting-neural network method. The computing module uses a random dynamics model and a simulated annealing algorithm to find the optimized main chain represented by the minimum point on the SCUBA energy surface after calculation, starting from the initial main chain structure.
[0038] In some embodiments, the calculation module further includes a parameter setting module for changing the search range for conformations by controlling parameters of the optimization algorithm (eg, the temperature range of simulated annealing, the simulation time length).
[0039] In some embodiments, the main chain structure generation module also includes a local resampling and optimization module, which is used to use a computational method to regenerate a large number of closed loop region structures in the loop region of the main chain structure obtained by continuous simulation optimization, optimize SCUBA energy, and select the optimized main chain with low energy.
[0040] In some embodiments, the amino acid sequence selection and main chain correction iterative module includes a second storage module, which stores an ABACUS program for designing an amino acid sequence using the generated main chain structure.
[0041] In some embodiments, the amino acid sequence selection and main chain correction iterative module also includes an optimization module, which is used to optimize the SCUBA energy of the overall structure after replacing the side chain with the designed amino acid sequence, generate a main chain structure corrected according to the designed amino acid sequence, and redesign the amino acid sequence with the corrected main chain structure. After multiple rounds of amino acid sequence design-main chain structure correction iterations, the optimal amino acid sequence is screened, and the optimal amino acid sequence is preferably selected based on the lowest ABACUS sequence energy.
[0042] definition
[0043] ABACUS: A set of algorithms and programs for designing fixed-mainchain amino acid sequences based on statistical energy functions. It is used to select the appropriate amino acid type for each residue position in a given protein backbone, so that the designed amino acid chain can spontaneously fold into the given protein backbone structure in the appropriate physical and chemical environment.
[0044] SCUBA: A statistical energy and neural network-based protein backbone optimization and design algorithm and program that focuses on the backbone and is independent of specific amino acid side chains. It can sample and optimize non-designable protein backbones, independent of specific side chain types, to generate designable protein backbone structures. Combined with a fixed backbone protein sequence design program, it enables template-independent de novo protein design.
[0045] Beneficial effects
[0046] The main distinguishing features of the present invention from the background art Rosetta Design are as follows: First, Rosetta Design requires existing protein structure fragments as templates and adopts a fragment splicing method to produce a main chain structure with high designability. The present invention is based on the inventor's original statistical energy function SCUBA that focuses on the main chain. It does not use the splicing of existing structure fragments, but instead continuously optimizes and samples the main chain structure to generate a main chain structure with high designability that can be used for automatic selection of amino acid sequences from scratch. Secondly, the present invention uses a statistical energy model that is completely different from the Rosetta Design energy model principle. When the sequence is pending, it can automatically optimize the total energy to obtain a high-precision main chain structure that can be used for automatic selection of subsequent amino acid sequences.
[0047] The main differences between the theoretical background and preliminary computational verification of the SCUBA method disclosed by the inventors in the present invention and the preprint are as follows:
[0048] (1) The background technology only includes the SCUBA algorithm principle, but does not include the specific parameters of the neural network energy term. The SCUBA model using specific parameters is unique to the present invention.
[0049] (2) Although the secondary structure stacking of the main chain structure obtained by a round of simulated annealing optimization using the algorithm in the background art is highly designable, in local areas such as the loop region, it is likely that it is only in the local high energy area on the energy surface and has not been fully sampled and optimized, resulting in the overall structure being undesignable. Therefore, it is necessary to go through the steps of local resampling and optimization proposed in the present invention to truly obtain a main chain with high overall designability.
[0050] (3) The background art only compares the similarity between the main chain structure generated by calculation and the natural structure, and does not describe how to further perform sequence design to obtain a real amino acid sequence that can be experimentally verified to be actually foldable. The present invention points out that after optimizing the main chain under the condition that the sequence is to be determined, it is necessary to use a computational method such as ABACUS to select the amino acid sequence based on the main chain structure, and then use SCUBA to optimize the main chain with side chains, so as to further improve the accuracy of the main chain structure and sequence design, and obtain the final amino acid sequence through multiple rounds of iteration. Therefore, the technical solution disclosed by the present invention is a complete protein design tool chain, which has been experimentally proven to be able to obtain an amino acid sequence that can ultimately spontaneously fold into the target structure.
[0051] Compared to RosettaDesign, the only publicly verified background technology outside the present invention that can design backbones from scratch, the present invention uses a unique statistical learning method to obtain statistical energy terms that can accurately characterize the high-dimensional correlations in designable backbone structures. This allows for continuous and extensive searches of the backbone structure space under the premise of uncertain amino acid sequences, generating highly designable backbone structures. In practical applications, the backbone design algorithm provided by the present invention is conducive to generating highly designable backbones that best meet specific application requirements (which can be defined by constraints in the optimization process or by screening conditions for the optimized backbone).
[0052] By using the method of the present invention, multiple artificial proteins of different structural types have been designed from scratch, six of which have completely new overall topological structures that do not exist in nature. The high-resolution crystal structures of these proteins are consistent with the designed models at the atomic level.
[0053] (4) Proteins are the main executors of life functions and are long-chain biomacromolecules composed of different types of amino acids connected in a specific order. The order in which amino acids are arranged in each protein, i.e., the amino acid sequence, determines whether the protein can form a stable three-dimensional structure, what kind of three-dimensional structure it will form, and what function it will ultimately perform. Randomly arranged amino acid sequences are difficult to fold into a stable three-dimensional structure. Currently, almost all proteins that can be folded into stable three-dimensional structures are natural proteins, and their amino acid sequences are formed through long-term natural evolution. The present invention can design a completely new protein structure from scratch, whose amino acid sequence can be folded into a stable protein structure and can be successfully expressed. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A sketch of the designed protein and the construction of the initial structure are shown. In the sketch structure, blue arrows and orange cylinders indicate the expected β-strand and α-helical segments and their locations, respectively. Bold numbers indicate their expected residue numbers, and semi-transparent rectangles indicate the planes in which the secondary structure segments lie. Double arrows indicate the distance between points or parallel lines, and black x marks indicate the locations of the endpoints of the secondary structure segments.
[0055] Figure 2 An example of the process of optimizing protein backbone structure using SCUBA SD simulation and optimizing SCUBA total energy using simulated annealing to obtain a highly designable backbone structure is shown.
[0056] Figure 3The main chain optimization-sequence design iteration process according to an embodiment of the present invention is shown. The M, N, C, and S listed in the flowchart are configurable parameters. In the design of the α+β protein in the example, M=5 (each initial optimized main chain uses different random initial velocities in SD to obtain 5 relaxed main chain structures), N=10 (each main chain structure obtains 10 different loop region optimized structures), C=3 (three main chain optimization-side chain design iterations), and S=1 (each main chain structure selects a designed amino acid sequence with the lowest energy).
[0057] Figure 4 The left image shows the superposition of the crystal structure (blue) and the designed structure (gray and green) of the native protein EXTD-3. The precisely designed H3-H4 loop region is marked in the middle The upper right figure shows the H3 and H6 helices in the initial model of the designed structure. During the subsequent main chain optimization, the H3 helix spontaneously formed a complete helix. The lower right figure shows the natural protein (purple) that is most similar to the designed protein crystal structure (blue).
[0058] Figure 5 The left and right sides of the upper half of the figure show the superposition effect of the crystal structure (blue) of the de novo designed protein XM2H (left) and AM2M (right) with the designed structure (green), and the RMSD is (XM2H) and (AM2M). The geometric figure in the middle of each figure represents the topological connection of the secondary structure of the α+β protein. ⊙(△) and ⊕(▽) indicate outward and inward helices (β chains), respectively, and the arrows indicate the direction of the loop regions connecting adjacent fragments (C-terminus to N-terminus). The left and right sides of the lower half of the figure respectively show the crystal structures (blue) of the XM2H (left) and AM2M (right) protein loop regions and the superposition of their electron density with the designed structure (green), showing that the structure of the loop region is precisely designed; the middle of the lower half of the figure shows the superposition of the crystal structures of XM2H (blue) and AM2M (brown), showing that their secondary structure regions are highly overlapping, but the structures of the loop regions are very different.
[0059] Figure 6 The results of de novo design of all-α proteins are shown. The upper part of the figure shows the superposition effect of the crystal structure (blue) of the de novo designed proteins H4A1R (left), H4A2S (middle) and H4C2R (right) with the designed structure (green), and the RMSD is (H4A1R), (H4A2S) and (H4C2R). The geometric figure in the center of each figure represents the topological connections of the α-helical segments of the all-α protein. ⊙ and ⊕ indicate outward and inward α-helices, respectively, and arrows indicate the direction of the loop regions connecting adjacent segments (C-terminus to N-terminus). The lower half of the figure shows the crystal structure (blue) of the loop region of the H4A1R (left), H4A2S (center), and H4C2R (right) proteins, and the superposition of their electron density with the designed structure (green), showing that the structure of the loop region is precisely designed.
[0060] Figure 7 The results of random helical protein design are shown. The upper part of the figure shows the superposition effect of the crystal structure (blue) and the designed structure (green) of the de novo designed random helical protein D12 (left), D22 (middle) and D53 (right), respectively. The RMSD are (D12), (D22) and The top half of the figure shows the superposition effect of the designed structures (green) of the designed proteins D12 (left), D22 (center), and D53 (right) and the natural protein structure (purple) that is most similar to them. DETAILED DESCRIPTION
[0061] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0062] Example Construction of SCUBA Energy Function and Application of the Present Invention Successfully Designed De Novo Proteins
[0063] 1. SCUBA Energy Function
[0064] 1.1. Energy term
[0065] The overall SCUBA energy is defined as the sum of the covalent, steric, and statistical terms,
[0066] E total (r) = E covalent (r)+E steric (r)+E statistical (r), (S1)
[0067] Where r represents the coordinates of the main chain non-hydrogen atoms. The covalent bond energy term E covalent (r) is the sum of harmonic functions, which depends on individual bond lengths, bond angles, or inappropriate dihedral angles. For backbone atoms, the steric repulsion energy term E steric (r) is the sum of the Lenard-Jones potentials for the atoms, with the tail of the attractive term removed. The statistical energy term E statistical(r) is the sum of different types of statistical energies derived using the Neighbor Counting-Neural Network (NC-NN) method. Trained on a dataset of 12,465 non-redundant natural protein structures, each NC-NN energy type describes the sum of the effects of different types of interactions on the relative distribution of a given set of sequential or spatially localized structural variables.
[0068]
[0069] in
[0070]
[0071] here express The ith subset of structural variables on which the term depends. The index i varies at all main-chain positions if the energy type depends on sequential local variables, or at pairs of main-chain positions if the energy type depends on spatial local variables. The types of NC-NN energy functions and the structural variables on which they depend are shown in Table 1, providing a comprehensive coverage of determining main-chain designability interactions in a way that is independent of or insensitive to side-chain type. We note that, unlike and compared to, The energy term inherently carries redundant information about the local main-chain conformation. However, The term explicitly depends on the interatomic distances and thus improves the main chain hydrogen bond distance distribution in the helix. and energy term, the distribution of main-chain hydrogen bond distances in the helix will be somewhat too broad.
[0072] Table 1 NC-NN statistical energy types in SCUBA
[0073]
[0074]
[0075] Initially, the SCUBA energy function was defined without considering any explicit side chains. However, we found that the main chain atom positions optimized on such a main chain-only energy surface were too "fuzzy" for subsequent amino acid sequence design (too many positions on such main chains had alanine or glycine selected). Therefore, we included explicit side chains and, in addition to the usual covalent bond energy, there were two additional energy terms. One is the one listed in Table 1. The sum of the terms, and the other is the side chain stacking term E SC-packing(r), the side chain stacking term is the sum of the side chain-side chain and side chain-main chain terms, and each atom pair term is related only to simple interatomic distances, consisting of a repulsive part of the Lennard-Jones form and an attractive part of the inverted Gaussian form. The resulting single potential well distance correlation function has two adjustable parameters. The first is the minimum energy distance, which is related to the optimal stacking distance and has been treated as atom type specific and taken from the previous ABACUS sequence design model. The second is the well depth, which for simplicity is treated the same for all atom pairs and is incorporated into the weight parameter w SC-packing In order to combine this side chain stacking term with other energy terms, the side chain stacking term is added to the main chain sampling and optimization before sequence design. During the main chain sampling and optimization before sequence design, explicit side chain-related terms are either not considered or calculated using relatively uncharacterized side chain types (such as LVG sequences). These side chains are then used to average out various side chain-related factors that affect the physical plausibility or designability of the main chain structure. These factors include the chiral covalent structure of the side chain, the main chain-dependent side chain conformational preferences, and the side chain excluded volume in the main chain-side chain and side chain-side chain stacking.
[0076] 1.2. Constructing NC-NN energy function
[0077] 1.2.1. Overall Approach
[0078] For a given set of structural variables, the NC-NN energy function is learned from a set of training structures in two stages. In the first, or NC, stage, a single-point estimate of the statistical energy is obtained in the high-dimensional space of structural variables using simple nonparametric, kernel-density estimation. In the second, or NN, stage, a neural network representing the statistical energy as an analytical function of the structural variables is trained on a large set of single-point energies estimated by the NC.
[0079] 1.2.2. Statistical energy is defined as the ratio between probability densities
[0080] To formally describe the NC estimation of single-point energies, for a set of correlated structural variables Q, the effective energy function learned from the observed data is defined as
[0081]
[0082] where ρ observed (Q) is the probability density of the observed data distributed in the Q space, and ρ reference (Q) is the probability density of the same set of variables, the difference from the former is that ρ reference (Q). The data points come from an ideal reference system, and there is no interaction between the data points described by e(Q). Table 1 shows the various types of detection points Q used in the NC-NN energy function of SCUBA. probe distribution of their respective reference systems.
[0083] 1.2.3. Single-point statistical energy estimation based on neighbor counting
[0084] In Q space, we can calculate Get from the data A single point estimate of Represents the Q probe The number of adjacent observation data points, and Represents the Q probe The number of neighboring reference data points, all of which are distributed according to the reference distribution ρ reference (Q) is calculated. In fact, given the observation data point and the reference data point, It can be approximated with the help of kernel function as
[0085]
[0086] Because the observed and reference neighbor counts are estimated using the same kernel function, the ratio is insensitive to the exact choice of kernel function,
[0087]
[0088] If necessary, the similarity criterion (or the radius of the kernel) in Eq. S6 can be adaptively selected to optimize the Q in the region with dense observation data. probe More stringent, Q for loose areas probe More relaxed. On the one hand, a smaller kernel radius results in a higher resolution model; on the other hand, since fewer points are counted as neighbors, the resulting model is accompanied by greater statistical uncertainty. Therefore, adaptive kernels can strike a balance between resolution and statistical uncertainty based on the observed data.
[0089] 1.2.4. Training a Neural Network to Represent a Statistical Energy Function
[0090] Each neural network in SCUBA consists of an input layer, a hidden layer, and a single-node output layer. Nodes in adjacent layers are fully connected. The nodes in the hidden layer are activated using a logistic function. Mathematically, each network performs the following transformation from input to output:
[0091]
[0092] Where N1 is the input dimension size or number of nodes in the input layer, and N2 is the number of nodes in the hidden layer. and are the weights of the connections from the first layer to the second layer and from the second layer to the third layer, and b2→3 are the biases of the second and third layer nodes, respectively.
[0093] In general, the input vector x encodes the structural variables on which the effective energy contained in Q depends, as follows: each angle variable θ is encoded as a series of trigonometric function values (sin kθ, cos kθ, k = 1, 2, 4); each interatomic distance variable d is converted into a series of c i is the center and the standard deviation is σ i Gaussian function
[0094] As in many machine learning practices, the scheme and parameters used to encode the input and the number of nodes in the neural network are manually selected through trial and error, while the weight and bias parameters are learned from a large number of single-point encodings and their energy values estimated by the aforementioned NC.
[0095] For each final trained NC-NN energy function, we have verified that the NN model can not only satisfactorily reproduce the energy estimated by the NC, but also The distribution of the sampled data points is calculated to closely mimic the corresponding distribution of the observed data points in Q-space.
[0096] 1.3. Calibration of energy weight parameters
[0097] To sample or optimize protein structures on the SCUBA energy surface, we applied stochastic dynamics (SD) simulations, in which the Langevin equations of motion are integrated to obtain the temporal trajectories of atoms moving in a friction medium at a given temperature. Simulations of a test set of natural proteins starting from natural X-ray crystal structures have been used to calibrate the entire SCUBA energy function (i.e., w in Eq. S2). type ) in the few undetermined energy weight parameters. The goal of calibration is to stabilize the native conformational state. In addition, we want the overall interaction strength to be as weak as possible, as long as the native structure can be r =1 to maintain its stability.
[0098] We perform calibration using an approach that is conceptually similar to force field refinement using thermodynamic cycles in conformational space. Briefly, to refine a set of experimental parameters, the test protein is simulated with two sets of experimental parameters. One set consists of unrestricted simulations, in which the structure is allowed to deviate freely from the starting native structure according to an uncalibrated energy function. The other set consists of restricted simulations, in which sampling in conformational space is limited to a region adjacent to the native state (here we constrain the RMSD of the secondary structure backbone atoms from the native structure to be less than ). Various energy terms averaged over the conformations sampled in the two sets of simulations are then compared. If the energy term is systematically lower in the conformation with the larger RMSD, its weight is decreased in the next round of test simulations. Otherwise, its weight can be kept constant or experimentally increased.
[0099] We estimated an initial set of weights by first considering only two small, globular test proteins: one entirely α (PDB ID 3132) and the other entirely α+β (PDB ID la6j, chain A). The resulting weights were then further refined using 33 manually selected test proteins. These 33 proteins are relatively small, soluble, globular proteins lacking disulfide bonds and belonging to diverse fold types. Furthermore, the PDB structures for these 33 proteins are all relatively high-resolution X-ray crystal structures.
[0100] After obtaining a set of calibrated energy weights ( w site-pair =0.32, w local-HB =0.6, w rotamer =2.4, w sc-packing =3.1) after we r = 1, each weight was individually and systematically scaled between 0.2 and 2.0, the other weights were fixed, or the total energy of the simulation was scaled between 0.25 and 2.5. The results confirmed that the native structures of most (26 out of 33) tested proteins are close to the minimum of the SCUBA energy (the average backbone RMSD of the structures sampled from each native structure in the SCUBA SD simulation is less than ), the weights are chosen to provide sufficiently strong interactions to r =1, the stability of the natural structure is maintained.
[0101] After calibrating the protein weights using specific native side chains as described above, we performed another set of simulations in which we varied the side chain types using LVG sequences (leucine in the α-helix, valine in the β-strand, and glycine in the loop region) to test whether the native backbone structure remained stable with non-native side chains in SCUBA SD simulations. The results showed that even with LVG sequences (i.e., non-native side chains), the native backbone structure remained relatively stable in SCUBA SD simulations.
[0102] 2. Redesign of the main chain
[0103] 2.1. Constructing the initial structure
[0104] The basic strategy for designing backbones from scratch using SCUBA is to start with an initial structure constructed from a set of user-defined structure sketches for SCUBA SD simulations. Figure 1 Based on a consistent initial structure, the portion of conformational space associated with the given sketch is searched for a physically plausible backbone.
[0105] In the present work, the sketches were described according to the "periodic table" model of protein structure described by Taylor et al. (WR Taylor, A "periodic table" for protein structure. Nature 416, 657-662 (2002)). In this model, well-folded proteins are abstracted into secondary structure elements, organized into roughly parallel layers, each layer containing multiple β sheets or approximately parallel or antiparallel helices. For each sketch, we predefined the number of secondary structure fragments, their approximate sizes, and their order and orientation in different abstraction layers. Based on the sketch, we first generated peptide segments in the local conformation of the helix or β strand at approximate relative positions, and then constructed loops to connect them (see Figure 1 ).
[0106] Secondary structure fragments are placed in one of two ways. The first method is to place the starting or ending atom of each fragment at the intersection of a two-dimensional planar grid, and then grow the entire fragment perpendicular to the grid plane (perpendicular to the plane of the secondary structure layer). The endpoints of fragments in the same secondary structure layer fall on the same straight line, and the lines corresponding to different layers are parallel and spaced apart. The distance between adjacent spirals in the same layer is approximately The distance between adjacent β-strands is approximately To grow each fragment, internal backbone torsion angles are randomly sampled according to the fragment's assigned secondary structure type. A second method for placing secondary structure fragments is to use a computer graphics system to interactively place idealized helices or β-strands at approximate locations. In our examples, the initial structure of the H2E4 protein was generated using the first method, and the initial structures of the EXTD and H4 proteins were generated using the second method.
[0107] To generate loop structures to connect two pre-placed secondary structure fragments, we first used backbone torsion angles randomly selected according to the Ramachandran backbone angle distribution of protein helices. Generate an unclosed loop structure and then try to close the loop using a kinematic loop closure algorithm (EA Coutsias, C. Seok, MP Jacobson, K A Dill, A kinematic view of loop closure. Journal of Computational Chemistry 25, 510-528 (2004)). The loop length is first set to 3 and then gradually increased until a solution for the closed loop can be found in 1000 loop generation attempts for a given loop length. Note that the sizes of the α-helix and β-strand segments defined in the sketch are only approximate, because the extension or contraction of the predefined secondary structure segments can occur spontaneously in subsequent structure sampling and optimization.
[0108] 2.2 Optimizing the main chain structure through simulation
[0109] After the initial backbone was constructed, the backbone was optimized using SCUBA-driven SD simulations in two sub-stages, with a backbone-only model used in sub-stage 1 and an LVG sequence used in sub-stage 2.
[0110] In sub-phase 1, the local conformation of the specified helical or β-strand segment is constrained by the secondary structure type and the steric interactions involving the atoms in the loop region are multiplied by a factor of 0.01. These treatments simplify the overall energy map and reduce the search space, so that approximate SCUBA energy minima that meet the user-defined sketch can be effectively located. The simulations in sub-phase 1 include the simulation at low temperature (relative temperature T r = 0.1) and large friction coefficient (friction coefficient γ = 5ps -1 ). This simulation was used to remove any possible stress in the initial backbone. The backbone obtained in the previous step was then optimized by a 60 ps simulated annealing simulation. In 6 cycles, the relative humidity was varied between 2.0 and 0.5, with each cycle including T r = 2.0 4ps, followed by T r 1ps gradually decreases from 2.0 to 0.5, then T r = 0.5 for 5 ps. To compensate for the effects of thermal expansion at higher temperatures during simulated annealing, the radius of gyration of the entire structure was constrained. Using this scheme, most simulations starting with different initial structures constructed for the same sketch can produce backbone structures that meet the sketch constraints. Furthermore, the SCUBA energy generally converges within the last few simulated annealing cycles.
[0111] In sub-stage 2 of the main-chain optimization, local conformational secondary structure type restraints were removed and interactions involving loop residues were restored to full strength. The SD simulations in this sub-stage included a 4 ps relaxation (γ = 5 ps -1 , T r= 0.2), followed by 10 cycles of simulated annealing for a total of 120 ps, in which the temperature was varied between 2.0 and 0.5, while applying a radius of gyration constraint. The backbone structure was finally refined by another 10 cycles of simulated annealing (SD) for a total of 120 ps (temperature cycled between 0.5 and 0.2). In most of the final SD simulations, the backbone structure and the total SCUBA energy fluctuated with relatively small amplitudes, indicating that a stable minimum of the SCUBA energy surface was reached (see Figure 2 ).
[0112] 2.3. Ring area resampling and optimization
[0113] The designed backbone undergoes further extensive loop resampling and optimization to produce a designable backbone. The loop resampling and optimization are based on the above two sub-stage optimization of the overall structure.
[0114] During the loop resampling process, regular secondary structure regions were fixed. The start and end positions of each loop were systematically varied within three residues before or after the secondary structure fragments at both ends. For the H2E4 backbone, the loop length was gradually increased from 3 residues to the loop length before resampling. For the H4 backbone, we considered the minimum loop length that could find a loop closure solution, which was one residue longer than the respective minimum length. For each loop, SCUBA SD simulation (T r = 0.5) were optimized individually for 1000 randomly generated and closed starting ring configurations until the energy fluctuations became small.
[0115] From the loop sampling results, a set of candidate loop structures was selected for each loop. For the H2E4 protein design, this set of loop structures consisted of the 10 lowest-energy non-redundant structures for the optimal loop length (meaning that the per-residue SCUBA energy of the lowest-energy structure for that loop length was lower than the corresponding energies for other loop lengths). Subsequently, from the non-redundant loop structures sampled for loops one residue shorter than the optimal loop length, the highest-energy loop structure among the 10 lowest-energy structures selected from the set with a single-residue SCUBA energy lower than the optimal loop length was selected as a candidate structure for that loop. To obtain a complete backbone structure encompassing all loops, different combinations of candidate loop structures were sampled, and the resulting total SCUBA energies were compared between different combinations. For each backbone used for loop resampling and optimization, the 10 final backbone structures with the lowest energy were selected for subsequent sequence selection. This resulted in 500 H2E4 backbone structures for subsequent sequence selection (50 backbones optimized from different initial structures × 10 loops reoptimized).
[0116] For the design of H4 protein, loop resampling and optimization were performed from 116 main-chain structures of 30 different topologies of initial main-chain optimization. For each H4 protein main-chain containing 3 loop regions, 1000 loop region optimized complete main-chain structures were generated by systematic combination of 10 candidate loop region structures. From the total 116,000 main-chain structures obtained, 3,382 structures with lower main-chain interaction energy (based on the total number of loop region residues) were selected. energy) for subsequent sequence selection.
[0117] 3. Use ABACUS2 for amino acid sequence selection and Rosetta biased forward folding method for screening
[0118] 3.1. Main chain relaxation and sequence selection iteration
[0119] The backbone optimized using the LVG sequence (and after loop resampling for the H2E4 and H4 proteins tested in the final experiments) was used for ABACUS2 sequence selection in the first iteration. The backbone was then relaxed using SCUBA-driven SD simulations and the specific amino acid sequence designed for ABACUS2.
[0120] In repeatedly applying this backbone relaxation and sequence selection process, we found that, in addition to the sequences selected in the first iteration, sequences selected for backbone relaxation in different later iterations had greater than 50% amino acid sequence similarity and similar ABACUS2 energies. On the other hand, the frequency of alanine and glycine in the selected sequences gradually increased with increasing iteration number. This phenomenon is likely due to the fact that small side-chain residues introduced early in the iterative process may cause the side-chain space at the corresponding position to be compressed in subsequent backbone relaxation, and then larger residues are no longer selected for this position in later iterations.
[0121] Since more and more small side-chain residues will be introduced with increasing iteration number, and since the sequences selected in different iterations have relatively high similarity, we limited the number of backbone relaxation and sequence selection iterations to two to three.
[0122] More specifically, the protocol used to generate the H2E4 sequence is as follows: Given an optimized target backbone, a round of amino acid sequence design is performed using ABACUS2. In this round of sequence design, the side chain atomic radius parameter of ABACUS2 is multiplied by 0.9. This operation aims to introduce larger amino acid side chains into the protein structure than normal. The backbone is then relaxed using a SCUBA-driven SD simulation. The simulation process is divided into two steps. The first step is a 4 ps low-temperature SD simulation (T r =0.2,γ=5ps -1), followed by 20 ps of simulated annealing SD(T r From 1.0 to 0.2, γ = 0.5ps -1 ). Five different sets of random initial velocities were used for the SD simulation of each backbone to generate five different relaxed backbone structures. Each structure was subjected to two iterations of ABACUS2 sequence selection (without reduction of the atomic radius parameter), followed by backbone relaxation using the same scheme as above (but using only one set of random initial velocities). The five backbones obtained were used for the final ABACUS2 sequence selection, and 20 sequences were selected for each backbone, with the one with the lowest energy retained (see the procedure for the procedure). Figure 3 During iterative designs using the standard ABACUS2 parameter set, we found that threonine was overrepresented at sites in the solvent-accessible surface region of the β-strand, while alanine was overrepresented at sites in the α-helix, potentially increasing the aggregation tendency of the designed proteins. To test this hypothesis, we designed additional sequences in which the ABACUS2 residue-type-specific reference energy parameters were adjusted in three different ways to increase the probability of selecting alternative polar residues at these positions. Because the sequences designed by the different ABACUS2 reference energies differed only at the solvent-exposed positions of regular secondary structure elements, we believe that the differently selected sequences would be equally plausible in subsequent computational screening and experimental characterization.
[0123] The specific protocol for generating EXTD and H4 sequences is as follows: Given an optimized backbone, 10 sequences were designed using ABACUS2 (standard parameters), and the one with the lowest energy was retained. The backbone was then relaxed through a 4 ps simulated degradation cycle, decreasing the temperature from 1.0 to 0.2. The backbone relaxation sequence selection was then repeated three times using the same SD protocol. For each designed backbone, the sequence with the lowest ABACUS2 energy across all iterations (not necessarily from the last iteration) was retained.
[0124] 3.2. Sequence of computational screening designs
[0125] Before experimental characterization of the sequences designed for the H2E4 and H4 backbones with reoptimized loop regions, computational screening was performed using the Rosetta biased forward folding method. In these screening processes, fragment conformations predicted in advance based on the sequence are assembled into ensemble structures to predict whether a given sequence is likely to fold into a specific target structure, using only three predicted conformations per fragment that have the lowest RMSD of all predicted conformations with the corresponding fragment of the target structure. Due to this limited selection of possible fragment conformations, only a small number of ensemble structures need to be generated and the lowest RMSD of these structures from the target can be used as a judgment criterion. For each ABACUS2 selected sequence of the H2E4 (H4) structural protein, 50 (200) folding simulation structures were generated. The lowest RMSD with Rosetta biased forward folding and the Rosetta energy score of the lowest RMSD model were considered as screening criteria for the final experimental testing. For the H2E4 structural protein, 2000 selected sequences (4 sequences were selected for each of the 500 main-chain structures using different ABACUS2 reference energies) were computationally screened by folding simulations. From each set of sequences designed using different ABACUS2 reference energies, 6 or 9 sequences were selected for experimental identification after folding simulation screening. From the 3382 ABACUS2 designed sequences covering the four different H4 structural drafts, 8 sequences (1 for each draft) were selected by folding simulation screening. Figure 2 sequences) were used for experimental characterization.
[0126] 3.3. Design of novel helical proteins
[0127] First, 10,000 initial helical segment arrangements were generated by randomly sampling the dihedral angles of the helical region in the Ramachandra graph, randomly generating 6 peptide segments with a length of 10 to 25 residues, randomly rotating them, and then randomly translating them so that their center of mass falls within a radius of inside the sphere.
[0128] Each of these initial structures was then optimized using two stages of SCUBA-driven simulated annealing, obtaining a variety of reasonably packed helical arrangements using the following protocol: The first stage consisted of a 50 ps SD simulation with the temperature varied between 1.0 and 0.5. In this stage, only backbone atoms were included; since the initial structures contained many interatomic clashes, steric repulsion in SCUBA was disabled. The second stage consisted of a 50 ps SD simulation with the temperature varied between 2.0 and 0.5. During this stage, the amino acid type of all residues was set to leucine.
[0129] The resulting six-segment helical arrangement was used to construct proteins with novel folding types that allow for designing backbones. For a given configuration, we considered each pair of helical segments in that configuration in turn and attempted to find ways to connect them to a loop structure of three or four residues, allowing for truncation of either segment. This was achieved by considering all possible combinations of N-side attachment (truncation) points on one helix with C-side attachment (truncation) points on another helix.
[0130] The plausibility of a pair of junction (truncation) points was scored empirically as follows. We extracted all helix-loop-helix motifs with three to four loop residues from non-redundant PDB structures and collected the following set of six-atom coordinates: if i is a residue that starts a loop on helix A and j is a residue that ends a loop on helix B, then we collected the coordinates of six Cα atoms for residues i, i-2, and i-4 from helix A and residues j, j+2, and j+4 from helix B. We applied kernel density estimation to obtain the local density in the relative geometric space between the six atoms. The local density of the native loops was estimated, ranked, and then used as a reference value for scoring the junction points.
[0131] We established a pool of plausible pairwise junctions between helical segments in the SCUBA-optimized configuration by retaining all pairwise junctions that resulted in a local density greater than 50% of that of the native loop. We then enumerated all possible N-terminal to C-terminal segment ordering schemes, ensuring that as many possible combinations could be sequentially connected via the retained junction pairs. For each ordering and connection scheme, we constructed closed loops of 3 to 4 residues and optimized loop backbones using SCUBA, obtaining single-chain mainchain configurations. The mainchain configurations were further relaxed using 50 ps of SCUBA SD simulated annealing at temperatures varying between 1.0 and 0.5, with all helical residues consisting of leucine and loop residues devoid of side chains. ABACUS2 sequence design was then applied to the backbone, followed by two iterations of SCUBA backbone relaxation, and the ABACUS2 sequence was reselected. The final design was the sequence with the lowest ABACUS2 energy.
[0132] In the above process, the screening criteria at each stage were as follows: the connected single chain should include at least five (truncated) helices, all connected helices must form a single compact structure, the estimated DaliZ-score must be lower than 6.0 to ensure structural novelty, the average per-residue ABACUS2 energy must be lower than -0.5, and the Rosetta forward folding results should be consistent with the designed model (RMSD lower than ).
[0133] Finally, we selected 13 design results for experimental identification based on a comprehensive consideration of structural novelty, ABACUS2 sequence energy, and Rosetta forward folding results.
[0134] 4. Experimental Methods
[0135] Preparation of recombinant proteins
[0136] DNA sequences encoding the designed proteins were synthesized by Qingke Biotechnology and General Biotechnology and cloned into the NdeI and Xhol sites of pET-22b(+). D12, D22, and D53 were subcloned into the MBP expression vector V28E2 or V28E4, with slight modifications to the N-terminal residues for crystallization. The plasmids were transformed into Escherichia coli BL21(DE3) and the OD 600 When the protein level was about 0.8, 0.5 mM IPTG was used for induction at 16°C for 20 hours. 1 H- 15 N HSQC NMR was performed by using an inorganic medium (24 g / L KH2PO4, 5 g / L NaOH, 0.5 g / L 15 E. coli was cultured in NH4Cl, 2.2 mM MgSO4, 0.1 mM CaCl2 and 2.5 g / L glucose to prepare uniform 15 N-labeled proteins were used 15 NH4Cl was used as an isotope source, and cells were collected and sonicated in a buffer containing 20 mM Tris and 500 mM NaCl (pH 7.8). The expression level and solubility of the designed proteins were evaluated by SDS-PAGE. The soluble supernatant was purified by Ni 2+ Affinity chromatography was performed using 500 mM NaCl, 20 mM Tris-HCl, and 350 mM imidazole. The eluted protein was concentrated in a buffer containing 20 mM Tris, 300 mM NaCl, and 1 mM EDTA, pH 7.8, and purified on a Superdex 75 column. Gel filtration was performed using a purification system (GE Healthcare), and monomer fractions were collected for structural characterization.
[0137] 4.2 NMR data acquisition
[0138] All NMR data were collected on a Bruker DMRX500 spectrometer equipped with a triple resonance, self-shielded z-axis gradient probe at 298 K. NMR samples typically contained 0.35 mM 15The N-labeled protein was treated with 20 mM NaH2PO4, 50 mM NaCl, 1 mM EDTA (pH 6.2) and 10% (v / v) D2O. The NMR data were processed using NMRDraw / NMRPipe and SPARKY.
[0139] Crystallization and X-ray diffraction analysis
[0140] Crystallization screening was performed at 289 K using the sitting drop vapor diffusion method and commercial screening reagents (Hampton Research). 10-20 mg / ml of purified protein was added to the corresponding crystallization buffer. The crystals were briefly soaked in the reservoir solution supplemented with 15% glycerol (v / v) and then flash frozen in liquid nitrogen. Diffraction data were acquired on a Rigaku XtaLABPRO 007HF. The collection and processing were performed using the CrysAlisPro software suite (CrysAlisPro Software System, Version 1.171.39.35c, Rigaku). The designed structures were used as search models for molecular replacement using PHENIX. Model reconstruction was performed in Coot, and the final structure was refined by PHENIX.
[0141] Circular dichroism
[0142] Circular dichroism (CD) data were obtained from an Applied Photophysics Chirascan TM Fluorescence was collected using a 1 mm pathlength quartz cuvette on a V100 spectrophotometer. Protein samples were prepared in 20 mM Na₂HPO₄ buffer (pH 8.0), and the protein concentration was adjusted to 0.6-0.8 mg / ml prior to measurement. CD spectra were scanned in the far-UV range (200-260 nm), and thermal denaturation CD data were obtained by measuring ellipticity at λ = 218 nm from 20°C to 95°C at 5°C increments. CD spectral data were processed, and secondary structure content was estimated using the built-in software suites Pro-Data Viewer v4.5 and Deconvolution v2.1, respectively.
[0143] After the above steps, the present invention obtained the following successful artificially designed protein and its amino acid sequence:
[0144] 1) For the partial modification and design of natural proteins, an artificially designed protein numbered EXTD-3 was obtained. This protein uses a natural protein (PDB ID 5wd2, an all-α helical protein) as the structural basis. By extending its H3 and H6 helices and artificially constructing another all-helical supersecondary structure, the structure was optimized using SCUBA and the amino acid sequence was designed using ABACUS2. Finally, the artificially designed protein was successfully expressed and the crystal structure was solved. The structural deviation between the crystal structure of the designed protein and the design target is small (the root mean square deviation of the main chain atomic position is )(See Figure 4 The amino acid sequence of EXTD-3 protein is shown in Table 2.
[0145] 2) Two successful artificially designed proteins, numbered XM2H and AM2M, were obtained from the de novo design of α+β proteins. Their crystal structures were analyzed by X-ray crystallography, and the structural deviations from the design targets were very small (the root mean square deviations of the main chain atomic positions were and )(See Figure 5 ). In addition, circular dichroism data showed that the thermal denaturation temperature of XM2H was above 95°C, indicating that this de novo designed protein has high thermal stability. The amino acid sequences of these two artificially designed proteins are shown in Table 2.
[0146] 3) All-α protein was designed from scratch to obtain three successful artificially designed proteins numbered H4A1R, H4A2S and H4C2R. Their crystal structures were analyzed by X-ray crystallography, and the structural deviations from the design targets were very small (the root mean square deviations of the main chain atomic positions were as well as )(See Figure 6 Circular dichroism data also showed that the thermal denaturation temperature of H4A1R was above 95°C, indicating that this de novo designed protein possesses high thermal stability. The amino acid sequences of these three artificially designed proteins are shown in Table 2.
[0147] 4) Three successful artificially designed proteins, numbered D12, D22, and D53, were designed from scratch by randomly placing helical proteins. MBP protein tags were used to assist protein crystallization, and their crystal structures were then analyzed by X-ray crystallography. The structural deviations from the design targets were small (the root mean square deviations of the main chain atomic positions were 1.0 and 2.0, respectively). as well as )(See Figure 7 ), where the rigid attachment of the MBP tag to the N-terminus of the designed protein peptide chain increases structural deviation. The amino acid sequences of these three artificially designed proteins are shown in Table 2.
[0148] Table 2. Amino acid sequences of artificially designed proteins
[0149]
[0150]
[0151] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for de novo protein design, characterized in that: The method comprises: a. Produce a highly designable main chain structure; b. Iterative amino acid sequence selection and backbone modification; Among them, the generation of a highly designable main chain structure includes: using the main chain-focused statistical energy function SCUBA as the objective function, optimizing the main chain structure with a computational method under the condition that the sequence is undetermined, and finding a stable main chain structure with minimal total SCUBA energy. The optimization of the main chain structure using computational methods includes sub-stage 1 and sub-stage 2. In the sub-stage 1, the local conformation of a given helix or β-strand segment is constrained by the secondary structure type, and the spatial interactions involving the atoms in the loop region are multiplied by 0.01; In sub-stage 2, local conformational secondary structure type restraints were removed and interactions involving loop residues were restored to full strength.
2. The method according to claim 1, characterized in that The energy term in SCUBA is a continuous analytical function obtained by training non-redundant natural protein structure data using the neighbor counting-neural network method. Using a stochastic kinetic model and a simulated annealing algorithm, the optimized main chain represented by the minimum point on the SCUBA energy surface is found after calculation, starting from the initial main chain structure.
3. The method according to claim 2, characterized in that The search range for conformations can be changed by controlling the parameters of the optimization algorithm.
4. The method according to claim 3, characterized in that The parameters of the optimization algorithm are the temperature variation range of simulated annealing or the simulation time length.
5. The method according to claim 3, characterized in that During the optimization process, user-defined constraint energy terms are added to make the generated main chain structure meet specific user requirements.
6. The method according to claim 5, characterized in that Local resampling and optimization are performed. For the ring regions in the main chain structure obtained by continuous simulation optimization, a large number of closed ring region structures are regenerated using computational methods, the SCUBA energy is optimized, and the optimized main chain with low energy is selected.
7. The method according to claim 6, characterized in that The model for optimizing the main chain structure SCUBA energy includes residues with simple side chain structures, which play a role in side chain space occupation.
8. The method according to claim 7, characterized in that The residues with simple side chain structures are leucine or valine.
9. The method according to claim 7, characterized in that The generated main chain structure was used to design the amino acid sequence, and the ABACUS method was used to design the amino acid sequence.
10. The method according to claim 9, characterized in that After replacing the side chains with the designed amino acid sequence, the SCUBA energy of the overall structure is optimized to produce a main chain structure modified according to the designed amino acid sequence.
11. The method according to claim 10, characterized in that The amino acid sequence was redesigned using the corrected main chain structure, and after multiple rounds of amino acid sequence design-main chain structure correction iterations, the optimal amino acid sequence was screened.
12. The method according to claim 11, characterized in that The optimal amino acid sequence was selected based on the lowest energy of the ABACUS sequence.