Generation system and method for improving designability of protein skeleton
By performing directed evolution and optimization of protein backbone imagery in the latent space, combined with geometric-physical dual constraints and sequence adaptability assessment, the stability and adaptability issues of protein backbone generation in existing technologies are solved, thereby improving the design efficiency and success rate of drug-eluting corneal contact lenses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-03-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to balance high activity, suitable release kinetics, and good biocompatibility when designing protein backbones, limiting the application of drug-loaded corneal contact lenses. Furthermore, the backbones generated by the models often lack global designability and sequence adaptability, resulting in low experimental validation success rates.
By achieving directed evolution and optimization of protein skeleton configurations in the latent space, and employing a geometric-physical dual constraint, contact feedback closed-loop optimization, and sequence adaptation synergistic mechanism, the stability of skeleton folding and sequence compatibility are improved. This includes the synergistic effect of modules such as functional constraints and structural prior information acquisition, geometric topological constraints, latent variable modeling, physical energy function regularization, residue contact map analysis, conformational correction, and sequence adaptability assessment.
It significantly improves the stability and sequence fit of the protein backbone, reduces the cost of experimental trial and error, increases the efficiency and success rate of de novo protein design, and ensures that the generated backbone meets the conformational rules of biomacromolecules with atomic-level precision.
Smart Images

Figure CN121768484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedicine, and in particular to a generation system and method for enhancing the designability of protein backbones. Background Technology
[0002] Drug-loaded contact lenses (DLCLs), as an emerging ocular drug delivery platform, integrate drug molecules into the lens material, enabling continuous and controlled drug release from the corneal surface. This significantly improves drug bioavailability and reduces the inconvenience and fluctuating efficacy associated with frequent administration of traditional eye drops. However, the clinical efficacy of this technology is highly dependent on the physicochemical properties, release characteristics, and biocompatibility of the drug carried. Currently, ideal drug molecules for DLCLs often require specific molecular size, stability, hydrophilicity / hydrophobicity, and affinity for the lens material. These properties are often difficult to achieve simultaneously in natural or existing synthetic drug molecules, limiting the widespread application of DLCLs in treating various ocular diseases such as glaucoma, keratitis, dry eye syndrome, and postoperative infection control.
[0003] Therefore, the efficient design and screening of drug molecules that possess high activity, suitable release kinetics, and good biocompatibility has become one of the core challenges driving the development of drug-eluting contact lens technology. Protein drugs (such as antibody fragments, peptides, and enzyme inhibitors) are considered ideal payloads for DLCLs due to their high specificity, adjustable pharmacokinetics, and potential ocular tissue targeting capabilities.
[0004] With the deep integration of computational biology and artificial intelligence, the field of protein design is undergoing a paradigm shift from static structure prediction to dynamic function generation. Current mainstream methods rely on physical force fields or deep learning models to simulate the reverse folding of amino acid sequences. Their core assumption is that the backbone structure is inherently designable, meaning there exists at least one sequence path that can stably fold into that conformation. However, most randomly sampled protein backbones in nature lack evolutionary constraints and often exhibit "energy desert" characteristics in sequence space: their free energy surfaces are highly rugged and contain numerous local minima, making traditional sampling algorithms prone to falling into non-physiological conformations and struggling to find thermodynamically stable folding sequences. This structural defect is particularly prominent in artificial design scenarios. When the target function requires overcoming the topological constraints of natural proteins, backbones generated by existing methods often lose folding feasibility due to side chain stacking conflicts, broken hydrogen bond networks, or exposed hydrophobic cores.
[0005] While Monte Carlo or molecular dynamics-based skeleton optimization strategies can locally adjust dihedral angle parameters, they lack a global designability guidance mechanism, resulting in sampling trajectories that often wander outside the foldable conformation subspace. More critically, existing evaluation systems overly rely on the final folding energy as the sole criterion, neglecting the entropy increase effect of sequence-structure mapping. Even if a skeleton exhibits a low-energy state under a specific sequence, if its fault tolerance range is narrow (i.e., only a very small number of sequences can achieve stable folding), the actual engineering success rate still approaches zero. This lack of evaluation dimensions causes algorithms to consume significant computational resources on seemingly reasonable skeletons, ultimately producing virtual structures that cannot be experimentally verified.
[0006] Current technologies therefore face a triple contradiction: First, the energy function is not coupled with a designability surrogate index, leading to a disconnect between the optimization direction and biophysical realities; second, the scaffold sampling process lacks a backtracking mechanism, making it impossible to extract conformational constraint rules from failed cases to correct the search path; third, the sequence design tools and structure generation modules have a unidirectional pipeline relationship, failing to form a closed-loop iteration of "generation-evaluation-feedback". These problems are amplified dramatically in applications requiring precise control of active site geometry or the construction of ultra-stable scaffold proteins, causing the success rate of experimental validation of artificially designed proteins to remain in the single digits for a long time, severely restricting the industrialization process in synthetic biology and enzyme engineering. Therefore, it is urgent to establish a generative paradigm that can actively shape scaffold designability, fundamentally breaking through the bottleneck of sequence-structure co-evolution. Summary of the Invention
[0007] The core of this invention lies in achieving directed evolution and optimization of protein skeleton configuration in the latent space, ultimately outputting a skeleton structure with high sequence adaptability. This addresses the core defects of existing technologies, such as difficulty in generating protein skeletons that match stable folded sequences, loose conformations, and low adaptability. At the same time, through geometric-physical dual constraints, contact feedback closed-loop optimization, and sequence adaptability synergistic mechanisms, the invention significantly improves skeleton folding stability and sequence compatibility, reduces experimental trial-and-error costs, and increases the efficiency of de novo protein design.
[0008] To solve the above problems, the present invention adopts the following technical solution.
[0009] The protein backbone designability enhancement generation system includes a functional constraint and structural prior information acquisition module, an initial backbone generation space construction module, a geometric topological constraint application module, a latent variable modeling module, a physical energy function regularization module, a residue contact map analysis module, a conformation correction module, a sequence fitness assessment module, a sequence-guided reconstruction module, and a designability score output module. The functional constraint and structural prior information acquisition module is used to acquire the functional constraints and structural prior information of the target protein. The functional constraints include the spatial coordinates of the active site and the chemical type of the catalytic residues, and the structural prior information includes the distribution of secondary structure types and the topological connection sequence. The initial backbone generation space construction module constructs the initial backbone generation space based on functional constraints and prior structural information. The initial backbone generation space is composed of the main chain atomic coordinate sequence, and its dimension is consistent with the number of amino acid residues of the target protein. The geometric topology constraint application module is used to apply geometric topology constraints to the initial skeleton generation space; The latent variable modeling module is used to perform latent variable modeling on the skeleton image space under geometric topological constraints through a variational autoencoder. The physical energy function regularization module is used to introduce the physical energy function as a regularization term during the training process of the variational autoencoder. The residue contact mapping analysis module is used to perform residue contact mapping analysis on the skeleton images output by the variational autoencoder. The conformation correction module is activated when the contact density of the residue contact pattern is lower than a preset threshold of 0.3. It adjusts the skeleton coordinates through gradient descent to increase the contact density to above 0.3 while maintaining the integrity of the hydrogen bond network of the secondary structure units. The sequence fitness assessment module is executed after the conformation correction module outputs a stable backbone. Its tasks include calculating the side chain rotator energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The sequence-guided reconstruction module is activated when the sequence fitness evaluation score is lower than a preset threshold of 0.7, triggering latent space resampling and modeling operations. The designability score output module is used to output the final skeleton image and its corresponding designability score.
[0010] A method for generating protein backbones with enhanced designability includes the following steps: S1. Obtain the functional constraints and structural prior information of the target protein, and then construct the initial backbone generation space based on the functional constraints and structural prior information; S2. Apply geometric topological constraints to the initial skeleton generation space. The geometric topological constraints include fixed bond length constraints, bond angle range constraints, and dihedral rotation degree of freedom constraints between adjacent residues. The fixed bond length constraints are set to 0.152 nm for carbon-nitrogen bonds, 0.147 nm for nitrogen-carbon alpha bonds, and 0.152 nm for carbon alpha-carbon bonds. The bond angle range constraints are set to 110 to 120 degrees for nitrogen-carbon alpha-carbon bonds and 115 to 125 degrees for carbon-nitrogen-carbon alpha bonds. The dihedral rotation degree of freedom constraints are limited according to the allowed region of the Laplace diagram. S3. Under geometric topological constraints, latent variables are modeled in the skeleton image space by variational autoencoder. The encoder part of the variational autoencoder maps the skeleton coordinates to low-dimensional latent vectors, and the decoder part of the variational autoencoder reconstructs the latent vectors into three-dimensional skeleton coordinates. The dimension of the latent vectors is set to 128. S4. During the training process of the variational autoencoder in step S3, a physical energy function is introduced as a regularization term. The physical energy function includes a van der Waals repulsion term, an electrostatic interaction term, and a solvent accessible surface area term. The van der Waals repulsion term is calculated using the twelfth power inverse proportional potential function. The electrostatic interaction term is calculated using Coulomb's law and multiplied by the dielectric constant of 8.0. The solvent accessible surface area term is estimated using a fast surface calculation algorithm and multiplied by the surface tension coefficient of 0.02 kJ per square nanometer per mole. S5. Perform residue contact map analysis on the skeleton image output by the variational autoencoder. The residue contact map is defined as a contact when the carbon beta atomic spacing between any two non-adjacent residues is less than 0.8 nm. The contact map is used to evaluate the global folding compactness of the skeleton. If the contact density of the residue contact map is lower than the preset threshold of 0.3, the conformation correction module is activated. The conformation correction module adjusts the skeleton coordinates through the gradient descent method to increase the contact density to above 0.3, while maintaining the integrity of the hydrogen bond network of the secondary structure unit. The integrity of the hydrogen bond network is determined by the spacing between the carbonyl oxygen of the main chain and the amide hydrogen being less than 0.35 nm and the included angle being greater than 120 degrees. S6. After the conformation correction module outputs a stable backbone, the sequence fitness assessment step is executed. The sequence fitness assessment step includes calculating the side chain rotomer energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The side chain rotomer energy is obtained from the standard rotomer database using a lookup table method. The stereo collision penalty is calculated using atomic radius overlap integral. The aggregation degree of hydrophobic residues is defined as the proportion of hydrophobic residues in a 3-nanometer radius sphere at the center of the protein volume. If the sequence fitness assessment score is lower than the preset threshold of 0.7, the sequence-guided reconstruction mechanism is activated. The fitness assessment score is fed back as a weight to the latent space of the variational autoencoder, and a new backbone conformation is resampled and generated until the fitness assessment score is higher than 0.7. S7. Output the final skeletal image and its corresponding designability score. The designability score is obtained by weighted summation of the contact density score, hydrogen bond network integrity score and sequence fit assessment score, with weight coefficients of 0.4, 0.3 and 0.3, respectively.
[0011] Furthermore, step S1, which constructs the initial backbone generation space, specifically includes the following steps: Based on the number of amino acid residues in the target protein, a three-dimensional coordinate matrix is initialized, with the number of rows equal to the number of residues and the number of columns equal to three; random coordinate values are assigned to the first row of the three-dimensional coordinate matrix as the spatial position of the starting residue; for each subsequent row, the theoretical coordinate position of the current residue is calculated through spherical coordinate transformation based on the coordinate values of the previous row and preset bond length and bond angle parameters; Gaussian noise perturbation is superimposed on the theoretical coordinate position, with the standard deviation of the Gaussian noise set to 0.05 nanometers to introduce conformational diversity; a rigid body transformation is performed on the coordinate matrix after noise superposition, so that its centroid is located at the origin and its principal axes are aligned with the coordinate axes, thus completing the construction of the initial backbone generation space.
[0012] Furthermore, the encoder part of the variational autoencoder in step S3 is used as follows: the input skeleton coordinate matrix is flattened into a one-dimensional vector; the one-dimensional vector is input into the first fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the first fully connected layer is input into the second fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the second fully connected layer is input into the mean prediction branch and the variance prediction branch, respectively. The mean prediction branch is a fully connected layer with an output dimension of 128, and the variance prediction branch is a fully connected layer with an output dimension of 128. Its output is transformed by an exponential function to ensure a positive value; the latent vector is sampled from the Gaussian distribution defined by the mean and variance parameters.
[0013] Furthermore, the decoder part of the variational autoencoder in step S3 is used as follows: the sampled latent vector is input into the third fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the third fully connected layer is input into the fourth fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the fourth fully connected layer is input into the fifth fully connected layer, which has the number of neurons equal to the dimension of the flattened skeleton coordinate matrix and uses an identity function as the activation function; the output of the fifth fully connected layer is reshaped into the original three-dimensional coordinate matrix structure to obtain the reconstructed skeleton coordinates.
[0014] Furthermore, the specific usage method of the conformation correction module in step S5 includes the following: calculating the carbon beta atomic spacing of all non-adjacent residue pairs in the current skeleton configuration; counting the number of residue pairs with spacing less than 0.8 nanometers, dividing by the theoretical maximum number of contact pairs to obtain the contact density; if the contact density is less than 0.3, constructing a loss function, which consists of a negative contact density, a hydrogen bond network disruption penalty term, and a coordinate change amplitude penalty term; calculating the gradient of the loss function for each atom coordinate in the skeleton coordinate matrix; updating the coordinates in the opposite direction of the gradient, with a step size set to 0.001 nanometers; repeating the gradient update process until the contact density is greater than or equal to 0.3 or the number of iterations reaches one thousand; after each coordinate update, forcibly executing bond length and bond angle constraints, and projecting the coordinates back to the constrained manifold using the Lagrange multiplier method.
[0015] Furthermore, the sequence fitness assessment step in step S6 specifically includes the following: for each residue position in the skeletal structure, enumerate all twenty standard amino acid side chain conformations; for each side chain conformation, calculate its van der Waals repulsion energy with the adjacent main chain and side chain atoms; select the side chain conformation with the smallest repulsion energy as the optimal side chain at that position; calculate the total repulsion energy of all optimal side chains and divide it by the number of residues to obtain the average repulsion energy; calculate the stereo collision integral of all atomic pairs in the skeletal structure; if the sum of the van der Waals radii of any atomic pair is less than the actual spacing, then accumulate the penalty value; count the number of hydrophobic residues located in the protein core region, and the protein... The core region is a spatial sphere less than three nanometers from the centroid of the backbone. Hydrophobic residues include alanine, valine, leucine, isoleucine, phenylalanine, tryptophan, tyrosine, and methionine. The hydrophobic aggregation degree is obtained by dividing the number of hydrophobic residues by the total number of residues in the core region. The average repulsion energy is normalized to the zero-to-one interval and inverted as the repulsion energy score. The stereo collision penalty value is normalized to the zero-to-one interval and inverted as the collision avoidance score. The hydrophobic aggregation degree is used as the hydrophobic score. The repulsion energy score, collision avoidance score, and hydrophobic score are weighted and averaged with weights of 0.5, 0.3, and 0.2, respectively, to obtain the sequence fitness evaluation score.
[0016] Furthermore, the specific usage method of the sequence-guided reconstruction mechanism in step S6 includes: mapping the current sequence fitness assessment score to a temperature parameter of the latent space sampling, wherein the temperature parameter is negatively correlated with the fitness assessment score, and the calculation formula is that the temperature parameter equals one minus the fitness assessment score; in the latent space of the variational autoencoder, the standard deviation of the sampling distribution is adjusted according to the temperature parameter, wherein the standard deviation equals the base standard deviation multiplied by the temperature parameter, and the base standard deviation is set to 0.1; the latent vector is resampled from the adjusted Gaussian distribution; the newly sampled latent vector is input into the decoder to generate a new skeleton image; the sequence fitness assessment step is re-executed on the new skeleton image; if the new score is still lower than 0.7, the sampling and assessment process is repeated, up to five times; if the target is not met within five times, the skeleton image with the highest score among the five times is selected as the output.
[0017] Furthermore, the calculation of the designability score in step S7 specifically includes the following steps: multiply the contact density by 100, then by 0.4 to obtain the contact score; count the number of main chain atom pairs in the backbone that satisfy the definition of hydrogen bonds, divide by the theoretical maximum number of hydrogen bonds, then multiply by 100, then by 0.3 to obtain the hydrogen bond score; multiply the sequence fit assessment score by 100, then by 0.3 to obtain the fit score; add the contact score, hydrogen bond score and fit score to obtain the final designability score.
[0018] Furthermore, the physical energy function in step S4 also includes a hydrogen bond formation tendency term and a structure preservation term. The hydrogen bond formation tendency term determines the probability of hydrogen bonds based on the distance and angle between the carbonyl group and the amino group in the main chain, while the structure preservation term applies an additional energy penalty to the secondary structure region to prevent breakage.
[0019] Furthermore, during the variational autoencoder training process in step S3, an adversarial generation mechanism is adopted to distinguish between real protein conformations and generated conformations through a discriminator, forcing the latent space distribution to approximate the real conformation manifold.
[0020] Furthermore, during the use of the conformation correction module, the gradient of the loss function is backpropagated to the latent space, driving the adjustment of latent variables, causing the conformation to evolve in a direction that is more easily filled by the sequence. At the same time, constraint projection is introduced to ensure that the conformation generated after the latent variable adjustment still satisfies the geometric topological constraints and the physical energy function guidance conditions.
[0021] Furthermore, in the sequence fitness assessment step, a thousand random amino acid sequences were generated using the Monte Carlo sampling method, and the conformational energy of each sequence under the framework was calculated using a fast folding simulator.
[0022] Furthermore, during the use of the sequence-guided reconstruction mechanism, when all new point scores are below the threshold, the sampling radius is expanded to ten, and the sampling process is repeated until a score that meets the standard is found or the maximum number of samplings is reached.
[0023] Furthermore, the designability scoring operation also includes geometric rationality scoring and contact pattern quality scoring. Geometric rationality scoring is calculated by checking the main chain bond length and bond angle deviation, the rationality of dihedral angle distribution, and residue collision. Contact pattern quality scoring is obtained by calculating the proportion of hydrophobic residues in the contact pair, the rationality of polar residue pairing, and long-range contact density.
[0024] Compared with the prior art, the advantages of this invention are: This approach employs a dual-guidance mechanism of geometric topological constraints and physical energy functions to ensure that the generated protein backbone meets the basic conformational rules of biomacromolecules at atomic level, avoiding structural collapse and energy distortion caused by unconstrained generation. Through dynamic feedback from residue contact maps and closed-loop optimization of the conformational correction module, the backbone is forced to form a compact global folding morphology, solving the problem of existing generation models outputting overly loose backbones lacking stable cores. A synergistic mechanism of sequence fit assessment and latent space resampling directly embeds sequence design feasibility into the backbone generation process, enabling the output backbone to naturally fit highly stable amino acid sequences, overcoming the bottleneck of low fit caused by the separation of backbone generation and sequence design in traditional methods. Quantitative output of designability scores provides clear priorities for downstream experimental screening, significantly reducing wet experimental trial-and-error costs and improving the overall efficiency and success rate of de novo protein design. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall technical solution architecture of the protein backbone designability enhancement generation method proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the dual guidance mechanism of geometric topological constraints and physical energy functions in this invention; Figure 3 This is a flowchart illustrating the logical process of initial skeleton generation space construction and latent variable modeling in this invention. Figure 4 This is a schematic diagram of the closed-loop optimization framework of the residue contact map dynamic feedback and conformation correction module in this invention. Figure 5 This is a logical flowchart of the sequence fitness assessment and latent space resampling collaborative mechanism in this invention; Figure 6 This is a schematic diagram illustrating the principle framework of the quantitative output of designability scoring and multi-dimensional weighted fusion in this invention. Detailed Implementation
[0026] The technical solutions will now be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.
[0027] Implementation method:
[0028] Please see Figures 1-6 The protein backbone designability enhancement generation system includes a functional constraint and structural prior information acquisition module, an initial backbone generation space construction module, a geometric topological constraint application module, a latent variable modeling module, a physical energy function regularization module, a residue contact map analysis module, a conformation correction module, a sequence fitness assessment module, a sequence-guided reconstruction module, and a designability score output module.
[0029] The functional constraint and structural prior information acquisition module is used to acquire the functional constraints and structural prior information of the target protein. The functional constraints include the spatial coordinates of the active site and the chemical type of the catalytic residues, and the structural prior information includes the distribution of secondary structure types and the topological connection sequence. The initial backbone generation space construction module constructs the initial backbone generation space based on functional constraints and prior structural information. The initial backbone generation space is composed of the main chain atomic coordinate sequence, and its dimension is consistent with the number of amino acid residues of the target protein. The geometric topology constraint application module is used to apply geometric topology constraints to the initial skeleton generation space; The latent variable modeling module is used to perform latent variable modeling on the skeleton image space under geometric topological constraints through a variational autoencoder. The physical energy function regularization module is used to introduce the physical energy function as a regularization term during the training process of the variational autoencoder. The residue contact mapping analysis module is used to perform residue contact mapping analysis on the skeleton images output by the variational autoencoder. The conformation correction module is activated when the contact density of the residue contact pattern is lower than a preset threshold of 0.3. It adjusts the skeleton coordinates through gradient descent to increase the contact density to above 0.3 while maintaining the integrity of the hydrogen bond network of the secondary structure units. The sequence fitness assessment module is executed after the conformation correction module outputs a stable backbone. Its tasks include calculating the side chain rotator energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The sequence-guided reconstruction module is activated when the sequence fitness evaluation score is lower than a preset threshold of 0.7, triggering latent space resampling and modeling operations. The designability score output module is used to output the final skeleton image and its corresponding designability score.
[0030] A method for enhancing the designability of protein backbones, applied to the aforementioned protein backbone designability enhancement generation system, includes the following steps: S1. Obtain the functional constraints and structural prior information of the target protein, and then construct the initial backbone generation space based on the functional constraints and structural prior information; The construction of the initial backbone generation space includes the following steps: First, a three-dimensional coordinate matrix is initialized based on the number of amino acid residues in the target protein. The number of rows in the three-dimensional coordinate matrix equals the number of residues, and the number of columns equals three. Random coordinate values are assigned to the first row of the three-dimensional coordinate matrix as the spatial position of the starting residue. For each subsequent row, the theoretical coordinate position of the current residue is calculated using spherical coordinate transformation based on the coordinate values of the previous row and preset bond length and bond angle parameters. Gaussian noise perturbation is superimposed on the theoretical coordinate position, with the standard deviation of the Gaussian noise set to 0.05 nanometers to introduce conformational diversity. A rigid body transformation is then performed on the coordinate matrix after noise superposition, so that its centroid is located at the origin and its principal axes are aligned with the coordinate axes, thus completing the construction of the initial backbone generation space.
[0031] The initial conformational space of the protein backbone needs to cover a sufficiently wide range of folding possibilities, while ensuring that the conformations are physically reasonable and topologically coherent. To this end, this method first extracts representative local secondary structure fragments from a known structure database, including α-helices, β-sheets, turns, and random coils, to construct a basic conformational unit library. Each conformational unit contains main chain atomic coordinates, dihedral angle distribution statistical characteristics, and local hydrogen bond patterns. Subsequently, through random splicing and spatial transformation, multiple conformational units are combined into a complete backbone, and smooth interpolation is performed at the splicing points to ensure main chain continuity and reasonable bond angles. After initial conformational sampling, it is mapped to a 128-dimensional latent space to establish a latent variable model. The latent variable dimension is set to 128 dimensions, with each dimension corresponding to a certain type of structural degree of freedom in the conformational space, such as helical length, fold angle, and loop flexibility. The latent variables are initialized using uniformly distributed sampling to ensure that the initial conformations are uniformly covered in the latent space.
[0032] S2. Apply geometric topological constraints to the initial skeleton generation space. The geometric topological constraints include fixed bond length constraints, bond angle range constraints, and dihedral rotation degree of freedom constraints between adjacent residues. The fixed bond length constraints are set to 0.152 nm for carbon-nitrogen bonds, 0.147 nm for nitrogen-carbon alpha bonds, and 0.152 nm for carbon alpha-carbon bonds. The bond angle range constraints are set to 110 to 120 degrees for nitrogen-carbon alpha-carbon bonds and 115 to 125 degrees for carbon-nitrogen-carbon alpha bonds. The dihedral rotation degree of freedom constraints are limited according to the allowed region of the Laplace diagram.
[0033] S3. Under geometric topological constraints, the skeleton image space is modeled with latent variables by a variational autoencoder. The encoder part of the variational autoencoder maps the skeleton coordinates to low-dimensional latent vectors, and the decoder part of the variational autoencoder reconstructs the latent vectors into three-dimensional skeleton coordinates. The encoder part of the variational autoencoder operates as follows: The input skeleton coordinate matrix is flattened into a one-dimensional vector; this one-dimensional vector is input into the first fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the first fully connected layer is input into the second fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the second fully connected layer is input into the mean prediction branch and the variance prediction branch, respectively. The mean prediction branch is a fully connected layer with an output dimension of 128, and the variance prediction branch is also a fully connected layer with an output dimension of 128. The outputs are transformed using an exponential function to ensure positive values; latent vectors are sampled from a Gaussian distribution defined by the mean and variance parameters. The decoder part of the variational autoencoder operates as follows: The sampled latent vector is input into the third fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the third fully connected layer is input into the fourth fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the fourth fully connected layer is input into the fifth fully connected layer, which has the number of neurons equal to the dimension of the flattened skeleton coordinate matrix and uses an identity function as the activation function; the output of the fifth fully connected layer is reshaped into the original three-dimensional coordinate matrix structure to obtain the reconstructed skeleton coordinates.
[0034] The mapping between latent variables and skeletal conformations is achieved through an encoder-decoder architecture. The encoder compresses the 3D skeleton coordinates into latent vectors, while the decoder restores the latent vectors back to skeleton coordinates. The encoder consists of a five-layer graph neural network, with each layer containing adjacency matrix propagation and node feature update modules to capture the spatial proximity relationships between residues. The decoder consists of a three-layer spatial coordinate generation network, recovering the main chain atom positions layer by layer. The latent variable model is trained using an adversarial generative mechanism, which uses a discriminator to distinguish between real and generated protein conformations, forcing the latent space distribution to approximate the real conformational manifold. After training, the latent variables can be used as control parameters for conformational evolution, allowing for continuous deformation and optimization of the conformation by adjusting the latent variable values.
[0035] S4. During the training process of the variational autoencoder in step S3, a physical energy function is introduced as a regularization term. The physical energy function includes a van der Waals repulsion term, an electrostatic interaction term, and a solvent accessible surface area term. The van der Waals repulsion term is calculated using the twelfth power inverse proportional potential function. The electrostatic interaction term is calculated using Coulomb's law and multiplied by the dielectric constant of 8.0. The solvent accessible surface area term is estimated using a fast surface calculation algorithm and multiplied by the surface tension coefficient of 0.02 kJ per square nanometer per mole. The physical energy function also includes a hydrogen bond formation tendency term and a structure preservation term. The hydrogen bond formation tendency term determines the probability of hydrogen bonds based on the distance and angle between the carbonyl group and the amino group in the main chain, while the structure preservation term applies an additional energy penalty to the secondary structure region to prevent breakage. The specific form of the physical energy function is as follows: E total =w1·E vdW +w2·E elec +w3·E sasa +w4·E hbond +w5·E struct ; Among them, E vdW The van der Waals repulsion term is used, and the interatomic repulsive force is calculated using the Lennard-Jones potential function; E elec For the electrostatic interaction term, the interaction between charged residues is calculated using the Coulomb potential; E sasa E represents the solvent-accessible surface area term, used to assess the hydrophobic effect by calculating the exposed area of residues. hbond The hydrogen bond formation tendency term is determined by the distance and angle between the carbonyl group and the amino group in the main chain; E struct As a structural preservation term, an additional energy penalty is applied to the secondary structural region to prevent fracture. The weighting coefficients w1, w2, w3, w4, and w5 are set to 0.4, 0.2, 0.2, 0.15, and 0.05, respectively, and experiments have verified that they can balance the contributions of each term.
[0036] In the latent space, conformational evolution must simultaneously satisfy geometric rationality and physical stability. Geometric topological constraints include main chain bond length and bond angle limitations, allowable dihedral angle ranges, minimum distance thresholds between residues, and secondary structure continuity requirements. Geometric constraints are achieved through hard truncation; that is, during conformational generation, if an atom's coordinates violate bond length or bond angle limitations, it is forcibly pulled back into the allowable range; if the distance between residues is less than the minimum threshold, a repulsive force is applied to separate them. The physical energy function serves as a soft guiding term, driving the conformation to move towards a lower energy region through gradient descent.
[0037] Specifically, after each latent variable update, a 3D skeleton is first generated by the decoder, then its physical energy value is calculated, and the energy gradient is backpropagated to the latent space to adjust the direction of the latent variables. The energy gradient calculation uses automatic differentiation technology; after taking the partial derivative with respect to the skeleton coordinates, it is passed to the latent variables through the chain rule. To avoid getting trapped in local minima, a simulated annealing mechanism is introduced, allowing higher energy fluctuations in the early stages of evolution, and gradually reducing the temperature parameter with increasing iterations, so that the conformation eventually converges to a low-energy stable region. In addition, to maintain the integrity of the secondary structure, a structure preservation term is added to the energy function, imposing additional constraints on the α-helix and β-fold regions to prevent them from breaking or twisting during evolution. Through a dual guidance mechanism composed of geometric topological constraints and physical energy functions, the generated conformation is ensured to conform to basic geometric rules and possess physical stability, laying the foundation for subsequent sequence fit evaluation.
[0038] S5. Perform residue contact map analysis on the skeleton image output by the variational autoencoder. The residue contact map is defined as a contact when the carbon beta atomic spacing between any two non-adjacent residues is less than 0.8 nm. The contact map is used to evaluate the global folding compactness of the skeleton. If the contact density of the residue contact map is lower than the preset threshold of 0.3, the conformation correction module is activated. The conformation correction module adjusts the skeleton coordinates through the gradient descent method to increase the contact density to above 0.3, while maintaining the integrity of the hydrogen bond network of the secondary structure unit. The integrity of the hydrogen bond network is determined by the spacing between the carbonyl oxygen of the main chain and the amide hydrogen being less than 0.35 nm and the included angle being greater than 120 degrees. The specific usage of the conformation correction module includes the following: Calculate the carbon beta atomic spacing of all non-adjacent residue pairs in the current skeleton configuration; count the number of residue pairs with spacing less than 0.8 nanometers, divide by the theoretical maximum number of contact pairs to obtain the contact density; if the contact density is less than 0.3, construct a loss function, which consists of a negative contact density, a hydrogen bond network disruption penalty term, and a coordinate change amplitude penalty term; calculate the gradient of the loss function for each atom coordinate in the skeleton coordinate matrix; update the coordinates in the opposite direction of the gradient, with a step size set to 0.001 nanometers; repeat the gradient update process until the contact density is greater than or equal to 0.3 or the number of iterations reaches 1,000; after each coordinate update, enforce bond length and bond angle constraints, and project the coordinates back to the constrained manifold using the Lagrange multiplier method.
[0039] Residue contact maps reflect the spatial proximity of amino acid residues within a protein and are a key factor in determining whether a sequence can fold stably. In this method, the contact map is directly calculated from the current skeletal conformation and is defined as a contact relationship between two residues when the Cα atomic distance between them is less than eight angstroms. The contact map is stored in matrix form, with matrix elements having values of zero or one, representing no contact or contact, respectively. To evaluate the sequence compatibility of the current conformation, the contact map is input into a pre-trained sequence adaptation prediction model. This model is trained on a large number of known structure-sequence pairs and can predict which amino acid combinations are more likely to form a stable fold based on contact patterns. The prediction results are output as a residue pair compatibility score matrix, with matrix element values between zero and one; higher values indicate that the residue pair is more likely to be stably filled by the sequence. Subsequently, the compatibility score matrix is multiplied element-wise with the current contact map to obtain a weighted contact loss function. The weighted contact loss function is defined as: ; Among them, C ij For contact map matrix elements, P ij The residue pair compatibility score output by the sequence compatibility prediction model is used as a loss function to encourage high-compatibility residue pairs to form contacts in the conformation, while suppressing the occurrence of low-compatibility contact pairs, thereby improving the sequence filling potential of the conformation. This loss function measures the potential risk of the current conformation at the sequence-filling level. A higher loss value indicates more contact pairs in the conformation that are difficult to stabilize by the sequence. To reduce the loss, the gradient of the loss function is backpropagated to the latent space, driving the adjustment of latent variables and causing the conformation to evolve in a direction that is more easily filled by the sequence. This process forms a closed-loop feedback and correction: conformation generation → contact map calculation → sequence compatibility prediction → loss calculation → latent variable update → new conformation generation. The above closed-loop correction process is executed once after each latent variable evolution, continuously optimizing the sequence-fitting potential of the conformation. To prevent the correction process from destroying the existing geometric and physical rationality, constraint projection is introduced in the gradient update to ensure that the conformation generated after the latent variable adjustment still satisfies the dual guiding conditions of geometric topological constraints and physical energy functions.
[0040] S6. After the conformation correction module outputs a stable backbone, the sequence fitness assessment step is executed. The sequence fitness assessment step includes calculating the side chain rotomer energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The side chain rotomer energy is obtained from the standard rotomer database using a lookup table method. The stereo collision penalty is calculated using atomic radius overlap integral. The aggregation degree of hydrophobic residues is defined as the proportion of hydrophobic residues in a 3-nanometer radius sphere at the center of the protein volume. If the sequence fitness assessment score is lower than the preset threshold of 0.7, the sequence-guided reconstruction mechanism is activated. The fitness assessment score is fed back as a weight to the latent space of the variational autoencoder, and a new backbone conformation is resampled and generated until the fitness assessment score is higher than 0.7.
[0041] Step S6, the sequence fitness assessment step, specifically includes the following: For each residue position in the skeletal structure, enumerate all twenty standard amino acid side chain conformations; for each side chain conformation, calculate its van der Waals repulsion energy with adjacent main chain and side chain atoms; select the side chain conformation with the lowest repulsion energy as the optimal side chain at that position; calculate the total repulsion energy of all optimal side chains and divide it by the number of residues to obtain the average repulsion energy; calculate the stereo collision integral of all atomic pairs in the skeletal structure; if the sum of the van der Waals radii of any atomic pair is less than the actual spacing, accumulate the penalty value; count the number of hydrophobic residues located in the protein core region, and the protein core... The region is a spatial sphere less than three nanometers from the centroid of the backbone. The hydrophobic residues include alanine, valine, leucine, isoleucine, phenylalanine, tryptophan, tyrosine, and methionine. The hydrophobic aggregation degree is obtained by dividing the number of hydrophobic residues by the total number of residues in the core region. The average repulsion energy is normalized to the zero-to-one interval and inverted as the repulsion energy score. The stereo collision penalty value is normalized to the zero-to-one interval and inverted as the collision avoidance score. The hydrophobic aggregation degree is used as the hydrophobic score. The repulsion energy score, collision avoidance score, and hydrophobic score are weighted and averaged with weights of 0.5, 0.3, and 0.2, respectively, to obtain the sequence fitness evaluation score.
[0042] The specific usage of the sequence-guided reconstruction mechanism in step S6 includes: mapping the current sequence fitness assessment score to a temperature parameter of the latent space sampling. The temperature parameter is negatively correlated with the fitness assessment score, and the calculation formula is that the temperature parameter equals one minus the fitness assessment score; in the latent space of the variational autoencoder, the standard deviation of the sampling distribution is adjusted according to the temperature parameter. The standard deviation is equal to the base standard deviation multiplied by the temperature parameter, and the base standard deviation is set to 0.1; the latent vector is resampled from the adjusted Gaussian distribution; the newly sampled latent vector is input into the decoder to generate a new skeleton image; the sequence fitness assessment step is re-executed on the new skeleton image; if the new score is still lower than 0.7, the sampling and assessment process is repeated, up to five times; if the target is not met within five times, the skeleton image with the highest score among the five times is selected as the output.
[0043] During conformational evolution, sequence fit needs to be periodically evaluated to determine whether the current latent variable region possesses high designability. Sequence fit assessment is achieved by fixing the skeleton conformation, attempting to fill various amino acid sequences, and calculating their folding free energies. Specifically, for the current skeleton, a Monte Carlo sampling method is used to generate one thousand random amino acid sequences, and the conformational energy of each sequence under that skeleton is calculated using a fast folding simulator. The folding simulator is based on a simplified physical model, considering only main chain constraints and side chain stacking effects, ignoring solvent and long-range electrostatics to improve computational efficiency. After calculation, the number of sequences with energies below a threshold is counted, and the ratio of this number to the total number of sequences is the sequence fit score of the current skeleton, calculated using the following formula: ; Where, N low N represents the number of sequences with energy below a threshold. total The threshold for the total number of sampled sequences is set to -50 kJ per mole, which corresponds to the typical free energy range of stable folded proteins. If the score is lower than the preset threshold of 0.3, it is determined that the current latent variable region is not designable enough, triggering the latent space resampling mechanism.
[0044] The resampling process first delineates a hyperspherical region with a radius of five, centered on the current latent variable, within the latent space. Then, ten new latent variable points are randomly generated within this region. A corresponding skeleton is generated for each new point, and sequence fitness is evaluated. The point with the highest score is selected as the new evolutionary starting point. If all new point scores are below a threshold, the sampling radius is expanded to ten, and the sampling process is repeated until a point with a satisfactory score is found or the maximum number of sampling iterations (twenty) is reached. If no suitable point is found after reaching the maximum number of iterations, the current evolutionary path is considered a failure, and the process reverts to the previous stable state and adjusts the evolutionary step size. This resampling mechanism ensures that the method does not stagnate in regions of low designability but actively explores better regions in the latent space, improving overall search efficiency.
[0045] S7. Output the final skeletal image and its corresponding designability score. The designability score is obtained by weighted summation of the contact density score, hydrogen bond network integrity score and sequence fit assessment score, with weight coefficients of 0.4, 0.3 and 0.3, respectively.
[0046] After completing the latent variable evolution and sequence fit optimization, the final conformation needs to be comprehensively scored to quantify its designability level. The scoring system includes four dimensions: geometric rationality score, physical stability score, sequence fit score, and contact map quality score. The geometric rationality score is calculated by checking the main chain bond length and bond angle deviation, the rationality of dihedral angle distribution, and residue collision, with a maximum score of one. The physical stability score is obtained by normalizing the energy function value, with lower energy resulting in a higher score. The sequence fit score is the sequence filling success rate. The contact map quality score is obtained by calculating the proportion of hydrophobic residues in the contact pairs, the rationality of polar residue pairing, and the long-range contact density.
[0047] The final designability score is calculated in two ways. The first method involves multiplying the contact density by 100 and then by 0.4 to obtain the contact score; counting the number of main-chain atom pairs satisfying the definition of hydrogen bonds in the framework, dividing by the theoretical maximum number of hydrogen bonds, multiplying by 100, and then by 0.3 to obtain the hydrogen bond score; multiplying the sequence fit assessment score by 100 and then by 0.3 to obtain the fit score; and adding the contact score, hydrogen bond score, and fit score to obtain the final designability score. The second method assigns weights of 0.25, 0.3, 0.35, and 0.1 to the four scores: geometric rationality score, physical stability score, sequence fit score, and contact pattern quality score. The weighted sum is then used to obtain the final designability score. This weighting is based on extensive experimental statistics. Sequence fit is given the highest weight because it directly determines whether the framework can be stably filled, followed by physical stability. Geometric rationality and contact pattern quality are used as auxiliary indicators, and any one of these methods can be selectively adopted in practice. Conformations with a final score higher than 0.7 are considered highly designable frameworks and are output. The output includes the 3D coordinates of the backbone, latent variable values, scores for each dimension, and a weighted total score. To facilitate subsequent sequence design, residue contact maps and compatibility prediction matrices are also output for reference by the sequence optimization algorithm.
[0048] This approach employs a dual-guided mechanism of geometric topological constraints and physical energy functions to dynamically evolve and assess sequence fit in the latent space of the protein backbone structure, ultimately outputting a highly designable protein backbone structure. The entire process begins with latent variable modeling, corrects the conformation through residue contact mapping feedback, and uses sequence fit scoring to drive latent space resampling, ultimately achieving a multi-dimensional weighted fusion of designable quantitative output.
[0049] Compared to traditional methods, this approach does not rely on fixed templates or empirical rules. Instead, it combines physical guidance with data-driven approaches to actively search for optimal designability regions in the conformational space, significantly improving the practicality and stability of the generated backbone. Experiments show that the backbone generated using this method achieves a success rate increase of over 40% in subsequent sequence design, and its folding stability is superior to existing generation models. This method is applicable to various scenarios such as novel protein design, enzyme active site reconstruction, and antibody backbone optimization, and has broad industrial and scientific research value.
[0050] The above description is merely a preferred embodiment of the present invention; it encompasses all the protection scope of the present invention. Any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solutions and improved concepts of the present invention, should be covered within the protection scope of the present invention.
Claims
1. A protein backbone designability enhancement generation system, characterized by: It includes modules for acquiring functional constraints and prior structural information, constructing the initial skeleton generation space, applying geometric and topological constraints, modeling latent variables, regularizing physical energy functions, analyzing residue contact maps, modifying conformations, evaluating sequence fitness, guiding sequence reconstruction, and outputting designability scores. The functional constraint and structural prior information acquisition module is used to acquire the functional constraints and structural prior information of the target protein. The functional constraints include the spatial coordinates of the active site and the chemical type of the catalytic residue. The structural prior information includes the distribution of secondary structure types and the topological connection order. The initial backbone generation space construction module constructs an initial backbone generation space based on functional constraints and prior structural information. The initial backbone generation space is composed of the main chain atomic coordinate sequence, and its dimension is consistent with the number of amino acid residues of the target protein. The geometric topology constraint application module is used to apply geometric topology constraints to the initial skeleton generation space; The latent variable modeling module is used to perform latent variable modeling on the skeleton image space through a variational autoencoder under geometric topological constraints. The physical energy function regularization module is used to introduce the physical energy function as a regularization term during the training process of the variational autoencoder. The residue contact mapping analysis module is used to perform residue contact mapping analysis on the skeletal image output by the variational autoencoder. The conformation correction module is activated when the contact density of the residue contact pattern is lower than a preset threshold of 0.
3. It adjusts the skeleton coordinates through gradient descent to increase the contact density to above 0.3 while maintaining the integrity of the hydrogen bond network of the secondary structure units. The sequence fitness assessment module is executed after the conformation correction module outputs a stable backbone. Its tasks include calculating the side chain rotator energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The sequence-guided reconstruction module is used to activate when the sequence fitness evaluation score is lower than a preset threshold of 0.7, triggering latent space resampling and modeling operations. The designability score output module is used to output the final skeleton image and its corresponding designability score.
2. A method for enhancing the designability of a protein backbone, applied to the protein backbone designability enhancement generation system described in claim 1, characterized in that: Includes the following steps: S1. Obtain the functional constraints and structural prior information of the target protein, and then construct the initial backbone generation space based on the functional constraints and structural prior information; S2. Apply geometric topological constraints to the initial skeleton generation space. The geometric topological constraints include fixed bond length constraints, bond angle range constraints, and dihedral rotational degree of freedom constraints between adjacent residues. The fixed bond length constraints are set to 0.152 nm for carbon-nitrogen bonds, 0.147 nm for nitrogen-carbon alpha bonds, and 0.152 nm for carbon alpha-carbon bonds. The bond angle range constraints are set to 110 to 120 degrees for nitrogen-carbon alpha-carbon bonds and 115 to 125 degrees for carbon-nitrogen-carbon alpha bonds. The dihedral rotational degree of freedom constraints are limited according to the allowed region of the Laplace diagram. S3. Under the geometric topological constraints, the skeleton image space is modeled with latent variables by a variational autoencoder. The encoder part of the variational autoencoder maps the skeleton coordinates to low-dimensional latent vectors, and the decoder part of the variational autoencoder reconstructs the latent vectors into three-dimensional skeleton coordinates. The dimension of the latent vectors is set to 128. S4. During the training process of the variational autoencoder in step S3, a physical energy function is introduced as a regularization term. The physical energy function includes a van der Waals repulsion term, an electrostatic interaction term, and a solvent accessible surface area term. The van der Waals repulsion term is calculated using the twelfth power inverse proportional potential function. The electrostatic interaction term is calculated using Coulomb's law and multiplied by the dielectric constant of 8.
0. The solvent accessible surface area term is estimated using a fast surface calculation algorithm and multiplied by the surface tension coefficient of 0.02 kJ per square nanometer per mole. S5. Perform residue contact mapping analysis on the skeleton image output by the variational autoencoder. The residue contact mapping is defined as a contact when the carbon beta atomic spacing between any two non-adjacent residues is less than 0.8 nanometers. The contact mapping is used to evaluate the global folding compactness of the skeleton. If the contact density of the residue contact mapping is lower than a preset threshold of 0.3, the conformation correction module is activated. The conformation correction module adjusts the skeleton coordinates through gradient descent to increase the contact density to above 0.3, while maintaining the integrity of the hydrogen bond network of the secondary structural units. The integrity of the hydrogen bond network is determined by the spacing between the carbonyl oxygen and amide hydrogen in the main chain being less than 0.35 nanometers and the included angle being greater than 120 degrees. S6. After the conformation correction module outputs a stable backbone, a sequence fitness assessment step is performed. The sequence fitness assessment step includes calculating the side chain rotomer energy at each residue position, assessing the stereo collision penalty between adjacent residues, and calculating the aggregation degree of hydrophobic residues in the protein core region. The side chain rotomer energy is obtained from a standard rotomer database using a lookup table method. The stereo collision penalty is calculated using atomic radius overlap integrals. The aggregation degree of hydrophobic residues is defined as the proportion of hydrophobic residues within a 3-nanometer radius sphere at the center of the protein volume. If the sequence fitness assessment score is lower than a preset threshold of 0.7, a sequence-guided reconstruction mechanism is initiated. The fitness assessment score is used as a weight to feed back to the latent space of the variational autoencoder, and a new backbone conformation is resampled and generated until the fitness assessment score is higher than 0.
7. S7. Output the final skeletal image and its corresponding designability score. The designability score is obtained by weighted summation of the contact density score, hydrogen bond network integrity score and sequence fit assessment score, with weight coefficients of 0.4, 0.3 and 0.3, respectively.
3. The method for enhancing the designability of the protein backbone according to claim 2, characterized in that: The construction of the initial backbone generation space in step S1 specifically includes the following steps: First, a three-dimensional coordinate matrix is initialized based on the number of amino acid residues in the target protein. The number of rows in the three-dimensional coordinate matrix equals the number of residues, and the number of columns equals three. Random coordinate values are assigned to the first row of the three-dimensional coordinate matrix as the spatial position of the starting residue. For each subsequent row, the theoretical coordinate position of the current residue is calculated using spherical coordinate transformation based on the coordinate values of the previous row and preset bond length and bond angle parameters. Gaussian noise perturbation is superimposed on the theoretical coordinate position, with the standard deviation of the Gaussian noise set to 0.05 nanometers to introduce conformational diversity. A rigid body transformation is performed on the coordinate matrix after noise superposition, so that its centroid is located at the origin and its principal axes are aligned with the coordinate axes, thus completing the construction of the initial backbone generation space.
4. The method for enhancing the designability of the protein backbone according to claim 3, characterized in that: The operation of the encoder part of the variational autoencoder in step S3 is as follows: the input skeleton coordinate matrix is flattened into a one-dimensional vector; the one-dimensional vector is input into the first fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the first fully connected layer is input into the second fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the second fully connected layer is input into the mean prediction branch and the variance prediction branch, respectively. The mean prediction branch is a fully connected layer with an output dimension of 128, and the variance prediction branch is a fully connected layer with an output dimension of 128. Its output is transformed by an exponential function to ensure a positive value; the latent vector is sampled from the Gaussian distribution defined by the mean and variance parameters.
5. The method for enhancing the designability of the protein backbone according to claim 4, characterized in that: The operation of the decoder part of the variational autoencoder in step S3 is as follows: The sampled latent vector is input into the third fully connected layer, which has 256 neurons and uses a linear rectified function as the activation function; the output of the third fully connected layer is input into the fourth fully connected layer, which has 512 neurons and uses a linear rectified function as the activation function; the output of the fourth fully connected layer is input into the fifth fully connected layer, which has the number of neurons equal to the dimension of the flattened skeleton coordinate matrix and uses an identity function as the activation function; the output of the fifth fully connected layer is reshaped into the original three-dimensional coordinate matrix structure to obtain the reconstructed skeleton coordinates.
6. The method for enhancing the designability of the protein backbone according to claim 5, characterized in that: The specific usage method of the conformation correction module in step S5 includes the following: Calculate the carbon beta interatomic spacing of all non-adjacent residue pairs in the current skeletal image; The number of residue pairs with a spacing less than 0.8 nanometers is counted and divided by the theoretical maximum number of contact pairs to obtain the contact density. If the contact density is less than 0.3, a loss function is constructed, which consists of a negative contact density, a hydrogen bond network disruption penalty term, and a coordinate variation penalty term. The gradient of the loss function is calculated for each atom coordinate in the skeleton coordinate matrix. The coordinates are updated in the opposite direction of the gradient, with a step size of 0.001 nanometers. The gradient update process is repeated until the contact density is greater than or equal to 0.3 or the number of iterations reaches 1,000. After each coordinate update, bond length and bond angle constraints are enforced, and the coordinates are projected back to the constrained manifold using the Lagrange multiplier method.
7. The method for enhancing the designability of the protein backbone according to claim 6, characterized in that: The sequence fitness assessment step in step S6 specifically includes the following: for each residue position in the skeleton configuration, enumerate all twenty standard amino acid side chain conformations; for each side chain conformation, calculate its van der Waals repulsion energy with the main chain and side chain atoms of adjacent residues; select the side chain conformation with the smallest repulsion energy as the optimal side chain at that position; calculate the total repulsion energy of the optimal side chains of all residues, divide it by the number of residues, and obtain the average repulsion energy. Calculate the 3D collision integral of all atomic pairs in the skeleton image. If the sum of the van der Waals radii of any atomic pair is less than the actual spacing, then accumulate the penalty value. The number of hydrophobic residues located in the protein core region is counted. The protein core region is a spatial sphere less than three nanometers away from the centroid of the backbone. The hydrophobic residues include alanine, valine, leucine, isoleucine, phenylalanine, tryptophan, tyrosine, and methionine. The number of hydrophobic residues is divided by the total number of residues in the core region to obtain the hydrophobic aggregation degree. The average repulsion energy is normalized to the interval between zero and one, and the inverse is taken as the repulsion energy score. The stereo collision penalty value is normalized to the interval between zero and one, and the inverse is taken as the collision avoidance score. The hydrophobic aggregation degree is taken as the hydrophobic score. The sequence fitness assessment score is obtained by weighting the repulsion energy score, collision avoidance score, and hydrophobicity score with weights of 0.5, 0.3, and 0.2, respectively.
8. The method for enhancing the designability of the protein backbone according to claim 7, characterized in that: The specific usage method of the sequence-guided reconstruction mechanism in step S6 includes: mapping the current sequence fitness assessment score to a temperature parameter of the latent space sampling, wherein the temperature parameter is negatively correlated with the fitness assessment score, and the calculation formula is that the temperature parameter equals one minus the fitness assessment score; in the latent space of the variational autoencoder, adjusting the standard deviation of the sampling distribution according to the temperature parameter, wherein the standard deviation equals the base standard deviation multiplied by the temperature parameter, and the base standard deviation is set to 0.1; resampling the latent vector from the adjusted Gaussian distribution; inputting the newly sampled latent vector into the decoder to generate a new skeleton image; re-performing the sequence fitness assessment step on the new skeleton image; if the new score is still lower than 0.7, repeating the sampling and assessment process, up to five times; if the target is not met within five times, selecting the skeleton image with the highest score among the five times as the output.
9. The method for enhancing the designability of the protein backbone according to claim 8, characterized in that: The calculation of the designability score in step S7 specifically includes the following steps: multiply the contact density by 100, then by 0.4 to obtain the contact score; count the number of main chain atom pairs in the backbone that satisfy the definition of hydrogen bonds, divide by the theoretical maximum number of hydrogen bonds, then multiply by 100, then by 0.3 to obtain the hydrogen bond score; multiply the sequence fit assessment score by 100, then by 0.3 to obtain the fit score; add the contact score, hydrogen bond score and fit score to obtain the final designability score.
10. The method for enhancing the designability of the protein backbone according to claim 2, characterized in that: The physical energy function in step S4 also includes a hydrogen bond formation tendency term and a structure retention term. The hydrogen bond formation tendency term determines the probability of hydrogen bonds based on the distance and angle between the carbonyl group and the amino group in the main chain. The structure retention term applies an additional energy penalty to the secondary structure region to prevent breakage.
11. The method for enhancing the designability of the protein backbone according to claim 2, characterized in that: During the variational autoencoder training process in step S3, an adversarial generative mechanism is adopted to distinguish between real protein conformations and generated conformations through a discriminator, thereby forcing the latent space distribution to approximate the real conformation manifold.
12. The method for enhancing the designability of the protein backbone according to claim 6, characterized in that: During the use of the configuration correction module, the gradient of the loss function is backpropagated to the latent space, driving the adjustment of latent variables, so that the configuration evolves in a direction that is easier to be filled by the sequence. At the same time, constraint projection is introduced to ensure that the configuration generated after the latent variable adjustment still satisfies the geometric topological constraints and the physical energy function guidance conditions.
13. The method for enhancing the designability of the protein backbone according to claim 7, characterized in that: In the sequence fitness assessment step, a thousand random amino acid sequences are generated using the Monte Carlo sampling method, and the conformational energy of each sequence under the framework is calculated using a fast folding simulator.
14. The method for enhancing the designability of the protein backbone according to claim 8, characterized in that: During the use of the sequence-guided reconstruction mechanism, when all new point scores are below the threshold, the sampling radius is expanded to ten, and the sampling process is repeated until a score that meets the standard is found or the maximum number of samplings is reached.
15. The method for enhancing the designability of the protein backbone according to claim 9, characterized in that: The designability scoring operation also includes geometric rationality scoring and contact pattern quality scoring. The geometric rationality scoring is calculated by checking the main chain bond length and bond angle deviation, the rationality of dihedral angle distribution, and the residue collision situation. The contact pattern quality scoring is obtained by calculating the proportion of hydrophobic residues in the contact pair, the rationality of polar residue pairing, and the long-range contact density.