Wheat germplasm protein structure accurate prediction method based on efficient calculation framework
Through the Cerebra-Triticum algorithm, the polymorphism and temperature dependence problems in wheat protein structure prediction were solved, high-precision protein structure prediction and processing technology optimization were achieved, providing an important tool for wheat quality research and product development.
Patent Information
- Application Number
- CN202510831225.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing protein structure prediction methods fail to fully consider the polymorphism, complex post-translational modification patterns and temperature-dependent conformational changes when dealing with wheat proteins, resulting in poor prediction results.
The end-to-end algorithm framework Cerebra-Triticum is used to accurately predict wheat germplasm protein through feature decoupling, dynamic optimization, and wheat-specific constraints, combined with a multi-task loss function.
It significantly improves the ability to characterize wheat protein structure, accurately predicts three-dimensional conformation, optimizes processing technology, improves prediction accuracy and computational efficiency, and supports the development of wheat protein products with high nutritional value or special functions.
Smart Images

Figure CN120708698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wheat protein structure prediction, and in particular to an end-to-end method for wheat germplasm protein structure prediction. Background Art
[0002] The present invention is aimed at characterization learning of the wheat proteome, and uses computer algorithms to convert the complex information of proteins into a form that can be understood and processed by computers, such as vectors, matrices, etc. Its significance lies in enabling us to use the powerful computing power of computers to study and understand the complexity of proteins, as well as to predict the behavior of proteins. Most existing protein representation methods come from self-supervised language models designed for natural language text. However, the structure and function of proteins are complex and may change in different biological environments. Therefore, how to effectively integrate the sequence, structure and function of proteins to obtain richer multimodal representation information, thereby improving the performance of downstream tasks, such as protein function and protein-protein binding prediction, is an important challenge, and the following problems still need to be faced:
[0003] (1) Wheat proteins are highly polymorphic, contain numerous repetitive sequence structures, and exhibit complex post-translational modification patterns;
[0004] (2) Traditional protein structure prediction methods are often inadequate when dealing with wheat proteins, mainly because they fail to fully consider the special structural characteristics of wheat proteins;
[0005] (3) Many wheat proteins exhibit significant temperature-dependent conformational changes, which are closely related to the processing characteristics of wheat products. Traditional protein structure prediction methods often do not consider this influencing factor;
[0006] In view of the above, the present invention proposes a method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework to solve the above problems. Summary of the Invention
[0007] In view of the above situation, in order to overcome the defects of the existing technology, the present invention provides a method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework. The method aims to propose an end-to-end algorithm framework Cerebra-Triticum for predicting the protein structure of wheat germplasm, which is used to accurately predict the distance between glutenin subunits, protein gluten elasticity, and protein structure under temperature and humidity environmental factors.
[0008] An end-to-end algorithm framework Cerebra-Triticum for predicting protein structure of wheat germplasm is characterized by comprising the following steps:
[0009] S1: In the input layer, the SNP data matrix is encoded and processed, the circular dichroism data is filtered, and the temperature series and humidity data are spliced;
[0010] S2: In the feature decoupling layer, features are automatically allocated through a learnable gating mechanism to reduce the feature overlap of the two paths after decoupling.
[0011] S3: A sliding window attention mechanism is used to predict β-turn dynamics and a 3D convolution kernel is used to scan the molecular surface to identify the distribution of hydrophobic patches.
[0012] S4: For disulfide bond prediction, a graph neural network (GNN) was used to model the relationship between cysteines (Cys);
[0013] S5: When fusion of data, micro-scale, meso-scale and macro-scale are processed respectively;
[0014] S6: Design the Cerebra-Triticum multi-task joint composite loss function, including the main task loss function, the decoupling module loss function, the pathway consistency loss function, and the total loss function.
[0015] The beneficial effects of the above technical solution are:
[0016] (1) The present invention adaptively encodes wheat protein germplasm features, integrates variety SNP data with 3DCNN to construct haplotype-aware residue embedding, and creatively constructs a wheat protein feature decoupling framework, effectively capturing the periodic interactions between repeating units, thereby accurately predicting the helical bundle structure of gliadin, significantly improving the ability to characterize wheat-specific protein structures, enabling the model to more accurately predict the three-dimensional conformation of wheat-specific proteins such as gliadin and glutenin, and adopting differentiated prediction strategies for different types of wheat proteins;
[0017] (2) By strengthening the geometric constraints on glutamine and proline residues, the present invention develops special periodic constraints for the polyglutamine repeat sequences in gliadin, ensuring that the predicted structure can reflect the true conformational characteristics of these repeat sequences, providing theoretical guidance for optimizing processing technology. By simulating the conformational changes of wheat protein under different processing conditions, key parameters in the dough formation process, such as stirring intensity and proofing time, can be precisely controlled.
[0018] (3) The present invention develops an innovative temperature-responsive dynamic optimization strategy. This design fully considers the actual behavior of wheat protein during food processing. The algorithm constructs a temperature-sensitive conformation sampling space, which can automatically adjust the conformation prediction preference according to the set temperature parameters and adjust the conformation prediction strategy according to the set water activity parameters, thereby improving the accuracy of the wheat protein conformation prediction task.
[0019] (4) By integrating conformational prediction results with physicochemical parameter calculations, the present invention can accurately predict functional indicators such as the solubility and surface hydrophobicity of wheat proteins, and can serve as an important auxiliary system for wheat quality research and product development;
[0020] (5) By developing a wheat-specific simplified representation method, the algorithm significantly reduces computational complexity while maintaining prediction accuracy, thus providing feasibility for large-scale wheat proteome analysis.
[0021] (6) This invention analyzes the structural differences in wheat proteins of different varieties and rapidly identifies key mutation sites related to quality characteristics, greatly accelerating the breeding process of high-quality wheat varieties. Cerebra-Triticum's reverse design module can predict wheat protein variants with ideal characteristics based on specific functional requirements, which is of great value in developing wheat protein products with high nutritional value or special functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a schematic diagram of the framework of the method for accurately predicting the protein structure of wheat germplasm of the present invention;
[0023] Figure 2 Schematic diagram showing the comparison of the prediction accuracy of various wheat proteins by Cerebra-Triticum with other mainstream algorithms in a specific embodiment;
[0024] Figure 3 Schematic diagram of the stability evaluation of the Cerebra-Triticum algorithm at different temperatures of the present invention;
[0025] Figure 4 This is a schematic diagram comparing the memory resources consumed under different sequence lengths in a specific implementation method. DETAILED DESCRIPTION
[0026] The aforementioned and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the embodiments with reference to the drawings of this application. The contents mentioned in the following embodiments are all based on the drawings of the specification as a reference.
[0027] Example 1, as Figure 1 As shown in the figure, this study proposes an end-to-end algorithm framework Cerebra-Triticum for wheat germplasm protein structure prediction.
[0028] At the input layer, the SNP matrix was processed using a three-channel encoding (0 = homozygous reference type, 1 = heterozygous type, 2 = homozygous variant type). Haplotype block characteristics were captured through 1D convolution (kernel size = 7, stride = 2), outputting a 256-dimensional germplasm fingerprint. Circular dichroism data were filtered using Savitzky-Golay filtering, and α-helix / β-sheet content was resolved using gradient boosted tree regression with an error setting of <3% to integrate experimental data.
[0029] Note: Circular dichroism is a technique that obtains information by measuring the difference in the absorption of left-handed and right-handed circularly polarized light by proteins. It is based on the circular dichroism effect and the phenomenon of chiral absorption. The secondary structure of proteins has a significant influence on circular dichroism spectra. The connection between circular dichroism (CD) data and single nucleotide polymorphisms (SNPs) is mainly reflected in the intersection of structural biology and functional genomics: by measuring the circular dichroism spectra of proteins or polypeptides, their secondary structures, such as the composition and conformational changes of α-helices and β-sheets, can be analyzed. If the SNP is located in the coding region, that is, a non-synonymous mutation, it may lead to amino acid substitutions, thereby changing the secondary structure of the protein. For example, mutations may destroy the stability of the α-helix (such as proline insertion). With the help of CD data, such conformational changes can be detected, indirectly reflecting the impact of the SNP on the protein structure.
[0030] At the same time, a bidirectional LSTM was used to process the growing season temperature series (time step = 30 days), and the final hidden state was concatenated with the humidity data to form the environmental feature vector.
[0031] At the feature decoupling layer, this project automatically assigns features through a learnable gating mechanism (for example, features with a proline content >15% are prioritized for the Gliadin pathway), reducing feature overlap between the two pathways after decoupling. At this point, proline ring correction uses a restricted Boltzmann machine (RBM) to impose energy constraints on the ring structure.
[0032] Example 2, based on Example 1, designs an energy function, specifically from the following five aspects:
[0033] 1. Decompose technical features. The core process is:
[0034] Conditional trigger: When the proline content of wheat germplasm is >15%, protein structure prediction prioritizes the gliadin pathway (because gliadin is rich in proline, it requires special treatment);
[0035] Proline ring correction: A restricted Boltzmann machine (RBM) is used to impose energy constraints on the proline ring structure to ensure conformational rationality.
[0036] 2. Design the energy function. The expression of the energy function (E) is:
[0037] E=∑(θ-θ exp ) 2 +λ * |ψ-ψ proline |
[0038] In the formula, the parameter definitions and functions are shown in the following table:
[0039]
[0040] 3. Biological logic analysis of energy function (E):
[0041] The first term (dihedral angle constraint): (∑(θ-θ exp ) 2 ) ensures that the protein main-chain conformation is consistent with experimental data and is applicable to all residues.
[0042] The second term (proline-specific constraint): (λ * |ψ-ψ proline |) For the ring structure of proline:
[0043] The ψ angle of proline is restricted by the five-membered ring, so its degree of freedom is significantly lower than that of other residues.
[0044] By (λ * ) weighting (e.g., λ*=2.0) to force the ψ angle to be close to the typical value and avoid unreasonable conformations at the energy minimum.
[0045] 4. Analysis of the role of RBM in structure prediction:
[0046] Restricted Boltzmann Machine (RBM): As a generative model, its hidden layer can learn the latent distribution of proline ring conformations. The energy function is used to:
[0047] Training phase: Optimize the weights so that the lowest energy state corresponds to the true proline ring conformation.
[0048] Prediction stage: Eliminate high-energy (unreasonable) proline conformations through energy constraints.
[0049] Difference from traditional MD: The latent variables of RBM can capture the global statistical properties of the proline ring, while molecular dynamics (MD) only relies on local force fields.
[0050] 5. Practical application examples:
[0051] Scenario: Predict the 3D structure of high-proline wheat Gliadin;
[0052] Input: Amino acid sequence (proline content 18%).
[0053] RBM correction:
[0054] If the simulated value of the ψ angle of a certain proline is -90°, the energy term (|-90-(-60)|=30) will generate a penalty;
[0055] Adjust the ψ angle by gradient descent to minimize the total energy (E);
[0056] Output: The lowest energy protein conformation whose proline ring matches experimental observations.
[0057] Example 3, based on Example 2, uses a sliding window attention mechanism (window size = 9 residues) to predict β-turn dynamics. The molecular surface was scanned to identify the distribution of hydrophobic patches (correlation with starch binding sites, r = 0.71). For spatial scale adaptation, a 9-residue window was used to match the hydrophobic core size of the starch binding domain. For function-guided screening, a strong correlation (r = 0.71) was established to directly link the distribution of hydrophobic patches with starch binding activity. Computational optimization was performed to balance detection sensitivity and false positive rate.
[0058] Example 4: Based on Example 3, for disulfide bond prediction, this project uses a graph neural network (GNN) to model the relationship between Cys residues, and the formula is used for edge feature extraction:
[0059]
[0060] The meanings of the parameters in the formula are shown in the following table:
[0061]
[0062]
[0063] Example 5, based on Example 4, when performing the fusion of SNP matrix, circular dichroism spectrum, and temperature / humidity sequence data, for the microscopic scale, the disulfide bond distance matrix of the Glutenin pathway is directly used to calculate the DS-score; for the mesoscopic scale, the 3D conformations of the two pathways are spliced and input into the voxel CNN (resolution ); for the macroscale, environmental factors and molecular features are integrated to predict farinograph parameters through a 6-layer Transformer.
[0064] Note: The DS-score (Distance-based Solvation Score) is a metric used to quantify the solvent accessibility or distribution of hydrophobic patches on a protein surface. It assesses the propensity of a residue to be exposed to solvent by calculating the spatial distance between the residue and surrounding atoms. It is commonly used to predict protein-ligand binding sites (such as the starch binding domain), analyze protein folding stability (such as the degree of burial of the hydrophobic core), and assist in the design of protein mutations (such as to enhance surface hydrophilicity).
[0065] Example 6: Cerebra-Triticum uses a multi-task joint optimization framework, and its composite loss function consists of three parts:
[0066] 1. Main task loss function:
[0067] The main task loss function is classification loss, and the improved Focal Tversky Loss function is used to optimize the class imbalance problem, as shown in the following formula:
[0068]
[0069] Where TP c FP c 、FN c They represent the true positive, false positive, and false negative of category c, respectively. α = 0.3, β = 0.7 are weights, the purpose of which is to determine through grid search and give higher penalties to false negatives; γ = 2 is the focusing coefficient, which suppresses the gradient of easy-to-classify samples. L cls is the classification loss, Tversky c is the Tversky index, which is used to measure the overlap between the predicted segmentation (such as the protein residue contact map) and the true label. It is a generalized form of the exponential coefficient.
[0070] 2. Decoupling module loss function:
[0071] The decoupling module loss function mainly refers to the feature separation loss. The Hilbert-Schmidt Independence Criterion (HSIC) is used to force different types of features to be orthogonal. The loss function formula is:
[0072]
[0073] K P =κ(F P ,F P ),K S =κ(F S ,FS )
[0074] Where, F P 、F S Represent different categories of feature matrices, κ(·,·) is the radial basis function (RBF) kernel function, is the centralization matrix, K p , K s are the kernel matrices used to calculate the joint similarity of two types of features. , L dec is the decoupling loss.
[0075] 3. Path consistency loss function:
[0076] The path consistency loss design uses bidirectional KL divergence to constrain the dual-path output distribution, and its loss function is:
[0077]
[0078] p i =softmax(z i / τ),τ=2.0
[0079] The parameters in the formula are shown in the following table:
[0080]
[0081] Note: Softmax (normalized exponential function) is a commonly used nonlinear activation function, which is mainly used to convert a set of real numbers (usually the output of a neural network) into a probability distribution so that the sum of all output values is 1 and each value is in the range [0,1].
[0082] The total loss function after combination:
[0083] When calculating the total loss function, the temperature coefficient λ is used to dynamically balance the loss terms:
[0084] L total =L cls +λ(t)(L dec +L con )
[0085]
[0086] In the formula, its hyperparameter settings are shown in the following table:
[0087] parameter value illustrate <![CDATA[λ max ]]> 1 Initial weight λmin 0.2 Final weight <![CDATA[T max ]]> 50 Decay cycles (epochs)
[0088] Implementation 7, such as Figure 2Figure 2 shows the Cerebra-Triticum algorithm's analysis of wheat protein feature decoupling. Traditional protein structure prediction methods often struggle with wheat proteins, primarily due to a failure to fully account for their unique structural characteristics. The Cerebra-Triticum algorithm creatively constructs a wheat protein feature decoupling framework, a design inspired by a deep understanding of wheat proteomics. The algorithm first establishes a feature representation space specifically for wheat proteins, decomposing their structural information into three interrelated yet relatively independent feature dimensions: gliadin feature space, glutenin feature space, and dynamic modification feature space.
[0089] The gliadin feature space focuses on analyzing the repeat sequence structure in wheat proteins, especially the repeat regions rich in glutamine and proline. Through the design of special convolution kernels, this feature space can effectively capture the periodic interactions between repeat units, thereby accurately predicting the helical bundle structure of gliadin. The glutenin feature space focuses on the disulfide bond network in wheat proteins, and by establishing geometric constraints between sulfur atoms, it accurately models the cross-linking structure between glutenin subunits. The dynamic modification feature space specifically handles post-translational modifications unique to wheat proteins, including glycosylation, phosphorylation, and other modification types that have a significant impact on protein conformation.
[0090] This feature decoupling design for wheat proteins brings many technical advantages. First, it significantly improves the ability to characterize the structure of wheat-specific proteins, enabling the model to more accurately predict the three-dimensional conformation of wheat-specific proteins such as gliadin and glutenin. Second, this decoupling strategy enables the model to adopt differentiated prediction strategies for different types of wheat proteins. For example, when dealing with highly repetitive gliadin, it focuses on periodic feature extraction, while when predicting disulfide-rich glutenin, it strengthens the geometric constraints between sulfur atoms. Most importantly, this feature decoupling enables the model to better handle the polymorphism of wheat proteins, providing a reliable computational tool for the study of protein structure variation in different varieties of wheat.
[0091] Implementation 8, such as Figure 3 Figure 4 shows the dynamic conformational optimization analysis of wheat proteins using the Cerebra-Triticum algorithm. Conformation prediction for wheat proteins presents a unique challenge: many exhibit significant temperature-dependent conformational changes, which are closely related to the processing properties of wheat products. The Cerebra-Triticum algorithm developed an innovative temperature-responsive dynamic optimization strategy that fully accounts for the actual behavior of wheat proteins during food processing.
[0092] The algorithm constructs a temperature-sensitive conformational sampling space, which can automatically adjust the conformational prediction preference according to the set temperature parameters. In the low temperature range (such as 4-25°C), the model tends to predict a relatively compact protein conformation, which is consistent with the actual state of wheat protein at room temperature. As the temperature rises, the model will gradually introduce more flexible conformational sampling to simulate the conformational relaxation of wheat protein during heating. Especially in the critical temperature range of 60-90°C, the algorithm will focus on optimizing the conformational transition path of the protein unfolding process. This feature is of great significance for understanding the rheological behavior of wheat dough.
[0093] Another key innovation of the dynamic optimization mechanism is its moisture responsiveness. The algorithm adjusts the conformational prediction strategy based on a set water activity parameter, simulating the conformational changes of wheat proteins under different hydration states. This feature enables the model to accurately predict the different conformational states of wheat proteins in low-moisture environments (such as dry storage) and high-moisture environments (such as dough formation). Experimental data show that this dynamic optimization strategy improves the model's accuracy in wheat protein conformation prediction by 25%. In particular, when simulating conformational changes during heating, the correlation coefficient between the predicted results and experimental observations reached 0.89.
[0094] Implementation 9, such as Figure 4 Figure 2 shows the biological plausibility analysis of the Cerebra-Triticum algorithm's wheat-specific constraints. Ensuring that predictions conform to the inherent properties of wheat proteins is a core design principle of the Cerebra-Triticum algorithm. By introducing multiple levels of wheat-specific constraints, the algorithm significantly improves the biological plausibility of predicted structures.
[0095] At the amino acid composition level, the algorithm specifically strengthens geometric constraints on glutamine and proline residues, as these two amino acids are present in a very high proportion in wheat proteins and have a decisive influence on conformation. Specifically, the algorithm develops special periodic constraints for the polyglutamine repeats in gliadin to ensure that the predicted structure reflects the true conformational characteristics of these repeats. Regarding disulfide bond processing, the algorithm not only considers conventional cysteine pairings but also pays special attention to special cross-linking forms such as trisulfide bonds, which are common in wheat proteins.
[0096] The algorithm incorporates extensive prior knowledge regarding post-translational modifications of wheat proteins. For glycosylation, the algorithm automatically adjusts local conformational constraints based on predicted glycosylation sites, accurately reflecting the impact of sugar chains on protein structure. For phosphorylation, the algorithm predicts conformational changes caused by the introduction of phosphate groups, particularly alterations in protein surface charge distribution. These wheat-specific constraints enable the algorithm to produce predictions that are not only geometrically sound but also biologically meaningful.
[0097] Implementation 10, such as Figure 2 、 3 As shown in Figures 4 and 5, the performance of the Cerebra-Triticum algorithm was analyzed and evaluated. Cerebra-Triticum demonstrated excellent predictive performance on a specially constructed wheat protein structure evaluation set. Compared to general protein prediction methods, the algorithm's advantages in wheat protein prediction are particularly evident. For complex proteins such as glutenin, the algorithm's prediction accuracy is over 30% higher than that of general methods. When predicting polymorphic variations in wheat proteins, the algorithm is able to accurately capture conformational changes caused by single amino acid variations, providing a molecular-level explanation for the quality differences between wheat varieties.
[0098] Of particular note is the algorithm's ability to predict the functional properties of wheat proteins. By integrating conformational predictions with calculated physicochemical parameters, the algorithm accurately predicts functional properties such as solubility and surface hydrophobicity of wheat proteins. These predictions are highly consistent with experimental measurements, with correlation coefficients generally exceeding 0.85. This makes the algorithm not only a structural prediction tool but also a valuable auxiliary system for wheat quality research and product development.
[0099] In terms of computational efficiency, Cerebra-Triticum is specifically optimized for the characteristics of wheat proteins. By developing a simplified representation method specific to wheat species, the algorithm significantly reduces computational complexity while maintaining prediction accuracy. On standard hardware, a structure prediction for a typical wheat protein (approximately 500 residues) takes only approximately 40 minutes, making it feasible for large-scale wheat proteome analysis.
[0100] In summary, Cerebra-Triticum has achieved a quantum leap in wheat protein structure prediction by integrating innovative technologies such as multi-scale feature characterization, dynamic conformational optimization, and wheat-specific constraints. The development of the Cerebra-Triticum algorithm has brought new opportunities to wheat-related research and industry. In agricultural breeding, the algorithm's high-precision prediction capabilities provide a powerful tool for molecular marker-assisted breeding. By analyzing the structural differences in wheat proteins across different varieties, breeders can rapidly identify key variants associated with quality traits, significantly accelerating the selection of high-quality wheat varieties.
[0101] In the flour processing industry, the algorithm's conformational prediction capabilities provide theoretical guidance for optimizing processing techniques. By simulating the conformational changes of wheat proteins under different processing conditions, key parameters in the dough formation process, such as mixing intensity and proofing time, can be precisely controlled. Preliminary applications have shown that algorithm-guided process optimization improves dough quality consistency.
[0102] The development of functional wheat proteins is another promising application area. Cerebra-Triticum's reverse design module predicts wheat protein variants with desirable properties based on specific functional requirements. This capability is invaluable in developing wheat protein products with high nutritional value or specialized functions. The algorithm has been successfully applied to several wheat protein design projects, including the development of hypoallergenic and highly emulsifying wheat proteins.
[0103] The above description is only for illustrating the present invention. It should be understood that the present invention is not limited to the above embodiments, and various variations that conform to the concept of the present invention are within the scope of protection of the present invention.
Claims
1. Based on an efficient computing framework, an end-to-end algorithm framework Cerebra-Triticum for wheat germplasm protein structure prediction is proposed, which is characterized by: The following steps are involved: S1: In the input layer, the SNP data matrix is encoded and processed, the circular dichroism data is filtered, and the temperature series and humidity data are spliced; S2: In the feature decoupling layer, features are automatically allocated through a learnable gating mechanism to reduce the feature overlap of the two paths after decoupling. S3: A sliding window attention mechanism is used to predict β-turn dynamics and a 3D convolution kernel is used to scan the molecular surface to identify the distribution of hydrophobic patches. S4: For disulfide bond prediction, a graph neural network (GNN) was used to model the relationship between cysteines (Cys); S5: When fusion of data, micro-scale, meso-scale and macro-scale are processed respectively; S6: Design the Cerebra-Triticum multi-task joint composite loss function, including the main task loss function, the decoupling module loss function, the pathway consistency loss function, and the total loss function.
2. The method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework according to claim 1, characterized in that: The step S1 comprises: S1-1: Three-channel encoding is applied to the SNP data matrix, including 0 = homozygous reference type, 1 = heterozygous type, and 2 = homozygous variant type. Haplotype block characteristics are captured by 1D convolution, where the convolution kernel size = 7 and the stride = 2, and a 256-dimensional germplasm fingerprint is output; S1-2: After the circular dichroism data are filtered by a polynomial smoothing filter (Savitzky-Golay), the α-helix / β-sheet content is analyzed using gradient boosting tree regression to achieve integration of the filtered circular dichroism data; S1-3: Use a bidirectional LSTM to process the growing season temperature series (time step = 30 days), and finally concatenate the LSTM output gate hidden state with the humidity data to form an environmental feature vector.
3. The method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework according to claim 2, characterized in that: The step S2 includes that when the proline content is greater than 15%, the feature preferentially enters the gliadin pathway, and the proline ring correction uses a restricted Boltzmann machine (RBM) to impose energy constraints on the ring structure, and the energy function is designed as: E=∑(θ-θ exp ) 2 +λ*|ψ-ψ proline | Where λ * refers to the penalty weight of the proline ring conformational deviation, ψ refers to the current proline residue angle, ψ proline refers to the theoretical angle of proline residues, θ refers to the simulated calculated value of the current dihedral angle in the protein backbone, and θ exp Refers to the actual measured dihedral angle reference value.
4. The method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework according to claim 1, characterized in that: The step S3 includes a window size of 9 residues and a correlation with starch binding sites of r=0.
71.
5. The method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework according to claim 1, characterized in that: The step S4 includes edge feature extraction using the formula: Among them, q i ,q j is the charge, S ij is the solvent accessible surface area, w ss is the secondary structure weight coefficient, is the interatomic distance, and ε0 is the dielectric constant of vacuum.
6. The method for accurately predicting the protein structure of wheat germplasm based on an efficient computing framework according to claim 1, characterized in that: The step S5 specifically includes: For the microscale, the distance-based solvation score (DS-score) was calculated directly using the disulfide distance matrix of the glutenin pathway; For the mesoscopic scale, the 3D conformations of the dual pathways of Gliadin and Glutenin were spliced and input into the voxel CNN; For the macroscale, environmental factors and molecular characteristics are integrated to predict the farinograph parameters through a 6-layer neural network Transformer.