Wheat protein structure prediction method, device and equipment and storage medium

By performing subgenomic sequence separation and weight fusion on the wheat protein structure prediction model, and combining iterative training with self-attention mechanism and inter-subunit distance constraint term, the problem of low accuracy in wheat protein structure prediction was solved, and the accuracy of protein structure prediction and the predictive ability of wheat quality formation were improved.

CN120808890APending Publication Date: 2025-10-17HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510912772.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively distinguish the structural differences between different subgenome subunits in wheat, resulting in low accuracy in wheat protein structure prediction and affecting the accuracy of protein-protein interaction interface prediction.

Method used

By acquiring wheat protein data, a pre-defined wheat prediction model was used to separate and weight subgenome sequences. Iterative training was then conducted using a self-attention mechanism and inter-subunit distance constraints to optimize the model architecture and improve prediction accuracy.

Benefits of technology

It improves the accuracy of wheat protein structure prediction, especially the prediction accuracy of protein-protein interaction interfaces, and enhances the predictive ability of wheat quality formation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808890A_ABST
    Figure CN120808890A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat protein structure prediction method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining to-be-processed wheat protein data; inputting the wheat protein data into a preset wheat prediction model, and performing prediction processing on a subgenome sequence of the wheat protein data based on the preset wheat prediction model to obtain a protein structure prediction result, the protein structure prediction result being obtained by separating the subgenome sequence to obtain a protein structure; and carrying out fusion processing on the separated subgenome sequences according to a weight ratio to obtain the subgenome. The technical effect of improving the prediction precision of the protein structure is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a wheat protein structure prediction method and device, equipment and a storage medium. BACKGROUND

[0002] Breakthroughs in protein structure prediction technology are reshaping the paradigm of crop functional genomics research. As one of the most important food crops in the world, the hexaploid genome characteristics of common wheat (Triticum aestivum L.) make its protein interaction network present unique complexity. The high homology among the A, B and D subgenomes of wheat provides a genetic basis for environmental adaptability, but also brings special challenges to protein structure prediction.

[0003] At present, when dealing with wheat proteins, the traditional multiple sequence alignment (MSA) method usually regards the hexaploid genome as a single information source, which leads to the serious dilution of the coevolution signals of homologous sequences among subgenomes. This way cannot effectively distinguish the structural differences of different subgenomic coding group subunits, which directly affects the prediction accuracy of protein-protein interaction interfaces that are crucial for wheat quality formation, resulting in low prediction accuracy of wheat protein structure. SUMMARY

[0004] The main purpose of the present application is to provide a wheat protein structure prediction method, device, equipment and storage medium, which aims to solve the technical problem in the related art that different subgenomic coding group subunits cannot be effectively distinguished, which directly affects the prediction accuracy of protein-protein interaction interfaces that are crucial for wheat quality formation, resulting in low prediction accuracy of wheat protein structure.

[0005] To achieve the above-mentioned purpose, the embodiments of the present application provide a wheat protein structure prediction method, comprising:

[0006] obtaining wheat protein data to be processed;

[0007] inputting the wheat protein data into a preset wheat prediction model, performing prediction processing on the subgenome sequences of the wheat protein data based on the preset wheat prediction model, and obtaining a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the subgenome sequences and then fusing the separated subgenome sequences according to a weight ratio.

[0008] In a possible implementation manner of the present application, the protein structure prediction result is obtained by performing prediction processing on the subgenome sequences of the wheat protein data based on the preset wheat prediction model, comprising:

[0009] Separate the sub-genome sequences in the wheat protein data based on the preset wheat prediction model to obtain a plurality of sub-genomic sequences;

[0010] Distribute weights to each of the sub-genomic sequences and fuse the sub-genomic sequences after weight distribution to obtain a fused gene sequence;

[0011] Perform prediction processing on the fused gene sequence to obtain a protein structure prediction result.

[0012] In a possible implementation of the present application, the separating the sub-genome sequences in the wheat protein data based on the preset wheat prediction model to obtain a plurality of sub-genomic sequences comprises:

[0013] Extracting a preset number of sub-genome sequences in the wheat protein data based on the preset wheat prediction model;

[0014] Obtaining homologous sequences from each of the sub-genome sequences;

[0015] Determining a plurality of sub-genomic sequences based on the homologous sequences.

[0016] In a possible implementation of the present application, the distributing weights to each of the sub-genomic sequences and fusing the sub-genomic sequences after weight distribution to obtain a fused gene sequence comprises:

[0017] Determining sequence conservation and gene expression of the sub-genomic sequences;

[0018] Calculating weight coefficients of each of the sub-genomic sequences based on the sequence conservation and the gene expression;

[0019] Fusing each of the sub-genomic sequences according to the weight coefficients to obtain a fused gene sequence.

[0020] In a possible implementation of the present application, before the inputting the wheat protein data into the preset wheat prediction model, the method further comprises:

[0021] Obtaining protein sample data and gene sequence label data;

[0022] Iteratively training a wheat prediction model to be trained based on the protein sample data and the gene sequence label data until a prediction index corresponding to the wheat prediction model to be trained reaches a preset accuracy standard to obtain a trained wheat prediction model.

[0023] In a possible implementation of the present application, the protein sample data and the gene sequence label data are repeatedly detected by a self-attention mechanism in the to-be-trained wheat prediction model to obtain a repeated region, and a bias term is calculated for the repeated region.

[0024] The protein sample data and the gene sequence label data are repeatedly detected by a self-attention mechanism in the to-be-trained wheat prediction model to obtain a repeated region, and a bias term is calculated for the repeated region.

[0025] A subunit distance constraint term is added to a loss function of the to-be-trained wheat prediction model, and the to-be-trained wheat prediction model is iteratively calculated multiple times by a preset strategy until the loss function converges, and a trained wheat prediction model is obtained.

[0026] In a possible implementation of the present application, the preset strategy includes a computation flow reorganization algorithm, and the to-be-trained wheat prediction model is iteratively calculated multiple times by the preset strategy, including:

[0027] Based on the computation flow reorganization algorithm, the activation values of each network level in the to-be-trained wheat prediction model are determined.

[0028] For a network layer node with a dependency degree of the activation value less than or equal to a preset threshold, the activation value of the current network layer node is discarded.

[0029] For a network layer node with a dependency degree of the activation value greater than a preset threshold, the activation value of the current network layer node is stored for iterative training of the wheat prediction model.

[0030] The present application also provides a wheat protein structure prediction device, including:

[0031] An acquisition module is configured to acquire wheat protein data to be processed.

[0032] A processing module is configured to input the wheat protein data into a preset wheat prediction model, perform prediction processing on a subgenomic sequence of the wheat protein data based on the preset wheat prediction model, and obtain a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the subgenomic sequence and then fusing the separated subgenomic sequence according to a weight ratio.

[0033] The application further provides a wheat protein structure prediction device, which is a physical node device, and comprises a memory, a processor and a program of the wheat protein structure prediction method stored on the memory and executable on the processor, which can realize the steps of the wheat protein structure prediction method as described above when executed by the processor.

[0034] To achieve the above object, the application further provides a storage medium having a wheat protein structure prediction program stored thereon, which can realize the steps of any of the above wheat protein structure prediction methods when executed by a processor.

[0035] The application provides a wheat protein structure prediction method, device, equipment and storage medium. In the related art, different subgenomic coding group subunits cannot be effectively distinguished in structure, which directly affects the prediction accuracy of the protein-protein interaction interface that is crucial for wheat quality formation, resulting in low prediction accuracy of wheat protein structure. In the present application, the wheat protein data to be processed is obtained, and the wheat protein data is input into a preset wheat prediction model. The subgenomic sequence of the wheat protein data is predicted based on the preset wheat prediction model to obtain a protein structure prediction result. The protein structure prediction result is obtained by separating the subgenomic sequence and then fusing the separated subgenomic sequence according to a weight ratio, so as to consider multiple dimensions of the subgenomic sequence, rather than regarding the homologous sequence in the subgenomic sequence as a single information source. The model is used to predict the multiple dimensions of the subgenomic sequence to obtain the protein structure prediction result, thereby improving the prediction accuracy of the protein structure. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 Flowchart of the first embodiment of the wheat protein structure prediction method of the present application;

[0037] Figure 2 Model prediction flowchart involved in the wheat protein structure prediction method of the present application;

[0038] Figure 3 Flowchart of the second embodiment involved in the wheat protein structure prediction method of the present application;

[0039] Figure 4 Training calculation flowchart involved in the wheat protein structure prediction method of the present application;

[0040] Figure 5 Device structure diagram of the hardware running environment involved in the embodiment of the present application. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are merely exemplary and do not limit the application.

[0042] The embodiment of the application provides a wheat protein structure prediction method, in the first embodiment of the wheat protein structure prediction method, referring to Figure 1 , the method comprises the following steps:

[0043] Step S10, acquiring wheat protein data to be processed;

[0044] Step S20, inputting the wheat protein data into a preset wheat prediction model, performing prediction processing on a sub-genome sequence of the wheat protein data based on the preset wheat prediction model, and obtaining a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the sub-genome sequence, and then performing fusion processing on the separated sub-genome sequence according to a weight ratio.

[0045] The embodiment aims to improve the prediction accuracy of protein structure.

[0046] The specific steps are as follows:

[0047] Step S10, acquiring wheat protein data to be processed;

[0048] As an example, the wheat protein structure prediction method can be applied to a wheat protein structure prediction device, and the wheat protein structure prediction device belongs to a wheat protein structure prediction system, and the wheat protein structure prediction system belongs to a wheat protein structure prediction equipment.

[0049] As an example, the wheat protein data can be wheat protein gene data obtained from a wheat protein database, for example, alpha-gliadin (Gliadin).

[0050] Step S20, inputting the wheat protein data into a preset wheat prediction model, performing prediction processing on a sub-genome sequence of the wheat protein data based on the preset wheat prediction model, and obtaining a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the sub-genome sequence, and then performing fusion processing on the separated sub-genome sequence according to a weight ratio.

[0051] As an example, the preset wheat prediction model can be the WheatFold model, which is built based on the open source framework of the AlphaFold2 model. In this model, the hexaploid genome characteristics are integrated into the structure prediction process. Different from the single sequence modeling of AlphaFold2, the WheatFold model designs a three-level information processing flow, such as Figure 2 As shown by Figure 2 It can be seen that the preset wheat prediction model mainly includes an input processing layer, a feature fusion core and a prediction engine. Among them, the input processing layer is mainly used to separate each subgenome, and the feature fusion core is used to fuse the classified subgenome sequences according to a certain weight ratio. Finally, the prediction processing is performed through the prediction engine, and then the protein structure prediction results are output.

[0052] As an example, the model architecture of the preset wheat prediction model follows the principle of modularity. Each component seamlessly connects to the original module through an interface, retaining the general prediction capabilities of AlphaFold2 while deeply optimizing the biological characteristics of hexaploid wheat. This architecture retains the core components of AlphaFold2 (MSA processing, Evoformer iteration, and structure generation) and adds three key modules:

[0053] Multi-subunit collaborative modeling module: By building an attention mechanism across subgenomes, it solves the systematic bias of traditional single-subunit prediction.

[0054] Repeat domain compensation network: Design attention bias items for the repetitive sequences commonly found in wheat proteins (such as the PolyQ region of alcohol-soluble proteins).

[0055] Memory optimization engine: uses dynamic gradient checkpoint technology to break through hardware resource limitations.

[0056] The step S20 of performing prediction processing on the subgenomic sequence of the wheat protein data based on the preset wheat prediction model to obtain a protein structure prediction result includes:

[0057] Step S21, based on the preset wheat prediction model, separating and processing the subgenomic sequences in the wheat protein data to obtain a plurality of subgenomic sequences;

[0058] As an example, the purpose of separation processing is mainly to divide homologous sequences into genes according to the type of subunits, so as to avoid treating different types of subgenomes as the same information source, where subgenomes are separated from the same subgenome sequence, for example, Figure 2 The A subgenome and B subgenome are subgene sequences.

[0059] The step S21 of separating the sub-genome sequences in the wheat protein data based on the preset wheat prediction model to obtain a plurality of sub-gene sequences comprises:

[0060] Extracting a preset number of sub-genome sequences in the wheat protein data based on the preset wheat prediction model;

[0061] As an example, A, B and D sub-genome sequences are extracted from the wheat protein database by the preset wheat prediction model, and the non-protein coding region is filtered by the TriAnnot annotation system.

[0062] Obtaining homologous sequences from each of the sub-genome sequences;

[0063] Determining a plurality of sub-gene sequences based on the homologous sequences.

[0064] As an example, for each target protein, homologous sequences (E-value <1e-5) are obtained from three sub-genomes respectively, a plurality of sub-gene sequences are obtained, and an independent MSA (multiple sequence alignment) matrix is constructed according to each sub-gene sequence.

[0065] The step S22 of assigning weights to each of the sub-gene sequences and fusing the sub-gene sequences after weight assignment to obtain a fused gene sequence;

[0066] As an example, the fusion processing can be performed in the following manner: first, the weight coefficients of each sub-gene sequence are calculated, then the sub-gene sequences are fused according to the weight coefficients, and finally, the fused gene sequence is obtained.

[0067] The step S22 of assigning weights to each of the sub-gene sequences and fusing the sub-gene sequences after weight assignment to obtain a fused gene sequence;

[0068] Determining the sequence conservation and gene expression of the sub-gene sequence;

[0069] Based on the sequence conservation and the gene expression, the weight coefficient of each sub-gene sequence is calculated;

[0070] As an example, the sub-genome weight coefficient is introduced in the MSA embedding stage, wherein the calculation method of the weight coefficient can be:

[0071]

[0072] Wherein:

[0073] - Sequence conservation of subgenomic I (based on Clustal Omega alignment score)

[0074] - Gene expression of subgenomic I (TPM value from wheat eFP browser)

[0075] - a = 0.6, b = 0.4: optimal weight coefficients determined by ablation experiments.

[0076] The subgenomic sequences are fused according to the weight coefficients to obtain a fused gene sequence.

[0077] As an example, the MSA matrix of the original AlphaFold2 model is re-fused according to the above weight coefficients, so that the optimized results of the present application can still be subjected to the next action by the remaining modules of the AlphaFold2 model.

[0078] Step S23, the fused gene sequence is subjected to a prediction process to obtain a protein structure prediction result.

[0079] As an example, the prediction process mainly includes geometric feature modeling by the evoformer improvement module, three-dimensional coordinate prediction of the gene sequence by the structure generator, and protein structure prediction result after subunit constraint optimization.

[0080] As an example, taking a-gliadin as an example, collect its homologous sequences in A / B / D subgenomic (500 each), construct a three-dimensional MSA tensor. Strengthen the contribution of D subgenomic in the Evoformer input layer through the weight distributor (because it is related to allergenicity), and finally the matching degree of the predicted IgE binding site with the experimental data reaches 82.3%.

[0081] The application provides a wheat protein structure prediction method, device, equipment and storage medium. In the related art, different subgenomic coding group subunits cannot be effectively distinguished in structure difference, which directly affects the prediction accuracy of the protein-protein interaction interface that is crucial for wheat quality formation, resulting in low prediction accuracy of wheat protein structure. In the application, the wheat protein data to be processed is obtained, and the wheat protein data is input into a preset wheat prediction model. The subgenomic sequence of the wheat protein data is predicted based on the preset wheat prediction model to obtain a protein structure prediction result. The protein structure prediction result is obtained by separating the subgenomic sequence and then fusing the separated subgenomic sequence according to a weight ratio, so as to consider multiple dimensions of the subgenomic sequence, instead of regarding the homologous sequence in the subgenomic sequence as a single information source. The model is used to predict the multiple dimensions of the subgenomic sequence to obtain the protein structure prediction result, thereby improving the prediction accuracy of the protein structure.

[0082] Further, with reference to Figure 3 Based on the first embodiment of the application, another embodiment of the application is provided, in which the wheat protein data is input into a preset wheat prediction model, and the method further comprises steps S11-S12 before the wheat protein data is input into the preset wheat prediction model.

[0083] In step S11, protein sample data and gene sequence label data are obtained.

[0084] As an example, before the model is used to predict the protein structure, the wheat prediction model needs to be iteratively trained, and the protein sample data and the gene sequence label data are obtained. The model is iteratively trained by using the protein sample data and the gene sequence label data.

[0085] As an example, in this embodiment, the AlphaFold2 model is fine-tuned using the wheat protein data, so that it can enhance the prediction accuracy of the wheat protein at the cost of reducing the prediction accuracy of other proteins.

[0086] As an example, before the training data is input into the model to be trained, first, data screening is performed. From the UniProt wheat protein library, 15,000 sequences that meet the following conditions are screened:

[0087] (1) Length 50-800 aa (covering 95% of wheat proteins).

[0088] (2) Pfam annotation covers 15 core families (such as PF00420 gliadin, PF13016 prolamin).

[0089] (3) Contains experimentally validated structures (PDB / cryo-EM) or high-confidence homology models.

[0090] Further, the protein data screened is subjected to data enhancement. The data amount of the screened protein data is too small to have a great impact on the original model parameters of AlphaFold2. In order to increase the influence of this part of data, a data enhancement strategy is constructed to make the original 15,000 data grow by more than ten times, increase the proportion of wheat data in the original AlphaFold2 training data, and thus enhance the prediction ability of wheat protein. Of course, this enhancement of ability is not without cost. The additional fine-tuning training will produce a forgetting effect, reducing the prediction ability of the original model for predicting the protein structure of other species.

[0091] The specific method is: for low-resolution structures (EMDB-10080, etc.), RosettaRelax is used for conformation optimization. Ten conformation variants are generated by MCMC sampling of AlphaFold2, and the conformation with pLDDT>90 is taken as the supplementary training data.

[0092] Step S12, based on the protein sample data and the gene sequence label data, the wheat prediction model to be trained is iteratively trained until the prediction index corresponding to the wheat prediction model to be trained reaches the preset accuracy standard, and a trained wheat prediction model is obtained.

[0093] As an example, the prediction index can be prediction confidence. When iteratively training the wheat prediction model to be trained, the prediction confidence can reach a certain value, or the loss function converges, and the iterative training of the model is stopped.

[0094] The step S12 of iteratively training the wheat prediction model to be trained based on the protein sample data and the gene sequence label data until the prediction index corresponding to the wheat prediction model to be trained reaches the preset accuracy standard to obtain the trained wheat prediction model comprises:

[0095] The protein sample data and the gene sequence label data are repeatedly detected by the self-attention mechanism in the wheat prediction model to be trained to obtain a repeated region, and a bias term is calculated for the repeated region.

[0096] As an example, in order to capture the unique rules of the protein sequence of the hexaploid gene, the embodiment of the application inserts a repeat domain perception module at the 6th and 12th layers of the Evoformer, and detects the repeat pattern through a sliding window (window length=12aa). For the detected repeated region, a bias term is added when calculating self-attention:

[0097]

[0098] wherein M repeat is a repeat mask matrix generated from the sliding window detection result.

[0099] In the original AlphaFold2 training data, human (diploid gene) protein data accounts for the majority. The repeat domain of human protein (diploid gene) is far less than the quantity of wheat (hexaploid gene), so there is no need to set and weight the repeat domain perception module. This module is not available in the original AlphaFold2.

[0100] A subunit distance constraint term is added to the loss function of the wheat prediction model to be trained, and the wheat prediction model to be trained is calculated multiple times through a preset strategy until the loss function converges, thereby obtaining a trained wheat prediction model.

[0101] As an example, in the loss function, the embodiment of the application adds a subunit distance constraint term to the original AlphaFold2 basic loss L base = L lddt + L dist + L angle The basic loss is added to the subunit distance constraint term:

[0102]

[0103] Wherein:

[0104] - represents the predicted subunit i / j inter-residue distance (Cβ atom)

[0105] - represents the contact distance verified by cryo-EM

[0106] -N contacts represents the number of interface residue pairs

[0107] The final loss function is:

[0108] L total = L base + λ·L interface

[0109] Wherein, λ = 0.3 (determined by ablation experiment, the experience value obtained by the present application through a large number of experiments is 0.3).

[0110] As an example, by optimizing the loss function and adding a subunit distance constraint term, the prediction accuracy of the model is improved.

[0111] The preset strategy includes a computational flow reorganization algorithm.

[0112] Based on the computational flow reorganization algorithm, the activation values of each network level in the wheat prediction model to be trained are determined.

[0113] As an example, in a large number of wheat protein prediction experiments, it is found through PyTorchMemoryProfiler quantification that the original AlphaFold2 has the following defects in protein prediction of hexaploid gene crops:

[0114] The Evoformer layer accounts for 68% of the peak video memory, and the MSA axial attention module (AxialAttention) single layer accounts for 12 GB of video memory.

[0115] The structure generation module accounts for 22%, mainly consuming in the calculation of the SE(3) transformation matrix.

[0116] The temporary cache accounts for 10%, which can be optimized by reorganizing the computational flow.

[0117] If the prediction target is mostly hexaploid gene crops, the above defects will cause considerable waste of GPU computing resources.

[0118] In view of this, the embodiment of the application proposes a "layered discarding-on-demand recalculation" strategy. The calculation parameters of each layer of the original AlphaFold are no longer statically reserved, and a small amount of recalculation is used to save a large amount of video memory.

[0119] As an example, the activation value can be a dynamic output value generated by the neural network in the forward propagation process, which is generated after weighted calculation, bias adjustment and activation function processing of input data,

[0120] For the network layer node with a dependency degree of the activation value less than or equal to a preset threshold, the activation value of the current network layer node is discarded;

[0121] As an example, for the MSA embedding layer, the initial state of the structure module, and the final folding result frequently referenced in subsequent calculations, they are marked as must be stored.

[0122] As an example, the preset threshold can be 0.7, 0.8, etc., and is not limited in particular.

[0123] As an example, for the low-dependency intermediate layer activation value, it is discarded after use

[0124] For the network layer node with a dependency degree of the activation value greater than a preset threshold, the activation value of the current network layer node is stored for iterative training of the wheat prediction model.

[0125] As an example, for the Evoformer intermediate layer activation value, only store the nodes with gradient dependence > 0.7 during forward propagation, and store the activation value of the current network layer node for iterative training of the wheat prediction model.

[0126] As an example, the success of the "hierarchical dropout-on-demand recomputation" strategy lies in on-demand recomputation. Therefore, a calculation flow reorganization algorithm is proposed in the present application to determine the temperature (i.e., the probability of needing to be recomputed) T recompute .

[0127]

[0128] wherein:

[0129] -L active : number of layers that need to be recomputed;

[0130] -L checkpoint : number of layers to store checkpoints;

[0131] -L total : total number of Evoformer layers;

[0132] Since the same algorithm is used to calculate the temperature during storage and calculation, it is easy to determine which layers to directly take the stored values and which layers to recompute in the original AlphaFold2 calculation process.

[0133] In practice, according to the values obtained by applying the above algorithm in the experience process (such as the wheat protein structure prediction in the present application, where the GPU card is NVIDIA TX3090), L checkpoint is determined, and T recompute is actually not calculated every time, but only once every L checkpoint layers to achieve the goal, which can further save GPU resources.

[0134] The actual calculation process is shown in the flowchart of Figure 4 In the present embodiment, by setting L checkpoint = 16 (store once every 3 layers), the memory usage is reduced to 33 GB (the upper limit of RTX3090), only 7.8% of the speed loss is introduced, so that the program that originally requires expensive large memory GPU cards (such as A100) can be efficiently run on consumer-level graphics cards.

[0135] As an example, the present application also makes some optimizations to the underlying CUDA code of AlphaFold2 to improve efficiency and further reduce the calculation workload of the GPU, such as:

[0136] - Adopt NVIDIA's Unified Memory management strategy to automatically migrate non-critical intermediate results to CPU memory;

[0137] - Design asynchronous computing flow: separate MSA processing from structure generation, and pre-load data during GPU idle period;

[0138] - Implement mixed precision training: use FP16 for Evoformer and maintain FP32 for structure modules to maintain geometric precision;

[0139] The present application is specifically designed for the calculation of wheat hexaploid, and AlphaFold2 is effectively improved, which has obvious innovation points. The main innovation points are:

[0140] 1. Field specificity:

[0141] The sub-genome weight formula (ωi) first integrates evolutionary conservation and expression;

[0142] Customize the interface constraint loss function according to the characteristics of wheat storage proteins;

[0143] 2. Technical depth:

[0144] The repeated domain bias attention mechanism is improved at three levels of MSA processing, model fine-tuning, and loss function;

[0145] The dynamic checkpoint algorithm realizes the Pareto optimization of memory-computing (31.4% memory saving and 92% speed maintenance);

[0146] 3. Application value:

[0147] The training data covers key breeding proteins (glutenin, gliadin, and disease-resistant proteins);

[0148] The memory optimization scheme enables single-card laboratory to handle whole-genome scale prediction (15,000 proteins <2000 GPU hours).

[0149] In this embodiment, by improving the fine-tuning process during model training, the trained model can better obtain the relevant knowledge of wheat protein structure, thereby improving the prediction accuracy of the model and reducing the computational amount of the model.

[0150] Reference Figure 5 , Figure 5 is a device structure schematic diagram of a hardware running environment related to the embodiment scheme of the present application.

[0151] As Figure 5As shown, the wheat protein structure prediction device can include a processor 1001, a memory 1005, a communication bus 1002. The communication bus 1002 is used to realize the connection communication between the processor 1001 and the memory 1005.

[0152] Optionally, the wheat protein structure prediction device can further include a user interface, a network interface, a camera, an RF (Radio Frequency) circuit, a sensor, a WiFi module, etc. The user interface can include a display, an input sub-module such as a keyboard, and the optional user interface can further include a standard wired interface, a wireless interface. The network interface can include a standard wired interface, a wireless interface (such as a WI-FI interface).

[0153] Those skilled in the art can understand that the wheat protein structure prediction device structure shown in the above embodiments is not a limitation on the wheat protein structure prediction device, and the wheat protein structure prediction device can include more or fewer components than the drawings, or combine certain components, or different component arrangements. Figure 5

[0154] As shown, the memory 1005 as a storage medium can include an operating system, a network communication module, and a wheat protein structure prediction program. The operating system is a program that manages and controls the hardware and software resources of the wheat protein structure prediction device, supports the operation of the wheat protein structure prediction program and other software and / or programs. The network communication module is used to realize the communication between the components in the memory 1005, and the communication between the other hardware and software in the wheat protein structure prediction system. Figure 5 In the wheat protein structure prediction device shown in the above embodiments, the processor 1001 is used to execute the wheat protein structure prediction program stored in the memory 1005, and realize the steps of the wheat protein structure prediction method described in any one of the above embodiments.

[0155] Figure 5 The specific embodiments of the wheat protein structure prediction device of the present application are basically the same as the above-described embodiments of the wheat protein structure prediction method, and will not be repeated here.

[0156] The specific embodiments of the wheat protein structure prediction device of the present application are basically the same as the above-described embodiments of the wheat protein structure prediction method, and will not be repeated here.

[0157] ​​It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0158] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0159] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, and of course, they can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) as described above, and includes a number of instructions for making a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) execute the methods described in various embodiments of the present application.

[0160] The above is only the preferred embodiment of the present application, and does not limit the application range of the present application. Any equivalent structure or equivalent process transformation made by using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, is also included in the application protection range of the present application.

[0161] It should be noted that: the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.

[0162] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for predicting wheat protein structure, characterized in that: The method comprises: Obtain wheat protein data to be processed; The wheat protein data is input into a preset wheat prediction model, and based on the preset wheat prediction model, the subgenomic sequence of the wheat protein data is predicted to obtain a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the subgenomic sequence and then fusing the separated subgenomic sequences according to a weight ratio.

2. The wheat protein structure prediction method according to claim 1, wherein The method of performing prediction processing on the subgenomic sequence of the wheat protein data based on the preset wheat prediction model to obtain a protein structure prediction result includes: Based on the preset wheat prediction model, the subgenomic sequences in the wheat protein data are separated and processed to obtain a plurality of subgene sequences; weighting the sub-gene sequences, and fusing the weighted sub-gene sequences to obtain a fused gene sequence; The fused gene sequence is subjected to prediction processing to obtain a protein structure prediction result.

3. The wheat protein structure prediction method according to claim 2, wherein: Based on the preset wheat prediction model, the subgenomic sequences in the wheat protein data are separated and processed to obtain multiple subgene sequences, including: Based on the preset wheat prediction model, extracting a preset number of subgenomic sequences from the wheat protein data; Obtaining homologous sequences from each of the subgenomic sequences; Based on the homologous sequences, multiple subgene sequences are determined.

4. The wheat protein structure prediction method according to claim 2, wherein: The weighting of each sub-gene sequence is performed, and the weighted sub-gene sequences are fused to obtain a fused gene sequence, including: Determining the sequence conservation type and gene expression level of the sub-gene sequence; Calculating the weight coefficient of each sub-gene sequence based on the sequence conservation type and the gene expression level; Each of the sub-gene sequences is fused according to the weight coefficient to obtain a fused gene sequence.

5. The wheat protein structure prediction method according to claim 1, wherein Before inputting the wheat protein data into a preset wheat prediction model, the method further comprises: Obtain protein sample data and gene sequence tag data; Based on the protein sample data and the gene sequence tag data, the wheat prediction model to be trained is iteratively trained until the prediction index corresponding to the wheat prediction model to be trained reaches a preset accuracy standard, thereby obtaining a trained wheat prediction model.

6. The wheat protein structure prediction method according to claim 5, wherein: The method of iteratively training the wheat prediction model to be trained based on the protein sample data and the gene sequence tag data until the prediction index corresponding to the wheat prediction model to be trained reaches a preset accuracy standard, thereby obtaining a trained wheat prediction model, comprising: Performing repeated detection on the protein sample data and the gene sequence label data through the self-attention mechanism in the wheat prediction model to be trained to obtain repeated regions, and calculating a bias term for the repeated regions; An inter-subunit distance constraint term is added to the loss function of the wheat prediction model to be trained, and the wheat prediction model to be trained is iteratively calculated multiple times through a preset strategy until the loss function converges, thereby obtaining a trained wheat prediction model.

7. The wheat protein structure prediction method according to claim 6, wherein: The preset strategy includes a calculation flow reorganization algorithm, and the wheat prediction model to be trained is subjected to multiple iterative calculations using the preset strategy, including: Determining activation values ​​of each network layer in the wheat prediction model to be trained based on the computational flow reorganization algorithm; For the network layer nodes whose activation value dependency is less than or equal to a preset threshold, discarding the activation value of the current network layer node; For the network layer nodes whose activation value dependence is greater than a preset threshold, the activation value of the current network layer node is stored for iterative training of the wheat prediction model.

8. A wheat protein structure prediction device, characterized in that: The wheat protein structure prediction device comprises: An acquisition module, used for acquiring wheat protein data to be processed; A processing module is used to input the wheat protein data into a preset wheat prediction model, and based on the preset wheat prediction model, perform prediction processing on the subgenomic sequence of the wheat protein data to obtain a protein structure prediction result, wherein the protein structure prediction result is obtained by separating the subgenomic sequence and then fusing the separated subgenomic sequences according to a weight ratio.

9. A wheat protein structure prediction device, characterized in that: The device includes: a memory, a processor, and a wheat protein structure prediction program stored in the memory and executable on the processor, wherein the wheat protein structure prediction program is configured to implement the steps of the wheat protein structure prediction method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that The computer storage medium stores a wheat protein structure prediction program, which, when executed by a processor, implements the steps of the wheat protein structure prediction method according to any one of claims 1 to 7.