Protein structure prediction method, electronic equipment, storage medium and program product

Through the proposed protein structure prediction method, the secondary structure prediction network is used to predict the protein secondary structure using the feature extraction network and the secondary structure prediction network, which solves the problems of low prediction efficiency and low accuracy in the prior art, and achieves efficient and accurate secondary structure prediction of protein.

CN119943156APending Publication Date: 2025-05-06FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020208.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-27
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is inefficient and has low accuracy in predicting secondary structures of proteins.

Method used

A protein structure prediction method is proposed. By obtaining the amino acid sequence to be tested and inputting it into the prediction model, the feature extraction network and the secondary structure prediction network are used to extract feature information and predict the secondary structure to reduce prediction errors.

Benefits of technology

It improves the efficiency and accuracy of protein secondary structure prediction, can quickly and accurately predict the secondary structure of protein, and reduces the dependence on protein sequence database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943156A_ABST
    Figure CN119943156A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a protein prediction method, electronic equipment, a storage medium and a program product, the method is applied to the electronic equipment, and the method comprises the following steps: obtaining a to-be-detected amino acid sequence; the to-be-detected amino acid sequence is input into a prediction model, and the prediction model is used for extracting feature information (such as at least one of evolutionary information, physical properties of amino acid, chemical properties of amino acid, a local pattern of the amino acid sequence or a global pattern of the amino acid sequence) of the to-be-detected amino acid sequence; and determining a protein secondary structure corresponding to the feature information according to an association relationship between the feature information and the protein secondary structure, thereby accurately determining the secondary structure of the protein and reducing prediction errors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application filed with the China Patent Office on November 27, 2024, with application number 202411728776.5 and application name “Protein structure prediction method, electronic device, storage medium and program product”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a protein structure prediction method, electronic equipment, storage medium and program product. Background Art

[0003] In recent years, high-throughput experimental methods have developed rapidly, resulting in the generation of a large number of new proteins. Protein function prediction has become a core issue in the field of molecular biology research.

[0004] The function of a protein is strongly related to its structure. Therefore, predicting the higher-order structure of a protein can help study how the protein performs its biological function and further study the interaction between the protein and other molecules.

[0005] Traditional protein function prediction methods can predict protein secondary structure based on molecular mechanics and molecular dynamics, but the process is time-consuming and expensive. Prediction methods that use statistics, such as statistical comparisons of protein databases, require querying a large amount of amino acid data in the database, which is also time-consuming. Summary of the invention

[0006] The embodiments of the present application provide a protein structure prediction method, electronic device, storage medium and program product, which solve the problems of low efficiency and low accuracy in predicting protein secondary structure.

[0007] In the first aspect, the present application proposes a protein structure prediction method for application to electronic devices, the method comprising: obtaining an amino acid sequence to be tested; inputting the amino acid sequence to be tested into a prediction model to obtain a protein secondary structure, wherein the prediction model is used to extract characteristic information of the amino acid sequence to be tested, and determine the protein secondary structure corresponding to the characteristic information based on the correlation between the characteristic information and the protein secondary structure.

[0008] It is understandable that the above-mentioned characteristic information may be, for example, evolutionary information, physical properties of amino acids, chemical properties of amino acids, local patterns of amino acid sequences, or global patterns of amino acid sequences, so as to accurately determine the secondary structure of the protein and reduce prediction errors.

[0009] In a possible implementation of the first aspect above, the prediction model includes a feature extraction network and a secondary structure prediction network, wherein: the feature extraction network is used to extract feature information of the amino acid sequence to be tested; the secondary structure prediction network is used to determine the protein secondary structure corresponding to the feature information based on the correlation between the feature information and the protein secondary structure.

[0010] It can be understood that the feature extraction network can be built based on a large language model architecture, for example, it can be a neural network based on a BERT (bidirectional encoder representations from transformers) architecture; the secondary structure prediction network can be built based on CNN and BiLSTM.

[0011] In a possible implementation of the first aspect above, the characteristic information includes at least one of evolutionary information, physical properties of amino acids, chemical properties of amino acids, local patterns of amino acid sequences, or global patterns of amino acid sequences.

[0012] It can be understood that the evolutionary information in amino acid sequences refers to the wide differences in the corresponding nucleic acid and protein sequence composition in different species. These differences are generated in the long-term evolution process and have relatively stable genetic characteristics. The evolutionary information in amino acid sequences can be used to understand the function, structure and biological evolution of proteins.

[0013] In a possible implementation of the first aspect above, the amino acid sequence to be tested is input into a prediction model to obtain a secondary structure of the protein, including: performing a feature extraction operation on the amino acid sequence to be tested through a feature extraction network to obtain feature information; and performing a classification operation on the feature information according to the correlation between the feature information and the secondary structure of the protein through a secondary structure prediction network to obtain a label characterizing the secondary structure of the protein.

[0014] It can be understood that the electronic device (such as the electronic device 1500 proposed below) can extract the corresponding features in the amino acid sequence to be tested based on the above prediction model, and match the feature information (such as evolutionary information) in the pre-trained amino acid sequence in the model based on the feature, and then determine the matched feature information, and determine the classification label corresponding to the protein secondary structure according to the correlation between the feature information and the protein secondary structure, and output the classification label. Among them, the amino acid sequence may include the evolutionary information in the amino acid sequence such as the physicochemical properties of the amino acids, the local and global patterns in the amino acid sequence, etc.

[0015] In a possible implementation of the first aspect, inputting the amino acid sequence to be tested into the prediction model includes: during the input of the amino acid sequence to be tested, randomly performing a masking operation on a preset proportion of amino acids.

[0016] Exemplarily, the preset ratio may be 15%.

[0017] In a possible implementation of the first aspect above, a feature extraction operation is performed on the amino acid sequence to be tested through a feature extraction network to obtain feature information, including: encoding the amino acid sequence to be tested into a one-hot vector based on relative position encoding through the feature extraction network, and extracting feature information based on the one-hot vector.

[0018] It can be understood that extracting feature information based on one-hot vectors is conducive to subsequent feature extraction processing, simplifies the content of the sequence to be processed, and can improve efficiency.

[0019] In a possible implementation of the first aspect above, a method for training a feature extraction network includes: obtaining a feature extraction network to be trained; selecting a first training set for a preset number of rounds of training, setting an initial learning rate, and performing verification on a randomly divided verification machine after each training round, and stopping the training of the feature extraction network when a preset accuracy condition is met to obtain a trained feature extraction network.

[0020] For example, the electronic device 1500 can select the data set UniRef50 containing a large number of protein sequences as the first training set, and train for about 25 rounds in total, with an initial learning rate of 1e-3. After each round, verification is performed on a randomly divided verification machine. In this way, the feature extraction network can learn accurate amino acid sequence evolution information, so that the electronic device 1500 can determine the feature information (such as evolution information) in the amino acid sequence based on the prediction model, thereby replacing the cumbersome database evolution information query operation.

[0021] In a possible implementation of the first aspect above, when a preset accuracy condition is met, training of the feature extraction network is stopped, including: when a first accuracy condition is met, dividing the current learning rate by 2 as the learning rate for the next round of training, wherein the first accuracy condition indicates a decrease in accuracy; when a second accuracy condition is met, training is stopped, wherein the second accuracy condition indicates that the total number of decreases in accuracy reaches a preset threshold.

[0022] It can be understood that the preset threshold of the total number of times the accuracy rate drops can be 4 times, then the training is stopped, so that the feature extraction network can learn accurate amino acid sequence evolution information.

[0023] In a possible implementation of the first aspect above, a training method for a secondary structure prediction network includes: obtaining a secondary structure prediction network to be trained; cleaning the secondary structures of the proteins in a second training set to obtain a cleaned second training set, wherein the cleaning includes removing invalid structures or redundant structures; obtaining secondary structure labels corresponding to the secondary structures of the proteins in the cleaned second training set using a preset algorithm; and training the secondary structure prediction network to be trained based on the secondary structure labels corresponding to the secondary structures of the proteins in the cleaned second training set.

[0024] For example, the electronic device 1500 can select a protein structure database, such as data in a protein data bank (PDB) as a second training set, and screen the secondary structures of the proteins in the second training set to remove invalid structures or redundant structures to obtain a screened second training set; then use a DSSP (dictionary of secondary structure of proteins) algorithm to obtain the secondary structure labels of the screened second training set.

[0025] In a possible implementation of the first aspect above, the training method of the secondary structure prediction network also includes: inputting an amino acid sequence whose secondary structure has not been determined into the secondary structure prediction network to be trained; storing the labels whose confidence is greater than the confidence threshold among the secondary structure labels predicted by the secondary structure prediction network to be trained as soft labels, and the soft labels are used to train the secondary structure prediction network to be trained, wherein the weight of the soft labels is less than the weight of the secondary structure labels of the second training set.

[0026] It can be understood that since the training set of the secondary structure prediction network depends on the protein sequence database, but the protein sequence database has only about 190,000 entries, and most of the amino acid sequences in the database have no experimentally determined structural data, it is possible to predict the protein secondary structure for the amino acid sequences whose secondary structures have not been determined. At this time, the electronic device 1500 can store the labels with higher confidence in the secondary structures predicted based on the amino acid sequences as soft labels, and add them back to the training set to train the secondary structure prediction network.

[0027] Exemplarily, soft labeling refers to setting the sample weight of this type of training sample label to be smaller than the training sample label weight of the known protein secondary structure. For example, if the training sample label weight of the known protein secondary structure is 1, then the amino acid sequence training sample label weight for which the secondary structure data has not been measured but has a higher confidence after testing can be 0.5, which facilitates the electronic device 1500 to perform a new round of optimization of the prediction model.

[0028] It is understandable that the electronic device 1500 can set a confidence threshold to facilitate the determination of soft tags, for example, storing tags with a confidence greater than 0.85 as soft tags.

[0029] In a possible implementation of the first aspect above, the method also includes: setting different random initialization parameters for the secondary structure prediction network to be trained, performing multiple trainings, and obtaining multiple different trained secondary structure prediction networks; and taking the average of the prediction results of the multiple trained secondary structure prediction networks as the prediction result of the trained secondary structure prediction network to determine the secondary structure of the protein.

[0030] It can be understood that the average of the prediction results of multiple trained secondary structure prediction networks is used as the prediction result of the trained secondary structure prediction network, which can effectively improve the prediction accuracy.

[0031] In a second aspect, the present application also provides an electronic device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the protein structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0032] In a third aspect, the present application further provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, causes the computer to execute the protein structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0033] In a fourth aspect, an embodiment of the present application discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the protein structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0034] The beneficial effects of the second to fourth aspects mentioned above can be referred to the relevant descriptions in the first aspect mentioned above and various possible implementations of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1A A schematic diagram showing a scenario of predicting protein structure;

[0036] Figure 1B A schematic diagram showing another scenario for predicting protein structure;

[0037] Figure 2 A schematic diagram showing a specific implementation process of a protein structure prediction method provided according to some embodiments of the present application is shown;

[0038] Figure 3A schematic diagram of a prediction model framework structure provided according to some embodiments of the present application is shown;

[0039] Figure 4 A schematic structural diagram of an electronic device provided according to some embodiments of the present application is shown. DETAILED DESCRIPTION

[0040] In order to facilitate understanding of the technical solutions provided by the embodiments of the present application, the meanings of some related field terms involved in the embodiments of the present application are explained below.

[0041] (1) Protein secondary structure refers to the specific conformation formed by the backbone atoms of the polypeptide main chain circling or folding along a certain axis, that is, the spatial arrangement of the backbone atoms of the peptide chain, without involving the side chains of amino acid residues. The main forms of protein secondary structure include α-helix, β-fold, β-turn and random coil. Due to the large molecular weight of protein, different peptide segments of a protein molecule can contain different forms of secondary structure. The main force maintaining the secondary structure is hydrogen bond. The secondary structure of a protein is not a simple α-helix or β-fold structure, but a combination of these different types of conformations.

[0042] (2) Relative position encoding mechanism: This is an encoding method used in the Transformer model and its derivative models to capture the relationship between different positions in a sequence. It allows the model to understand the relative distance between words or tokens, which is very important for understanding sentence structure and semantics.

[0043] (3) One-hot encoding: Also known as one-hot encoding, it is a method of converting categorical or nominal variables into a form that can be better processed by machine learning algorithms. In this encoding, each category value is represented as a binary vector, with all positions except one that is 1, indicating the category, and all other positions are 0.

[0044] (4) AdamW optimizer: It is a variant of the Adam optimizer (adaptive moment estimation). It adds weight decay to Adam to solve the problem of model overfitting. AdamW optimizer has faster convergence speed and robustness when training deep learning models.

[0045] (5) Position-specific scoring matrix (PSSM) is a commonly used sequence feature representation method in bioinformatics. It scores each amino acid site of the protein by analyzing the conservation of homologous protein sequences, thereby predicting the structure and function of the protein.

[0046] (6) HHM (HHblits-based feature) is a feature generated by the HHblits tool and is used for protein structure and function prediction. HHblits is a fast and accurate sequence alignment tool based on the hidden Markov model (HMM). In protein prediction, the HHM feature contains the sequence alignment results obtained through the HHblits search, which can reflect the homology and conservation information between protein sequences.

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions provided by the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0048] It can be understood that the electronic device 1500 in the embodiment of the present application can be a terminal device, which can also be called a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The terminal device can be a mobile phone, a smart TV, a wearable device, a tablet computer (pad), a computer with a wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a vehicle-mounted terminal, etc.

[0049] The following uses the electronic device 1500 as the terminal 100 as an example, and describes in further detail the protein structure prediction method provided in the embodiment of the present application in combination with relevant drawings.

[0050] Figure 1A A schematic diagram showing a scenario for predicting protein structure. Figure 1B Schematic diagram showing another scenario for predicting protein structure.

[0051] refer to Figure 1A The algorithm set in the terminal 100 reads the primary structure of the protein input by the user (hereinafter referred to as the amino acid sequence to be tested), and can return the predicted secondary structure of the protein to the user.

[0052] As mentioned above, the reliability and efficiency of terminal 100 in predicting protein secondary structure are relatively low.

[0053] For example, refer to Figure 1B , the terminal 100 reads the amino acid sequence to be tested and can send a query request to the protein sequence database 01. The protein sequence database 01 can be stored in the storage environment 201 provided by the server 200, and a large amount of protein primary structure (i.e., the amino acid sequence of the protein) and the corresponding secondary structure data are pre-stored in the protein sequence database 01. The terminal 100 can find out the protein secondary structure corresponding to the amino acid sequence to be tested by comparing the amino acid sequence to be tested with the amino acid sequence in the protein sequence database 01, and return the queried protein secondary structure to the user.

[0054] However, as the number of protein types continues to increase, many amino acid sequences in the protein sequence database 01 have not yet determined their secondary structures. Therefore, it is not necessarily possible to find the secondary structure corresponding to the amino acid sequence to be tested in the protein sequence database 01. Figure 1B The method shown has low robustness in predicting protein secondary structure. Moreover, corresponding to a large number of protein secondary structure prediction scenarios, a large number of amino acid sequences to be tested can generate high-concurrency query requests. Limited by the hardware performance of the terminal 100, the design and architecture of the protein sequence database 01, and the number of data connections between the two, the query efficiency of this method is relatively low.

[0055] In order to solve the problems of low efficiency and low accuracy in predicting protein secondary structure, an embodiment of the present application provides a protein prediction method, which is applied to electronic equipment, including: obtaining an amino acid sequence to be tested; inputting the amino acid sequence to be tested into a prediction model, the prediction model is used to extract characteristic information of the amino acid sequence to be tested (for example, evolutionary information, physical properties of amino acids, chemical properties of amino acids, local patterns of amino acid sequences, or at least one of global patterns of amino acid sequences), and determining the protein secondary structure corresponding to the characteristic information based on the correlation between the characteristic information and the protein secondary structure, thereby accurately determining the secondary structure of the protein and reducing prediction errors.

[0056] It can be understood that the prediction model can be a deep learning algorithm including a multi-layer neural network, which can continuously learn the characteristic information in the amino acid sequence (such as evolutionary information, physical properties of amino acids, chemical properties of amino acids, local patterns of amino acid sequences or at least one of the global patterns of amino acid sequences) through the existing amino acid sequence and the corresponding protein secondary structure, thereby replacing the cumbersome database evolution information query operation, saving time and effort.

[0057] The evolutionary information in amino acid sequences refers to the wide differences in the composition of corresponding nucleic acid and protein sequences in different species. These differences are generated during the long-term evolution and have relatively stable genetic characteristics. The evolutionary information in amino acid sequences can be used to understand the function, structure and biological evolution of proteins.

[0058] In some embodiments of the present application, in order to facilitate learning of feature information in an amino acid sequence, the above-mentioned prediction model may have a large language model architecture. The electronic device 1500 may use the amino acid sequence as a sentence based on the prediction model to analyze the association between the amino acid sequence and the secondary structure of the corresponding protein, so that the prediction model can learn to map the amino acid sequence to the secondary structure of the corresponding protein. For example, a large number of known amino acid sequences and corresponding known secondary structures of proteins may be used to train the prediction model, so that it can deeply learn the sequence content and evolutionary information in the amino acid sequence, and the association between the secondary structure of the protein, and then the electronic device 1500 can accurately predict the secondary structure of the protein with a high confidence based on the amino acid sequence to be tested and the evolutionary information based on the prediction model. Since the training is performed using amino acid sequences with known protein secondary structures, the prediction model can be applied to different proteins, for example, it can be applied to newly discovered amino acid sequences whose secondary structures have not been determined in the laboratory, and it has high robustness. Moreover, the prediction model can be applied to high-throughput amino acid sequence prediction scenarios based on the large language model architecture, and even if a large number of amino acid sequences to be tested are input, the prediction of the secondary structure can still be completed quickly, compared to the above. Figure 1B The traditional method of querying the protein sequence database 01 in the example effectively improves the prediction efficiency.

[0059] Combine the following Figure 2 The specific implementation process and principle of predicting protein structure provided in the examples of the present application are described in detail.

[0060] Figure 2 A schematic diagram of a specific implementation process of a protein structure prediction method provided according to some embodiments of the present application is shown.

[0061] Here, Figure 2 The execution subject of the specific implementation process shown may be the electronic device 1500, which will not be elaborated here.

[0062] refer to Figure 2 , the specific implementation process includes the following steps:

[0063] S201, obtaining an amino acid sequence to be tested.

[0064] It can be understood that the amino acid sequence to be tested is the amino acid sequence of the protein to be tested, which is the primary structure of the protein to be tested.

[0065] In some embodiments, the electronic device 1500 can obtain the amino acid sequence to be tested input by the user through an input control, and the input control can be, for example, an input box control.

[0066] S202, inputting the amino acid sequence to be tested into the prediction model, wherein the prediction model is used to extract characteristic information of the amino acid sequence to be tested, and determine the protein secondary structure corresponding to the characteristic information according to the correlation between the characteristic information and the protein secondary structure.

[0067] It can be understood that the electronic device 1500 can extract the corresponding features in the amino acid sequence to be tested based on the above prediction model, and match the feature information (such as evolutionary information) in the pre-trained amino acid sequence in the model based on the feature, and then determine the matched feature information, and determine the classification label corresponding to the secondary structure of the protein according to the association between the feature information and the secondary structure of the protein, and output the classification label. Among them, the amino acid sequence may include the physicochemical properties of amino acids, local and global patterns in the amino acid sequence, and other evolutionary information in the amino acid sequence.

[0068] It can be understood that since the electronic device 1500 has learned the relationship between the amino acid sequence, the amino acid sequence evolution information and the secondary structure of the protein corresponding to the amino acid sequence based on the prediction model, the electronic device 1500 can determine the relationship between the amino acid sequence and the secondary structure of the protein without searching the protein sequence database 01. While traditional PSSM, HHM and other information require a protein sequence database query, which is time-consuming. Since the embodiment of the present application does not require a protein sequence database query, it is conducive to large-scale protein secondary structure prediction. The calculation speed of the protein structure prediction method proposed in the present application can exceed the traditional method by more than 2 orders of magnitude, which is conducive to high-throughput data calculation.

[0069] S203, obtain the protein secondary structure.

[0070] It can be understood that the electronic device 1500 can determine the secondary structure of the protein based on the classification label query output by the prediction model.

[0071] Here, the electronic device 1500 can obtain the classification label output by the prediction model and can determine the secondary structure of the protein based on the classification label.

[0072] There are three basic types of protein secondary structures: α-helix, β-fold and random coil. The α-helix structure is a typical structural form of the protein main chain, which exists in both fibrous proteins and globular proteins. β-fold is two or more almost fully extended polypeptide chains gathered together laterally, connected to a lamellar structure by interchain hydrogen bonds. Random coil is some irregular conformations of the peptide plane in the polypeptide chain that are arranged irregularly. Its flexible conformation can change the direction of the peptide chain, which is conducive to connecting the relatively rigid α-helix and β-fold. Thus, the electronic device 1500 can obtain the first label corresponding to the α-helix, the second label corresponding to the β-fold and / or the third label corresponding to the random coil based on the prediction model.

[0073] According to the protein structure prediction method shown in the above steps S201 to S203, the embodiment of the present application provides a prediction model based on a large language model architecture and learned with amino acid sequence evolution information, so that the electronic device 1500 can efficiently and accurately predict the secondary structure of the protein based on the prediction model.

[0074] It can be understood that the prediction model proposed in the embodiment of the present application is designed for the prediction scenario of protein secondary structure, and can capture the characteristics of protein secondary structure more finely. Compared with existing prediction tools, such as AlphaFold2, it can have a higher accuracy in the field of secondary structure prediction, thus providing a basis for further improving the performance of structure prediction algorithms.

[0075] In some embodiments, in order to improve the reliability and practicality of the prediction results, the electronic device 1500 can also perform post-processing such as smoothing, noise removal, and error correction on the prediction results to obtain more accurate and stable secondary structure prediction results. At the same time, it can also be combined with other bioinformatics methods to further mine the structural and functional information of proteins.

[0076] In some embodiments of the present application, the above-mentioned prediction model can be a deep learning algorithm including a multi-layer neural network, and the neural network of the deep learning algorithm can include: feedforward neural networks (FNN), in which information flows only in one direction, from the input layer to the hidden layer, and then to the output layer; convolutional neural networks (CNN), which can extract local features through the convolution layer and reduce the spatial dimension of the features through the pooling layer; recurrent neural networks (RNN), which are used to process sequence data, such as time series or natural language, and can process the time dependency between input data; long short-term memory network (LSTM), a special RNN that can learn long-term dependency information and solve the gradient vanishing problem of traditional RNN; gated recurrent unit (GRU), which is similar to LSTM, but has a simpler structure and fewer parameters, and can sometimes be used as a substitute for LSTM; bidirectional long short-term memory network (bidirectional long short-term memory network, LSTM), which is similar to LSTM, but has a simpler structure and fewer parameters, and can sometimes be used as a substitute for LSTM; BiLSTM is a special recurrent neural network (RNN) that can process sequence data and maintain long-term memory. BiLSTM runs two LSTMs on the time series simultaneously, one from front to back and the other from back to front, so that it can consider both past and future information at the same time. This structure enables BiLSTM to better capture the contextual relationship in the sequence data. BiLSTM is widely used in natural language processing tasks, such as sentiment analysis, machine translation, text generation and question-answering systems. It can capture the temporal dependency between the source language and the target language, as well as the contextual information of the dialogue or text in the dialogue system or news summary. Generative adversarial networks (GAN), composed of generators and discriminators, are used to generate new data samples, such as images, music, etc. Variational autoencoders (VAE), used to generate new data samples and learn the potential representation of the input data at the same time; Transformer, a model based on the self-attention mechanism, is widely used in natural language processing tasks; deep residual networks (ResNet), solve the problem of difficult training of deep networks by introducing residual learning.

[0077] In some embodiments of the present application, the neural network architecture of the above prediction model can use a large language model as the core architecture. For example, the prediction model can include a neural network based on the BERT (bidirectional encoder representations from transformers) architecture, which includes 33 bidirectional Transformer layers.

[0078] Exemplarily, in this prediction model, a relative position encoding mechanism can be introduced to improve the input layer. Relative position encoding is a method of encoding a sequence according to the relative relationship between positions. Relative position encoding takes into account the relative distance and relationship between different positions in the sequence, and uses learnable parameters to model these relationships. Relative position encoding can capture the relative information between positions by calculating the offset or relative position difference between different positions. Compared with absolute position encoding, relative position encoding pays more attention to the relative order and distance between positions in the sequence, and it can better handle position information in long sequences. Furthermore, after the amino acid sequence to be tested is input into the input layer, the amino acid sequence to be tested can be encoded into a one-hot vector through relative position encoding, which is convenient for subsequent feature extraction processing. The one-hot vector can be used to extract features through 33 layers of Transformer layers to obtain a feature vector of 1280 dimensions.

[0079] In some embodiments of the present application, during the input of the amino acid sequence to be tested, the prediction model can also randomly perform a mask operation on a preset proportion (e.g., 15%) of amino acids, and the output layer of the prediction model can predict the amino acid type at the masked position.

[0080] In some embodiments of the present application, the electronic device 1500 can use the AdamW optimizer in combination with a large-scale protein sequence database to train the prediction model. The loss function applied in the training process can be a cross-entropy loss function, so that the prediction model can accurately capture the key information in the amino acid sequence, and then accurately predict the secondary structure of the protein.

[0081] The specific training process of the prediction model is described in detail below with reference to the relevant drawings.

[0082] In some embodiments, reference Figure 3The above prediction model may include a feature extraction network and a secondary structure prediction network. The electronic device 1500 may input the amino acid sequence to be tested into the feature extraction network to determine the feature vector, and input the obtained feature vector into the trained secondary structure prediction network to determine the secondary structure of the protein. The feature extraction network is used to extract the feature information of the amino acid sequence to be tested, and the secondary structure prediction network is used to determine the secondary structure of the protein corresponding to the feature information according to the correlation between the feature information and the secondary structure of the protein.

[0083] Exemplarily, the electronic device 1500 may perform feature extraction operations on the amino acid sequence to be tested through a feature extraction network to obtain feature information; and perform classification operations on the feature information according to the correlation between the feature information and the secondary structure of the protein through a secondary structure prediction network to obtain a label characterizing the secondary structure of the protein.

[0084] Exemplarily, the feature extraction network can be constructed based on a large language model architecture, for example, it can be a neural network based on a BERT (bidirectional encoder representations from transformers) architecture.

[0085] Exemplarily, the secondary structure prediction network can be constructed based on CNN and BiLSTM.

[0086] In some embodiments of the present application, the above-mentioned feature extraction operation may include: encoding the amino acid sequence to be tested into a one-hot vector based on relative position encoding through a feature extraction network, and extracting feature information according to the one-hot vector.

[0087] In some embodiments, the electronic device 1500 can train the feature extraction network of the prediction model in the following manner: obtain the feature extraction network to be trained, select the first training set for a preset number of rounds of training, set the initial learning rate, and perform verification on a randomly divided verification machine for each training round. When the preset accuracy condition is met, stop training the feature extraction network to obtain a trained feature extraction network. For example, the electronic device 1500 can select a data set UniRef50 containing a large number of protein sequences as the first training set, train for about 25 rounds in total, and the initial learning rate is 1e-3. After each round, verify on a randomly divided verification machine. Corresponding to the case where the first accuracy condition is met, for example, the accuracy rate decreases, the current learning rate is divided by 2 as the learning rate for the next round of training. Corresponding to the case where the second accuracy condition is met, for example, the total number of times the accuracy rate decreases reaches a preset threshold, for example, a total decrease of 4 times, then stop training. In this way, the feature extraction network can learn accurate amino acid sequence evolution information, so that the electronic device 1500 can determine the feature information (such as evolution information) in the amino acid sequence based on the prediction model, thereby replacing the cumbersome database evolution information query operation.

[0088] In some embodiments, the electronic device 1500 can train the secondary structure prediction network of the prediction model in the following manner: obtain the secondary structure prediction network to be trained; perform a cleaning process on the secondary structures of the proteins in the second training set to obtain a cleaned second training set, wherein the cleaning process includes removing invalid structures or redundant structures; obtain the secondary structure labels corresponding to the secondary structures of the proteins in the cleaned second training set using a preset algorithm; and train the secondary structure prediction network to be trained based on the secondary structure labels corresponding to the secondary structures of the proteins in the cleaned second training set. For example, the electronic device 1500 can select data in a protein structure database PDB as the second training set, and screen the secondary structures of the proteins in the second training set, remove invalid structures or redundant structures, and obtain a screened second training set; and then use the DSSP (dictionary of secondary structure of proteins) algorithm to obtain the secondary structure labels of the screened second training set.

[0089] Exemplarily, the amino acid sequence is input into a trained feature extraction network to obtain the features (L, 1280) of each residue in the amino acid sequence. Here, it is assumed that the sequence length is L and the feature of each residue is 1280. Then it is input into the secondary structure prediction network. In the secondary structure prediction network, the input feature vector is subjected to feature extraction by a multi-layer deep convolutional neural network (CNN). CNN can automatically learn local features in protein sequences and convert them into higher-level feature representations through convolution operations. The extracted features are sequence modeled by a recurrent neural network (BiLSTM) with an attention mechanism. BiLSTM can consider the dependency between adjacent sequential elements in the features extracted by CNN, thereby improving the accuracy of the prediction, and is used for further sequence modeling and classification prediction of the extracted features.

[0090] In other embodiments, the electronic device 1500 can also train the secondary structure prediction network of the prediction model in the following manner: set different random initialization parameters for the secondary structure prediction network to be trained, perform multiple trainings, obtain multiple different result values ​​predicted by the trained secondary structure prediction networks, and average the different result values ​​output by the trained secondary structure prediction networks as the output result of the trained secondary structure prediction network. For example, a total of 3 trainings are performed to obtain the prediction results of 3 different trained secondary structure prediction networks, and the average of the prediction results of the three trained secondary structure prediction networks is used as the final result, that is, as the prediction result of the trained secondary structure prediction network, to determine the secondary structure of the protein, which can effectively improve the prediction accuracy.

[0091] Here, since the training set of the secondary structure prediction network depends on the protein sequence database, but the protein sequence database has only about 190,000 entries, and most of the amino acid sequences in the library have no experimentally determined structural data, the protein secondary structure prediction can be performed on the amino acid sequence whose secondary structure has not been determined. At this time, the electronic device 1500 can save the label with a higher confidence in the secondary structure predicted according to this type of amino acid sequence as a soft label, and add it back to the training set to train the secondary structure prediction network. The electronic device 1500 can input the amino acid sequence whose secondary structure has not been determined into the secondary structure prediction network to be trained; the label with a confidence greater than the confidence threshold in the secondary structure label predicted by the secondary structure prediction network to be trained is saved as a soft label, and the soft label is used to train the secondary structure prediction network to be trained, wherein the weight of the soft label is less than the weight of the secondary structure label of the second training set. Exemplarily, soft labeling refers to setting the sample weight of this type of training sample label to be smaller than the training sample label weight of the known protein secondary structure. For example, if the training sample label weight of the known protein secondary structure is 1, then the amino acid sequence training sample label weight for which the secondary structure data has not been measured but has a higher confidence after testing can be 0.5, which facilitates the electronic device 1500 to perform a new round of optimization of the prediction model.

[0092] It is understandable that the electronic device 1500 can set a confidence threshold to facilitate the determination of soft tags, for example, storing tags with a confidence greater than 0.85 as soft tags.

[0093] In some embodiments, the electronic device 1500 inputs the amino acid sequence to be tested into the trained prediction model, and can obtain the feature vector (L, 1280) of each residue in the sequence, where it is assumed that the sequence length is L and the feature of each residue is 1280. Then it is input into the secondary structure prediction network. The electronic device 1500 extracts features from the input feature vector based on the multi-layer deep convolutional neural network (CNN) in the prediction model. Here, CNN can automatically learn the local features in the protein sequence and convert them into higher-level feature representations through convolution operations. The extracted features are sequence modeled through a recurrent neural network with an attention mechanism, such as a bidirectional long short-term memory network (BiLSTM). The bidirectional long short-term memory network can take into account the dependencies between adjacent elements in the sequence, thereby improving the accuracy of the prediction, and is used to further perform sequence modeling and classification prediction on the extracted features to determine a more accurate protein secondary structure. For example, the electronic device 1500 can perform feature extraction on the feature vector (L, 1280) based on a multi-layer deep convolutional neural network (CNN) and a recurrent neural network with an attention mechanism (such as BiLSTM) to implement multiple strategies to improve prediction accuracy, such as ensemble learning, knowledge distillation, etc. These strategies can make full use of the advantages of different models, reduce prediction errors, and ultimately predict one of the three corresponding protein secondary structures for each amino acid (L, 3).

[0094] According to the protein structure prediction method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, when the computer program code is executed on a computer, the computer implements the steps performed by the electronic device 1500 in any one of the above embodiments.

[0095] According to the protein structure prediction method provided in the embodiments of the present application, the present application also provides a computer-readable medium, which stores a program code. When the program code is executed on a computer, the computer implements the steps performed by the electronic device 1500 in any one of the above embodiments.

[0096] Figure 4 A schematic structural diagram of an electronic device 1500 provided according to some embodiments of the present application is shown.

[0097] like Figure 4As shown, the electronic device 1500 includes one or more processors 1501, a system memory 1502, a non-volatile memory (NVM) 1503, a communication interface 1504, an input / output (I / O) device 1505, and a system control logic 1506 for coupling the processor 1501, the system memory 1502, the non-volatile memory 1503, the communication interface 1504 and the input / output (I / O) device 1505. Among them:

[0098] The processor 1501 may include one or more processing units, for example, a data processing unit or processing circuit that may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor or a programmable logic device (field programmable gate array, FPGA), a neural network processor (neural-network processing unit, NPU), etc. may include one or more single-core or multi-core processors. In some embodiments, the processor 1501 can be used to execute instructions to implement the above-mentioned protein structure prediction method.

[0099] The system memory 1502 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 1502 can be used to store instructions, and can also be used to store original data objects and changed data objects.

[0100] The non-volatile memory 1503 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 1503 may include any suitable non-volatile memory such as a flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 1503 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 1503 may be used to store instructions, and may also be used to store original data objects and changed data objects.

[0101] In some embodiments, the system memory 1502 and the non-volatile memory 1503 may respectively include: a temporary copy and a permanent copy of the instruction 1507. The instruction 1507 may include: when executed by at least one of the processors 1501, the electronic device 1500 implements the protein structure prediction method provided in various embodiments of the present application.

[0102] The communication interface 1504 may include a transceiver for providing a wired or wireless communication interface for the electronic device 1500, so as to communicate with any other suitable device through one or more networks. In some embodiments, the communication interface 1504 may be integrated into other components of the electronic device 1500, for example, the communication interface 1504 may be integrated into the processor 1501. In some embodiments, the electronic device 1500 may communicate with other devices through the communication interface 1504, for example, the electronic device 1500 may establish a communication connection with other devices through the communication interface 1504, so as to send data change requests to other devices, obtain original data objects, and send changed data objects through the communication connection.

[0103] Input / output (I / O) device 1505 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. A user may interact with electronic device 1500 via input / output (I / O) device 1505. For example, business personnel may input / select content for data change via input / output (I / O) device 1505.

[0104] The system control logic 1506 may include any suitable interface controller to provide any suitable interface with other modules of the electronic device 1500. For example, in some embodiments, the system control logic 1506 may include one or more memory controllers to provide interfaces to the system memory 1502 and the non-volatile memory 1503.

[0105] In some embodiments, at least one of the processors 1501 may be packaged together with the logic of one or more controllers for the system control logic 1506 to form a system in package (SiP). In other embodiments, at least one of the processors 1501 may also be integrated with the logic of one or more controllers for the system control logic 1506 on the same chip to form a system-on-chip (SoC).

[0106] Understandably, Figure 4 The structure of the electronic device 1500 shown is only an example. In other embodiments, the electronic device 1500 may include more or fewer components than shown, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0107] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer module or module code executed on a programmable system, and the programmable system includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.

[0108] The module code may be applied to input instructions to perform the functions described herein and generate output information. The output information may be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0109] Module code can be implemented with high-level modular language or object-oriented programming language to communicate with the processing system. When necessary, module code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0110] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).

[0111] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0112] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.

[0113] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.

[0114] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0115] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0116] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.

[0117] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0118] References to "some embodiments" or "embodiments" in the specification mean that the specific features, structures, or characteristics described in conjunction with the embodiments are included in at least one exemplary implementation or technology disclosed according to the embodiments of the present application. The appearance of the phrase "in some embodiments" in various places in the specification does not necessarily all refer to the same embodiment.

[0119] In addition, the language used in this specification has been primarily selected for readability and instructional purposes and may not be selected to describe or limit the disclosed subject matter. Therefore, the present application embodiment disclosure is intended to illustrate rather than limit the scope of the concepts discussed herein.

Claims

1. A protein structure prediction method, characterized in that: Applied to electronic equipment, the method comprises: Obtaining an amino acid sequence to be tested; The amino acid sequence to be tested is input into the prediction model to obtain the protein secondary structure, The prediction model is used to extract the characteristic information of the amino acid sequence to be tested, and to determine the protein secondary structure corresponding to the characteristic information based on the correlation between the characteristic information and the protein secondary structure.

2. The method according to claim 1, characterized in that The prediction model includes a feature extraction network and a secondary structure prediction network, wherein: The feature extraction network is used to extract feature information of the amino acid sequence to be tested; The secondary structure prediction network is used to determine the protein secondary structure corresponding to the feature information according to the association relationship between the feature information and the protein secondary structure.

3. The method according to claim 1, characterized in that The characteristic information includes at least one of evolutionary information, physical properties of amino acids, chemical properties of amino acids, local patterns of amino acid sequences, or global patterns of amino acid sequences.

4. The method according to claim 2, characterized in that: The step of inputting the amino acid sequence to be tested into a prediction model to obtain a protein secondary structure comprises: Performing a feature extraction operation on the amino acid sequence to be tested through the feature extraction network to obtain feature information; The feature information is classified according to the association between the feature information and the secondary structure of the protein through a secondary structure prediction network to obtain a label representing the secondary structure of the protein.

5. The method according to claim 1, characterized in that The step of inputting the amino acid sequence to be tested into a prediction model comprises: During the input of the amino acid sequence to be tested, a masking operation is randomly performed on a preset proportion of amino acids.

6. The method according to claim 4, characterized in that The step of performing a feature extraction operation on the amino acid sequence to be tested by the feature extraction network to obtain feature information includes: The amino acid sequence to be tested is encoded into a one-hot vector based on relative position encoding through the feature extraction network, and the feature information is extracted according to the one-hot vector.

7. The method according to claim 2, characterized in that The training method of the feature extraction network includes: Obtain the feature extraction network to be trained; Select the first training set for a preset number of rounds of training and set the initial learning rate. Each round of training is verified on a randomly divided verification machine. When a preset accuracy condition is met, the training of the feature extraction network is stopped to obtain a trained feature extraction network.

8. The method according to claim 7, characterized in that When the preset accuracy condition is met, stopping the training of the feature extraction network includes: When the first accuracy condition is met, the current learning rate is divided by 2 as the learning rate for the next round of training, wherein the first accuracy condition indicates that the accuracy rate has decreased; When a second accuracy condition is met, the training is stopped, wherein the second accuracy condition indicates that the total number of times the accuracy drops reaches a preset threshold.

9. The method according to claim 2, characterized in that: The training method of the secondary structure prediction network includes: Obtaining a secondary structure prediction network to be trained; Performing a cleaning process on the protein secondary structures in the second training set to obtain a cleaned second training set, wherein the cleaning process includes removing invalid structures or redundant structures; Obtaining secondary structure labels corresponding to the secondary structures of the proteins in the cleaned second training set using a preset algorithm; The secondary structure prediction network to be trained is trained based on the secondary structure labels corresponding to the protein secondary structures in the cleaned second training set.

10. The method according to claim 9, characterized in that The training method of the secondary structure prediction network also includes: Inputting the amino acid sequence whose secondary structure has not been determined into the secondary structure prediction network to be trained; Among the secondary structure labels predicted by the secondary structure prediction network to be trained, the labels with confidence greater than the confidence threshold are stored as soft labels, and the soft labels are used to train the secondary structure prediction network to be trained. The weight of the soft label is less than the weight of the secondary structure label of the second training set.

11. The method according to claim 2, characterized in that The method further comprises: Different random initialization parameters are set for the secondary structure prediction network to be trained, and multiple trainings are performed to obtain multiple different trained secondary structure prediction networks; The average of the prediction results of the multiple trained secondary structure prediction networks is used as the prediction result of the trained secondary structure prediction network to determine the secondary structure of the protein.

12. An electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the protein structure prediction method according to any one of claims 1 to 11.

13. A computer readable medium, characterized in that The computer-readable medium stores instructions, which, when executed on a machine, enable the machine to perform the protein structure prediction method according to any one of claims 1 to 11.

14. A computer program product, characterized in that The invention comprises a computer program / instruction, which, when executed by a processor, implements the protein structure prediction method according to any one of claims 1 to 11.