Protein sequence property prediction method and device, equipment and storage medium
By constructing a multi-task classification model based on the ESM protein language model and a multi-task classifier, the problems of low computational efficiency and insufficient accuracy in protein function prediction are solved, and efficient and accurate prediction of protein properties is achieved.
Patent Information
- Application Number
- CN202511462149.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
Smart Images

Figure CN120932732A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for predicting the properties of protein sequences. Background Technology
[0002] With the rapid development of genome sequencing technology, protein sequence data is growing exponentially, but the experimentally validated functional annotation coverage is less than 5%. Traditional protein function prediction techniques rely on the following methods:
[0003] Homology-based methods infer sequence similarity, but their accuracy in predicting distant homologous proteins drops significantly. Manual feature engineering methods extract shallow features such as amino acid composition and physicochemical properties, combining them with machine learning models (e.g., SVM, random forest) for classification, but they cannot capture deep semantic relationships within sequences, and suffer from time-consuming feature engineering and low prediction accuracy. Furthermore, existing single-task models exhibit low computational efficiency and insufficient feature reuse capabilities in multi-label prediction scenarios. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, device, and storage medium for predicting protein sequence properties to address the aforementioned technical problems.
[0005] In a first aspect, embodiments of this application provide a method for predicting the properties of a protein sequence, the method comprising:
[0006] A multi-task classification model is constructed based on the ESM protein language model and a multi-task classifier.
[0007] The protein sequence dataset is input into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model.
[0008] The target protein sequence is input into the protein property prediction model for prediction, and the enzyme category, substrate type, and nucleic acid binding characteristics of the target protein are obtained.
[0009] In one embodiment, the multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including:
[0010] The first feature vector matrix and the corresponding label are input into the multi-task classifier for initial prediction to obtain the initial prediction results of the enzyme category, substrate type and nucleic acid binding characteristics of each specified protein, and a first vector matrix is constructed based on each initial prediction result.
[0011] Randomly replace some elements in the first vector matrix with real labels to obtain the second vector matrix;
[0012] Based on the second vector matrix and the first vector matrix, the error matrix is obtained;
[0013] The error matrix is fused with the first eigenvector matrix to obtain the second eigenvector matrix.
[0014] The second feature vector matrix is input into the multi-task classifier for prediction to obtain the enzyme category, substrate type, and nucleic acid binding characteristics of each specified protein.
[0015] In one embodiment, obtaining the error matrix based on the second vector matrix and the first vector matrix includes:
[0016] The error matrix is obtained based on the difference between the second vector matrix and the first vector matrix.
[0017] In one embodiment, the step of fusing the error matrix with the first feature vector matrix to obtain the second feature vector matrix includes:
[0018] The error matrix is input into the self-attention mechanism module for feature enhancement to obtain the enhanced error matrix;
[0019] The enhancement error matrix is fused with the first eigenvector matrix to obtain the second eigenvector matrix.
[0020] In one embodiment, the multi-task classifier includes a first classifier, a second classifier, and a third classifier, and the ESM protein language model is connected in parallel with the first classifier, the second classifier, and the third classifier; the first classifier is used to predict the enzyme class of the protein, the second classifier is used to predict the substrate class of the protein, and the third classifier is used to predict the binding properties of the protein.
[0021] In one embodiment, the multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including:
[0022] The first feature vector matrix is input into the first classifier, the second classifier, and the third classifier respectively for prediction, to obtain the predicted probability of the enzyme category, the predicted probability of the substrate category, and the predicted probability of the nucleic acid binding property.
[0023] Based on the predicted probability of the enzyme category and the first tag, calculate the first cross-entropy loss; based on the predicted probability of the substrate category and the second tag, calculate the second cross-entropy loss; based on the predicted probability of the nucleic acid binding characteristics and the third tag, calculate the third cross-entropy loss.
[0024] The parameters of the multi-task classification model are fine-tuned based on the average of the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss.
[0025] In one embodiment, before inputting the protein sequence dataset into the multi-task classification model for training, the method further includes:
[0026] The specified protein sequences are preprocessed and padded to a fixed length.
[0027] Secondly, embodiments of this application also provide a protein sequence property prediction device, the device comprising:
[0028] The model building module is used to build multi-task classification models based on the ESM protein language model and multi-task classifier.
[0029] The model training module is used to input the protein sequence dataset into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model.
[0030] The model prediction module is used to input the target protein sequence into the protein property prediction model for prediction, and obtain the enzyme category, substrate type, and nucleic acid binding characteristics of the target protein.
[0031] Thirdly, embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the method described in the first aspect above.
[0032] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method described in the first aspect above.
[0033] The aforementioned protein sequence property prediction method, apparatus, device, and readable storage medium construct a multi-task classification model based on the ESM protein language model and a multi-task classifier. A protein sequence dataset is input into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model extracts features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier trains the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model. A target protein sequence is input into the protein property prediction model for prediction to obtain the enzyme type, substrate type, and nucleic acid binding characteristics of the target protein. This enables the prediction of multiple protein properties, improving prediction efficiency and accuracy.
[0034] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0035] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0036] Figure 1 This is a hardware structure block diagram of a terminal device for a protein sequence property prediction method in one embodiment;
[0037] Figure 2 This is a flowchart illustrating a protein sequence property prediction method in one embodiment;
[0038] Figure 3 This is a schematic diagram of the training process of a multi-task classification model in one embodiment;
[0039] Figure 4 This is a structural block diagram of a protein sequence property prediction device in one embodiment;
[0040] Figure 5 This is a schematic diagram of the computer device structure in one embodiment. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0042] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0043] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal of the protein sequence property prediction method in this embodiment. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.
[0044] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the protein sequence property prediction method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module for wireless communication with the Internet.
[0046] This application provides a method for predicting the properties of protein sequences, which can be applied to... Figure 1 Taking the terminal in the example of this, for example... Figure 2 As shown, the method includes the following steps:
[0047] Step 201: Construct a multi-task classification model based on the ESM protein language model and a multi-task classifier.
[0048] ESM (Evolutionary Scale Modeling) protein language models are a class of advanced models based on deep learning technology that can effectively process long sequence data and capture long-range dependencies in protein sequences.
[0049] This application embodiment is based on the ESM protein language model and a multi-task classifier to construct a multi-task classification model for predicting various properties of proteins. Specifically, the ESM protein language model is used to extract features from protein sequences, capturing deep semantic relationships within the protein sequences to obtain high-dimensional semantic features shared by the multi-task classifier. The multi-task classifier then uses these shared high-dimensional semantic features to predict various protein properties and outputs the corresponding prediction results.
[0050] For example, the multi-task classification model includes an ESM protein language model and three independent classifiers connected in parallel with the ESM protein language model. The three independent classifiers are used to predict the enzyme class, substrate class, and nucleic acid binding characteristics of the protein, respectively.
[0051] Step 202: Input the protein sequence dataset into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model.
[0052] In this step, the constructed multi-task classification model is trained using a protein sequence dataset to obtain a protein property prediction model, which is used to predict the enzyme type, substrate type, and nucleic acid binding characteristics of a protein.
[0053] Step 203: Input the target protein sequence into the protein property prediction model for prediction to obtain the enzyme type, substrate type, and nucleic acid binding characteristics of the target protein.
[0054] In the aforementioned protein sequence property prediction method, a multi-task classification model is constructed based on the ESM protein language model and a multi-task classifier. The protein sequence dataset is input into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model. The target protein sequence is input into the protein property prediction model for prediction to obtain the enzyme category, substrate type, and nucleic acid binding characteristics of the target protein. This method enables the prediction of multiple protein properties, improving prediction efficiency and accuracy.
[0055] In one embodiment, the multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including the following steps:
[0056] Step 301: Input the first feature vector matrix and the corresponding label into the multi-task classifier for initial prediction to obtain the initial prediction results of the enzyme category, substrate type and nucleic acid binding characteristics of each specified protein, and construct the first vector matrix based on each initial prediction result.
[0057] Specifically, the first feature vector corresponding to each specified protein sequence and the corresponding real label are input into the multi-task classifier for initial prediction to obtain the initial prediction results of the enzyme category, substrate type and nucleic acid binding characteristics of each specified protein.
[0058] For example, the specified protein is a deaminase, wherein the corresponding enzyme category is divided into 19 categories, the substrate category into 2 categories, and the nucleic acid binding category into 6 categories. The initial prediction results include the prediction probabilities corresponding to the 19 enzyme categories, the prediction probabilities corresponding to the 2 substrate categories, and the prediction probabilities corresponding to the 6 nucleic acid binding categories. The prediction results of each classification task are concatenated to form a first vector matrix M1.
[0059] Step 302: Randomly replace some elements in the first vector matrix with real labels to obtain the second vector matrix.
[0060] Specifically, in the first vector matrix M1, a portion of the elements are randomly masked, and then the masked predicted probabilities are filled with the true labels, forming the second vector matrix M2.
[0061] Step 303: Based on the second vector matrix and the first vector matrix, obtain the error matrix.
[0062] Specifically, the error matrix is obtained by subtracting the first vector matrix M1 from the second vector matrix M2.
[0063] Step 304: Perform feature fusion between the error matrix and the first feature vector matrix to obtain the second feature vector matrix.
[0064] Step 305: Input the second feature vector matrix into the multi-task classifier for prediction to obtain the enzyme category, substrate type, and nucleic acid binding characteristics of each specified protein.
[0065] In one embodiment, obtaining the error matrix based on the second vector matrix and the first vector matrix includes: obtaining the error matrix based on the difference between the second vector matrix and the first vector matrix.
[0066] In one embodiment, the step of fusing the error matrix with the first feature vector matrix to obtain a second feature vector matrix includes: inputting the error matrix into a self-attention mechanism module for feature enhancement to obtain an enhanced error matrix; and fusing the enhanced error matrix with the first feature vector matrix to obtain a second feature vector matrix.
[0067] In this embodiment, a dynamic attention mechanism is constructed by using the error between the known label and the prediction result. The error matrix is input into the self-attention mechanism module for feature enhancement. Then, the enhanced error matrix is fused with the first feature vector matrix to obtain a second feature vector matrix. The second feature vector matrix is then fed into the network of the multi-task classifier. By fusing the initial first feature vector matrix with the error information, the feature representation is optimized, and finally, a more accurate classification prediction is output.
[0068] In one embodiment, such as Figure 3 As shown, the multi-task classification model can be divided into a feature extraction module, a classification module, and a feature enhancement module. The feature extraction module extracts features using the ESM protein language model. The classification module includes a first classifier, a second classifier, and a third classifier connected in parallel with the ESM protein language model. The first classifier predicts the enzyme category of the protein, the second classifier predicts the substrate category of the protein, and the third classifier predicts the binding properties of the protein. The first classifier, the second classifier, and the third classifier constitute the multi-task classifier.
[0069] During the training of the multi-task classification model, the protein sequence dataset is input into the ESM protein language model for feature extraction, resulting in a high-dimensional feature vector F1. This high-dimensional feature vector F1 is then input into the first classifier, the second classifier, and the third classifier, respectively, to obtain initial prediction results p1 for enzyme category, p2 for substrate type, and p3 for nucleic acid binding characteristics. Initial prediction results p1 include the prediction probabilities corresponding to each enzyme category, p2 includes the prediction probabilities corresponding to each substrate category, and p3 includes the prediction probabilities corresponding to each nucleic acid binding category. The initial prediction results for each classification task are concatenated to form a first vector matrix M1. A portion of elements in the first vector matrix M1 is randomly masked to obtain a third vector matrix M3. The masked elements are then filled with the true labels (p'1 / p'2 / p'3), resulting in a second vector matrix M2. The first vector matrix M1 is subtracted from the second vector matrix M2 to obtain the error matrix. The error matrix is input into the self-attention mechanism module for feature enhancement to obtain the enhanced error matrix; the enhanced error matrix is fused with the high-dimensional feature vector F1 to obtain the second feature vector matrix F2; the second feature vector matrix F2 is input into the first classifier, the second classifier and the third classifier to predict the final classification label, and the final prediction results p''1 for enzyme category, p''2 for substrate type and p''3 for nucleic acid binding characteristics are obtained.
[0070] In one embodiment, the multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including the following:
[0071] The first feature vector matrix is input into the first classifier, the second classifier, and the third classifier for prediction, respectively, to obtain the predicted probability of the enzyme category, the predicted probability of the substrate category, and the predicted probability of the nucleic acid binding characteristic. Based on the predicted probability of the enzyme category and the first label, a first cross-entropy loss is calculated; the first label is the true label of the enzyme category. Based on the predicted probability of the substrate category and the second label, a second cross-entropy loss is calculated; the second label is the true label of the substrate category. Based on the predicted probability of the nucleic acid binding characteristic and the third label, a third cross-entropy loss is calculated; the third label is the true label of the nucleic acid binding characteristic. Based on the average of the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss, the parameters of the multi-task classification model are fine-tuned.
[0072] In this embodiment of the application, during the training of the multi-task classification model, the training of the multi-task classification model is updated by the average value of the loss of the three classifiers, and the model parameters of the ESM protein language model and the parameters of each classifier are dynamically fine-tuned.
[0073] In one embodiment, before inputting the protein sequence dataset into the multi-task classification model for training, the method further includes: performing data preprocessing on each of the specified protein sequences, padding all the specified protein sequences to a fixed length.
[0074] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0075] This application also provides a protein sequence property prediction device, such as... Figure 4 As shown, the device includes:
[0076] Model building module 10 is used to build a multi-task classification model based on the ESM protein language model and a multi-task classifier;
[0077] The model training module 20 is used to input the protein sequence dataset into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model.
[0078] The model prediction module 30 is used to input the target protein sequence into the protein property prediction model for prediction, and obtain the enzyme category, substrate type and nucleic acid binding characteristics of the target protein.
[0079] In one embodiment, the model training module 20 is further configured to: input the first feature vector matrix and its corresponding labels into the multi-task classifier for initial prediction, obtain initial prediction results of the enzyme type, substrate type, and nucleic acid binding characteristics of each specified protein, and construct a first vector matrix based on the initial prediction results; randomly replace some elements in the first vector matrix with real labels to obtain a second vector matrix; obtain an error matrix based on the second vector matrix and the first vector matrix; perform feature fusion between the error matrix and the first feature vector matrix to obtain a second feature vector matrix; and input the second feature vector matrix into the multi-task classifier for prediction to obtain the enzyme type, substrate type, and nucleic acid binding characteristics of each specified protein.
[0080] In one embodiment, the model training module 20 is further configured to: obtain an error matrix based on the difference between the second vector matrix and the first vector matrix.
[0081] In one embodiment, the model training module 20 is further configured to: input the error matrix into the self-attention mechanism module for feature enhancement to obtain an enhanced error matrix; and fuse the enhanced error matrix with the first feature vector matrix to obtain a second feature vector matrix.
[0082] In one embodiment, the multi-task classifier includes a first classifier, a second classifier, and a third classifier, and the ESM protein language model is connected in parallel with the first classifier, the second classifier, and the third classifier; the first classifier is used to predict the enzyme class of the protein, the second classifier is used to predict the substrate class of the protein, and the third classifier is used to predict the binding properties of the protein.
[0083] In one embodiment, the model training module 20 is further configured to: input the first feature vector matrix into the first classifier, the second classifier, and the third classifier respectively for prediction, to obtain the predicted probability of the enzyme category, the predicted probability of the substrate category, and the predicted probability of the nucleic acid binding characteristic; calculate a first cross-entropy loss based on the predicted probability of the enzyme category and a first label; calculate a second cross-entropy loss based on the predicted probability of the substrate category and a second label; calculate a third cross-entropy loss based on the predicted probability of the nucleic acid binding characteristic and a third label; and fine-tune the parameters of the multi-task classification model based on the average of the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss.
[0084] In one embodiment, the apparatus further includes a data preprocessing module for preprocessing each of the specified protein sequences to fill them to a fixed length.
[0085] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.
[0086] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for predicting the properties of protein sequences.
[0087] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0088] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in any of the protein sequence property prediction method embodiments described above.
[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0090] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0091] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for predicting the properties of a protein sequence, characterized in that, The method includes: A multi-task classification model is constructed based on the ESM protein language model and a multi-task classifier. The protein sequence dataset is input into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model. The target protein sequence is input into the protein property prediction model for prediction, and the enzyme category, substrate type, and nucleic acid binding characteristics of the target protein are obtained.
2. The method according to claim 1, characterized in that, The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including: The first feature vector matrix and the corresponding label are input into the multi-task classifier for initial prediction to obtain the initial prediction results of the enzyme category, substrate type and nucleic acid binding characteristics of each specified protein, and a first vector matrix is constructed based on each initial prediction result. Randomly replace some elements in the first vector matrix with real labels to obtain the second vector matrix; Based on the second vector matrix and the first vector matrix, the error matrix is obtained; The error matrix is fused with the first eigenvector matrix to obtain the second eigenvector matrix. The second feature vector matrix is input into the multi-task classifier for prediction to obtain the enzyme category, substrate type, and nucleic acid binding characteristics of each specified protein.
3. The method according to claim 2, characterized in that, The process of obtaining the error matrix based on the second vector matrix and the first vector matrix includes: The error matrix is obtained based on the difference between the second vector matrix and the first vector matrix.
4. The method according to claim 2, characterized in that, The step of fusing the error matrix with the first feature vector matrix to obtain the second feature vector matrix includes: The error matrix is input into the self-attention mechanism module for feature enhancement to obtain the enhanced error matrix; The enhancement error matrix is fused with the first eigenvector matrix to obtain the second eigenvector matrix.
5. The method according to claim 1, characterized in that, The multi-task classifier includes a first classifier, a second classifier, and a third classifier. The ESM protein language model is connected in parallel with the first classifier, the second classifier, and the third classifier. The first classifier is used to predict the enzyme class of the protein, the second classifier is used to predict the substrate class of the protein, and the third classifier is used to predict the binding properties of the protein.
6. The method according to claim 5, characterized in that, The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels, including: The first feature vector matrix is input into the first classifier, the second classifier, and the third classifier respectively for prediction, to obtain the predicted probability of the enzyme category, the predicted probability of the substrate category, and the predicted probability of the nucleic acid binding property. Based on the predicted probability of the enzyme category and the first tag, calculate the first cross-entropy loss; based on the predicted probability of the substrate category and the second tag, calculate the second cross-entropy loss; based on the predicted probability of the nucleic acid binding characteristics and the third tag, calculate the third cross-entropy loss. The parameters of the multi-task classification model are fine-tuned based on the average of the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss.
7. The method according to claim 1, characterized in that, Before inputting the protein sequence dataset into the multi-task classification model for training, the method further includes: The specified protein sequences are preprocessed and padded to a fixed length.
8. A protein sequence property prediction device, characterized in that, The device includes: The model building module is used to build multi-task classification models based on the ESM protein language model and multi-task classifier. The model training module is used to input the protein sequence dataset into the multi-task classification model for training. The protein sequence dataset includes multiple specified protein sequences and corresponding labels. The ESM protein language model is used to extract features from each specified protein sequence to obtain a first feature vector matrix. The multi-task classifier is used to train the multi-task classification model based on the first feature vector matrix and the corresponding labels to obtain a protein property prediction model. The model prediction module is used to input the target protein sequence into the protein property prediction model for prediction, and obtain the enzyme category, substrate type, and nucleic acid binding characteristics of the target protein.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Drug and target interaction prediction method and device, equipment and storage medium
CN113160894A
Rapid screening method of comprehensive in-vitro antioxidant active polypeptide
CN117437982A
Thermal stability new enzyme design and transformation method based on deep learning model
CN118782150A
High-throughput screening method and device for antioxidant polypeptides
CN119560027A
Anti-fraud model training method based on deep learning
CN119599675A