A protein function annotation method and device based on pre-trained large language model

By using a protein functional annotation method based on a pre-trained large language model, the functional domain categories and locations on a complete protein sequence are solved, and efficient and accurate functional annotation is achieved.

CN119479836BActive Publication Date: 2025-05-06ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510058985.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing deep learning-based protein function annotation model can only classify protein functional domain fragments, making it difficult to accurately identify functional domains on complete protein sequences.

Method used

The protein function annotation method based on pre-trained large language model is adopted, and the protein sequence to be annotated is input into the trained protein functional domain classification model, predict the functional domain category, and then input the functional domain identification model to determine the location of the functional domain category, and finally perform functional annotation.

Benefits of technology

Accurate identification of the functional domain categories and locations of complete protein sequences is achieved, and the efficiency and accuracy of protein functional annotation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479836B_ABST
    Figure CN119479836B_ABST
Patent Text Reader

Abstract

The present application relates to a protein function annotation method and device based on a pre-trained large language model, which is applied to the field of artificial intelligence-driven computational biology, wherein the protein function annotation method comprises: inputting the protein sequence to be annotated into a target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated; inputting the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into a target protein functional domain recognition model to obtain the target position where the functional domain category of the protein sequence to be annotated is located; and functionally annotating the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located. Through the present application, the effect of accurately and efficiently identifying the functional domains on the complete protein sequence is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence-driven computational biology, and in particular to a protein function annotation method and device based on a pre-trained large language model. Background Art

[0002] Protein function annotation is an extremely important task in the fields of molecular biology and bioinformatics. The purpose of protein function annotation is to provide family and domain classification for protein sequences, which is critical for analyzing new genomes and metagenomes and guiding specific protein and system experimental work.

[0003] Generally, traditional methods are used to perform functional annotation of proteins, which will divide the proteins into multiple segments and then perform functional classification. However, the efficiency and accuracy of protein functional annotation still need to be improved. In recent years, with the development of deep learning technology, protein functional annotation has become more efficient and accurate. Protein functional annotation methods based on deep learning, such as ProtCNN and ProtENN, perform functional annotation by learning vector representations of amino acid sequences. These models can infer known evolutionary substitution patterns and effectively cluster sequences from unseen families. However, existing protein annotation models based on deep learning can only classify protein functional domain fragments. Therefore, how to accurately and efficiently identify functional domains on complete protein sequences is still a major problem in current research.

[0004] There is currently no effective solution to the problem of how to accurately and efficiently identify functional domains on complete protein sequences in related technologies. Summary of the invention

[0005] In this embodiment, a protein function annotation method and apparatus based on a pre-trained large language model are provided to solve the problem of how to accurately and efficiently identify functional domains on a complete protein sequence in the related art.

[0006] In a first aspect, a protein function annotation method based on a pre-trained large language model is provided in this embodiment, and the method comprises:

[0007] Inputting the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated; the target protein functional domain classification model is obtained by training a protein functional domain classification model based on a pre-trained protein large language model according to the complete protein sequence in a preset training set;

[0008] The functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated are input into a target protein functional domain recognition model to obtain a target position where the functional domain category of the protein sequence to be annotated is located; the target protein functional domain recognition model is obtained by training a protein functional domain recognition model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in a preset training set; the pre-trained protein large language model is a model for determining protein representations in a protein sequence;

[0009] Functional annotation is performed on the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located.

[0010] In some of the embodiments, the target protein functional domain classification model includes a pre-trained protein large language model layer and a classification layer;

[0011] The step of inputting the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated includes:

[0012] After inputting the protein sequence to be annotated into the target protein functional domain classification model, determining the protein representation of the protein sequence to be annotated based on the pre-trained protein large language model layer in the target protein functional domain classification model;

[0013] According to the classification layer of the target protein functional domain classification model, the functional domain category of the protein sequence to be annotated is predicted based on the protein representation.

[0014] In some of the embodiments, the target protein functional domain identification model includes: a pre-trained protein large language model layer, an interval representation layer, and a functional domain representation layer;

[0015] The step of inputting the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into a target protein functional domain recognition model to obtain a target position where the functional domain category of the protein sequence to be annotated is located comprises:

[0016] Determining the protein representation of the protein sequence to be annotated based on the pre-trained protein large language model layer;

[0017] Based on the functional domain representation layer, the functional domain categories contained in the protein sequence to be annotated are mapped to a preset latent space to obtain a latent space category representation;

[0018] Based on the interval representation layer, mapping the protein representation to a preset latent space to obtain a latent space position representation;

[0019] Based on the latent space category representation and the latent space position representation, a target position where the functional domain category of the protein sequence to be annotated is located is determined.

[0020] In some embodiments, determining the target position where the functional domain category of the protein sequence to be annotated is located based on the latent space category representation and the latent space position representation includes:

[0021] determining a similarity between the latent space category representation and the latent space position representation;

[0022] Determining, according to the similarity, the category probabilities of the functional domain categories corresponding to the multiple functional domain positions in the protein to be annotated;

[0023] The function domain category whose category probability exceeds a preset probability threshold is determined as a target function domain category, and the position corresponding to the target function domain category is determined as a target position.

[0024] In some embodiments, determining the position corresponding to the target functional domain category as the target position includes:

[0025] Acquire multiple functional domain positions corresponding to the target functional domain category;

[0026] The non-overlapping functional domain positions are determined as target positions.

[0027] In some embodiments, the method further comprises:

[0028] Setting a training set and a validation set including a plurality of protein sequences; the training set and the validation set including a plurality of complete protein sequences that do not overlap with each other;

[0029] Obtaining a first protein sequence in the training set, inputting the first protein sequence into the pre-trained protein large language model layer, and obtaining an initial category representation in the protein representation;

[0030] Predicting the initial category representation through the classification layer to determine a predicted functional domain category label of the first protein sequence;

[0031] Obtaining a true category label of a first protein sequence in the training set, and determining a label loss value according to the predicted functional domain category label and the true category label;

[0032] Updating the protein functional domain classification model according to the label loss value;

[0033] obtaining a second protein sequence in the verification set, and determining a verification value of the second protein sequence;

[0034] The updated protein functional domain classification model is verified according to the verification value to obtain the target protein functional domain classification model.

[0035] In some embodiments, the method further comprises:

[0036] Obtaining a first protein sequence in the training set, inputting the first protein sequence into the pre-trained protein large language model layer, and obtaining an initial position representation and an initial category representation of the protein sequence;

[0037] determining, from the initial position representation, a latent spatial position representation corresponding to the functional domain of the first protein sequence through an interval representation layer;

[0038] Obtaining a true category label of a first protein sequence in the training set, and determining an initial representation of a functional domain according to the true category label;

[0039] Based on the initial category representation of the protein sequence, determining the latent space category representation corresponding to the functional domain of the first protein sequence through the functional domain representation layer;

[0040] Training and updating the protein functional domain recognition model according to the latent space position representation and the latent space category representation;

[0041] obtaining a second protein sequence in the verification set, and determining a verification value of the second protein sequence;

[0042] The updated protein functional domain recognition model is verified according to the verification value to obtain the target protein functional domain recognition model.

[0043] In a second aspect, a protein function annotation device based on a pre-trained large language model is provided in this embodiment, the device comprising: a category determination module, a position identification module and an annotation module;

[0044] The category determination module is used to input the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated; the target protein functional domain classification model is obtained by training a protein functional domain classification model based on a pre-trained protein large language model according to the complete protein sequence in a preset training set;

[0045] The position identification module is used to input the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into the target protein functional domain identification model to obtain the target position where the functional domain category of the protein sequence to be annotated is located; the target protein functional domain identification model is obtained by training a protein functional domain identification model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in a preset training set; the pre-trained protein large language model is a model for determining protein representations in a protein sequence;

[0046] The annotation module is used to perform functional annotation on the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located.

[0047] In a third aspect, an electronic device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the protein function annotation method based on the pre-trained large language model described in the first aspect is implemented.

[0048] In a fourth aspect, in this embodiment, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the protein function annotation method based on the pre-trained large language model described in the first aspect above is implemented.

[0049] Compared with the related art, a protein function annotation method and device based on a pre-trained large language model is provided in this embodiment. By inputting the protein to be annotated into a trained protein functional domain classification model, the functional domain category contained in the protein sequence to be annotated is obtained, and then the functional domain category and the protein to be annotated are input into the protein functional domain recognition model to obtain the target position corresponding to the functional domain category, and functional annotation is performed on the protein to be annotated according to the functional domain category and the target position corresponding to the functional domain category, so as to achieve complete recognition of the protein to be annotated.

[0050] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0052] Figure 1It is a hardware structure block diagram of a terminal of a protein function annotation method based on a pre-trained large language model provided in an embodiment of the present application;

[0053] Figure 2 It is a flowchart of a protein function annotation method based on a pre-trained large language model provided in an embodiment of the present application;

[0054] Figure 3 is a flow chart of a protein function annotation method based on a pre-trained large language model and named entity recognition provided in this specific embodiment;

[0055] Figure 4 is a flowchart of the initial characterization calculation of protein functional domains provided in this specific embodiment;

[0056] Figure 5 is an architecture diagram of a pre-trained large language model provided in this specific embodiment;

[0057] Figure 6 is a flow chart of protein functional domain classification model training based on a pre-trained large language model provided in this specific embodiment;

[0058] Figure 7 It is a flow chart of protein functional domain recognition model training based on pre-trained large language model and named entity recognition provided by this specific embodiment;

[0059] Figure 8 It is a flowchart of protein function annotation reasoning based on pre-trained large language model and named entity recognition provided by this specific embodiment;

[0060] Fig. 9 It is a structural block diagram of a protein function annotation device based on a pre-trained large language model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0062] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0063] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 is a hardware structure block diagram of a terminal of a protein function annotation method based on a pre-trained large language model provided in an embodiment of the present application. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0064] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the protein function annotation method based on the pre-trained large language model in the present embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0065] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.

[0066] Protein function annotation is an extremely important task in the field of molecular biology and bioinformatics. The purpose of protein function annotation is to provide family and domain classification for protein sequences, which is critical for analyzing new genomes and metagenomes and guiding specific protein and system experiments. The limitations of current protein function annotation work include slow annotation growth and insufficient annotation coverage. At least one-third of microbial proteins cannot be annotated by comparing them with functional feature sequences. The above problems limit the use of data from different organisms.

[0067] In recent years, with the development of deep learning technology, protein function annotation has become more efficient and accurate. In particular, protein function annotation methods based on deep learning, such as deep learning models ProtCNN and ProtENN, perform functional annotation by learning vector representations of amino acid sequences. These models can infer known evolutionary substitution patterns and effectively cluster sequences from unseen families. At the same time, deep learning models combined with existing methods can significantly improve the accuracy of remote homology detection, indicating that deep models have learned complementary information. These models can not only accurately predict the functions of unaligned amino acid sequences, but also expand the coverage of databases used in protein research, thereby helping to overcome the limitations of current protein annotation work and improve the overall quality and coverage of functional annotations.

[0068] However, the core problem at present is that the existing protein annotation models based on deep learning can only classify protein functional domain fragments, while the reality is often that the complete protein sequence needs to be functionally annotated, that is, the functional domain categories and their corresponding positions contained therein need to be parsed. However, there are tens of thousands of families (functional labels) of seed data sets in the protein research database, and there are a large number of discontinuous functional domains. This complexity not only reflects the diversity and complexity of protein functional domains, but also puts higher demands on the model's functional domain recognition capabilities. Therefore, how to accurately identify the types and positions of functional domains on complete protein sequences has become a major problem in current research.

[0069] In this embodiment, a protein function annotation method based on a pre-trained large language model is provided, which decomposes the protein function annotation task into two subtasks: protein functional domain classification and protein functional domain identification. The protein functional domain classification model is used to predict the functional domain category contained in the protein sequence, and then the protein functional domain identification model is used to identify the location of the functional domain. In the embodiment of the present application, a pre-trained protein large language model is applied to the protein functional domain classification model and the protein functional domain identification model to obtain protein representation, and named entity recognition technology is applied to the field of protein functional identification. By learning complex sequence features and contextual relationships, these models can more accurately annotate functional domains in proteins. Figure 2 is a flow chart of a protein function annotation method based on a pre-trained large language model provided in an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:

[0070] Step S210: input the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated.

[0071] The target protein functional domain classification model is obtained by training a protein functional domain classification model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence in the preset training set. The pre-trained protein large language model is a model for determining protein representations in protein sequences.

[0072] In this step, the complete protein sequence to be identified, i.e., the protein sequence to be annotated, is input into a well-trained target protein functional domain classification model to obtain the functional domain category of the protein sequence to be annotated. The target protein functional domain classification model is trained based on the protein functional domain classification model of the pre-trained protein large language model.

[0073] On an evolutionary scale, protein sequences contain information about biological structure and function. The protein language model is obtained by unsupervised training of a large amount of protein sequence data, and has a very high number of parameters. Its core architecture is based on the encoder component of the Transformer, which is a neural network model that uses a self-attention mechanism and has proven its efficiency and strong representation capabilities in the field of natural language processing. During training, the protein language model uses a masked learning strategy to enhance its understanding of local and global sequence information. Specifically, the protein language model randomly selects a portion of amino acid residues and masks them (i.e., replaces them with special mask tags); then the protein language model needs to predict the amino acid residues at the masked position based on the unmasked residues. The training goal of the protein language model is to minimize the prediction error of the masked residues. This training method enables the protein language model to learn the long-range dependencies and local structural features in the protein sequence, thereby generating a richer intermediate representation. Through the process of learning and denoising of randomly masked protein sequences, the protein large language model can learn the evolutionary correlation and topological structure of proteins, so that the pre-trained protein large language model has the ability to evaluate the conservation of amino acids in proteins, that is, whether a specific residue type in a sequence conforms to the semantics and word order rules of protein language in nature. In addition, the pre-training process of the protein large language model is unsupervised, which means that the model does not require specific label information during training. The model learns useful representations from a large amount of protein sequence data through self-supervised learning. This unsupervised training method not only saves the cost of annotating data, but also enables the model to process larger-scale sequence data, thereby improving its generalization ability and applicability. Therefore, the protein large language model has the potential to learn the patterns of protein sequences in evolution.

[0074] Furthermore, the pre-trained protein language model is a model for determining protein representations in protein sequences. The core architecture of the pre-trained protein language model is based on the Transformer encoder and uses a large number of protein sequences for unsupervised pre-training, where the Transformer encoder is a neural network model that uses a self-attention mechanism.

[0075] In some of the embodiments, the pre-trained protein language model is obtained by training a protein language model based on a training set based on an embedding layer, a position embedding layer, a language model module, and a linear normalization layer. Wherein, the training set includes a complete protein sequence. Exemplarily, the Pfam database, which is widely used in protein research, is used as a training set; the Pfam database focuses on the classification and annotation of protein families and domains. It provides support for the analysis of new genomes and metagenomes, and guides experimental studies related to specific proteins and their systems. Pfam helps scientists deeply understand the structure and function of proteins by providing family-specific hidden Markov models (HMMs) and multiple sequence alignments. The database contains multiple protein families, each family represented by multiple sequence alignments and HMMs, corresponding to one or more functional regions in the protein, i.e., domains. Each family has a "seed alignment" containing a set of representative sequences, which are used to generate profile HMMs and then search the sequence database of Pfam. In addition, each entry of Pfam is manually curated and annotated, containing functional information from the literature, ensuring the high quality and reliability of the data. Therefore, the Pfam database plays a key role in protein sequence classification, functional annotation, and the study of structure-function relationships.

[0076] In the process of protein function annotation based on pre-trained large language models and named entity recognition, validation sets and test sets were used in addition to the training set; among them, the complete protein sequences were obtained from the UniprotKB database according to the sequence names in the Pfam-A seed database, and then the seed sequences with more than 10 sequences in the Pfam family were randomly divided into non-overlapping training sets, validation sets and test sets to ensure that there were no repeated protein sequences in the training sets, validation sets and test sets.

[0077] Furthermore, the language model module includes a multi-head attention module and a feedforward neural network module. The embedding layer converts amino acid residues into embedding vectors and captures position information through position encoding. The self-attention module calculates the relationship between each amino acid and all other amino acids in the protein sequence, enabling the model to focus on different parts of the sequence and dynamically adjust weights. The feedforward neural network further processes the intermediate representation to enhance the representation ability.

[0078] In this step, based on the target protein functional domain classification model obtained by training the training set, the protein sequence to be annotated is classified to obtain the functional domain category, which is conducive to improving the accuracy of determining the functional domain category.

[0079] Step S220, inputting the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into the target protein functional domain recognition model to obtain the target position of the functional domain category of the protein sequence to be annotated.

[0080] Among them, the target protein functional domain recognition model is obtained by training the protein functional domain recognition model based on the pre-trained protein large language model and the named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in the preset training set. The functional domain category obtained according to the target protein functional domain classification model and the protein sequence to be annotated are input into the target protein functional domain recognition model to obtain the target position where the functional domain category is located.

[0081] Furthermore, named entity recognition (NER) is a basic task in natural language processing (NLP), which aims to identify and classify entities with specific meanings from unstructured text, such as names of people, places, organizations, time, date, quantity, currency, percentage, product name, disease name, gene and protein name, etc. The basic process of named entity recognition includes data preprocessing, feature extraction, model training, model evaluation and application. Common models include hidden Markov model (HMM), conditional random field (CRF), support vector machine (SVM) and deep learning model (such as bidirectional LSTM and BERT). These models have a wide range of applications in many fields, such as information extraction, knowledge graph construction, sentiment analysis, biomedical text mining, automatic question answering system, search recommendation system, legal text processing and financial text analysis. At present, the research on named entity recognition based on deep learning is quite mature. The general lightweight named entity recognition GLiNER uses a bidirectional language model such as BERT or deBERTa. Its core concept is to regard the named entity recognition NER task as matching the entity type embedding with the text span representation in the latent space. This method solves the scalability problem of NER and allows bidirectional context processing, thereby achieving richer representation. Therefore, the existing named entity recognition model can be used to assist the protein function annotation task. By using the deep learning model as the core component of the protein function annotation tool, it is beneficial to improve the accuracy and efficiency of protein functional domain annotation.

[0082] Step S230 , performing functional annotation on the protein sequence to be annotated according to the target position of the functional domain category of the protein sequence to be annotated.

[0083] In this step, the protein sequence to be annotated is processed according to the target protein functional domain classification model and the target protein functional domain recognition model to obtain the functional domain category corresponding to the protein sequence to be annotated and the target position corresponding to the functional domain category, and then the protein sequence to be annotated is annotated according to the category and position.

[0084] Through the above steps, the protein to be annotated is input into the trained protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated, and then the functional domain category and the protein to be annotated are input into the protein functional domain recognition model to infer the target position corresponding to the functional domain category, and the protein to be annotated is functionally annotated according to the functional domain category and the target position corresponding to the functional domain category, so as to achieve complete recognition of the protein to be annotated, thereby improving the accuracy of annotation of the complete protein sequence.

[0085] In some of the embodiments, the target protein functional domain classification model and the target protein functional domain recognition model are both trained based on a pre-trained protein large language model.

[0086] In one possible embodiment, the training method of the target protein functional domain classification model includes: setting a training set and a validation set including multiple protein sequences; the training set and the validation set include multiple complete protein sequences that do not overlap with each other; obtaining the first protein sequence in the training set, inputting the first protein sequence to the pre-trained protein large language model layer, and obtaining an initial category representation in the protein representation; predicting the initial category representation through the classification layer to determine the predicted functional domain category label of the first protein sequence; obtaining the true category label of the first protein sequence in the training set, and determining the label loss value according to the predicted functional domain category label and the true category label; updating the protein functional domain classification model according to the label loss value; obtaining the second protein sequence in the validation set, and determining the validation value of the second protein sequence; validating the updated protein functional domain classification model according to the validation value to obtain the target protein functional domain classification model.

[0087] In this embodiment, the protein functional domain classification model to be trained includes a pre-trained protein large language model layer and a classification layer, wherein the classification layer includes a linear layer and an activation layer, and the output layer that finally outputs the functional domain category is a linear layer and a Sigmoid activation function, and the output value is limited to between 0 and 1, which is used to predict the functional domain category label of the protein, wherein the output value represents the probability of belonging to a certain category. The protein functional domain classification problem is a multi-label classification problem.

[0088] Among them, the protein functional domain classification model training process specifically includes: first setting the number of training rounds; inputting the first protein sequence in the training set to the pre-trained protein large language model layer, and obtaining the protein representation corresponding to the complete protein sequence in the training set, that is, the initial category representation, through the pre-trained protein large language model layer; then predicting the functional domain category label of the protein according to the protein representation through the classification layer, and calculating the binary cross entropy loss value between the predicted functional domain category label and the true category label corresponding to the prediction result; finally, according to the calculated cross entropy loss value and using the optimizer, the model parameters of the protein functional domain classification model are updated to obtain the trained target protein functional domain classification model.

[0089] Furthermore, the F1 index of the validation set classification is used to verify the protein functional domain classification model after each round of updating the model parameters, so as to determine the model parameters with the largest F1 value on the validation set as the ideal model parameters, so as to obtain the optimal target protein functional domain classification model according to the ideal model parameters. Among them, the F1 index (F1 Score) is an important indicator for evaluating model performance in the field of machine learning and information retrieval. It combines the performance of precision and recall. The F1 index can take into account the prediction accuracy and coverage of the positive class at the same time. It is mainly used to process data sets with imbalanced categories. The F1 index of the validation set classification is conducive to improving the classification accuracy of the target protein functional domain classification model, thereby improving the accuracy of subsequent protein functional domain annotations.

[0090] For example, the binary cross entropy loss function for a multi-label classification problem is expressed as:

[0091] ;

[0092] Where N is the number of samples, is the true label (0 or 1) of the i-th sample in the j-th category, is the predicted value of the i-th sample in the j-th category (the value after Sigmoid activation). This loss function can calculate the binary cross entropy loss label by label, and then average the losses of all samples and all labels.

[0093] In one possible embodiment, the training method of the target protein functional domain recognition model includes: obtaining the first protein sequence in the training set, inputting the first protein sequence into the pre-trained protein large language model layer, and obtaining the initial position representation and initial category representation of the protein sequence; determining the latent space position representation corresponding to the functional domain of the first protein sequence from the initial position representation through the interval representation layer; obtaining the true category label of the first protein sequence in the training set, and determining the initial representation of the functional domain according to the true category label; based on the initial category representation of the protein sequence, determining the latent space category representation corresponding to the functional domain of the first protein sequence through the functional domain representation layer; training and updating the protein functional domain recognition model according to the latent space position representation and the latent space category representation; obtaining the second protein sequence in the validation set, and determining the validation value of the second protein sequence; validating the updated protein functional domain recognition model according to the validation value to obtain the target protein functional domain recognition model.

[0094] In this embodiment, the protein functional domain recognition model training process specifically includes simultaneously inputting the first protein sequence in the training set, and the real functional domain category label and position label corresponding to the first protein sequence in the training set, wherein the position label corresponds to the interval where the functional domain is located. The protein representation of the first protein sequence is obtained through the pre-trained protein large language model layer, and then the potential space position representation corresponding to the interval protein is obtained based on the protein representation and the interval representation layer. At the same time, the corresponding calculated protein functional domain initial representation is found according to the input real functional domain category label, and the protein functional domain initial representation is passed through the domain representation layer to obtain the functional domain category representation, that is, the potential space category representation. The obtained potential space position representation is calculated with the potential space category representation. Thereafter, the binary cross entropy loss value is calculated according to the real functional domain category label and the position label. Finally, the model parameters of the protein functional domain recognition model are updated according to the calculated binary cross entropy loss value and the optimizer to obtain the trained target protein functional domain recognition model.

[0095] Furthermore, the model parameters are updated using the optimizer according to the calculated cross entropy loss. Afterwards, the F1 index of the second protein sequence classification in the validation set is used to verify the protein functional domain recognition model after each round of updating, and finally the model parameters with the largest F1 value on the validation set are determined as the ideal model parameters to obtain the optimal target protein functional domain recognition model.

[0096] Exemplarily, the above binary cross entropy loss function Loss is:

[0097] ;

[0098] Among them, S is the number of functional domain intervals, T is the number of functional domain categories, is the true label (0 or 1) of the i-th interval in the j-th category, is the predicted value of the i-th sample in the j-th category (the value after Sigmoid activation). This loss function can calculate the binary cross entropy loss label by label, and then average the loss values ​​of all intervals and all categories.

[0099] In some of the embodiments, the target protein functional domain classification model includes a pre-trained protein large language model layer and a classification layer; in step S210, the protein sequence to be annotated is input into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated, including: after the protein sequence to be annotated is input into the target protein functional domain classification model, based on the pre-trained protein large language model layer in the target protein functional domain classification model, determining the protein representation of the protein sequence to be annotated; according to the classification layer of the target protein functional domain classification model, predicting the functional domain category of the protein sequence to be annotated based on the protein representation.

[0100] In this embodiment, the test set is used to infer the protein functional domain classification model. The complete protein sequence to be annotated in the test set is sequentially passed through the pre-trained protein large language model layer and the classification layer, and finally activated by the Sigmoid activation function, and the protein functional domain category with an output value greater than a preset threshold is selected as the prediction result. Preferably, the preset threshold is set to 0.5.

[0101] In some embodiments, the target protein functional domain recognition model includes: a pre-trained protein large language model layer, an interval representation layer and a functional domain representation layer; in step S220, the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated are input into the target protein functional domain recognition model to obtain the target position of the functional domain category of the protein sequence to be annotated, including: based on the pre-trained protein large language model layer, determining the protein representation of the protein sequence to be annotated; based on the functional domain representation layer, mapping the functional domain category contained in the protein sequence to be annotated to a preset latent space to obtain a latent space category representation; based on the interval representation layer, mapping the protein representation to a preset latent space to obtain a latent space position representation; based on the latent space category representation and the latent space position representation, determining the target position of the functional domain category of the protein sequence to be annotated.

[0102] In this embodiment, the target protein functional domain recognition model includes a pre-trained protein large language model layer, an interval representation layer and a functional domain representation layer. The pre-trained protein large language model layer is used to obtain the protein representation of the protein sequence to be annotated. The interval representation layer is composed of two layers of feedforward neural networks, which are used to obtain the latent space representation corresponding to the interval protein from the protein representation. The functional domain representation layer is composed of two layers of feedforward neural networks, which are used to obtain the latent space representation corresponding to the functional domain category. The goal of the interval representation layer and the functional domain representation layer is to map the category representation and the interval representation to the same latent space, obtain the latent space category representation and the latent space position representation, and calculate the similarity between the latent space category representation and the latent space position representation, and then determine the target position based on the similarity.

[0103] In this embodiment, the protein sequence to be annotated of the complete sequence in the test set is input into the target protein functional domain classification model to obtain the predicted functional domain category. Thereafter, the predicted functional domain category and the protein sequence to be annotated are input into the target functional domain recognition model to obtain the target position where the predicted functional domain category is located. Then, according to the predicted functional domain category of the protein sequence to be annotated and the target position corresponding to the functional domain category, the protein sequence to be annotated is annotated, thereby improving the accuracy of identifying the functional domain category and position of the complete protein sequence.

[0104] In some of the embodiments, based on the latent space category representation and the latent space position representation, the target position of the functional domain category of the protein sequence to be annotated is determined, including: determining the similarity between the latent space category representation and the latent space position representation; determining the category probability of the functional domain category corresponding to multiple functional domain positions in the protein to be annotated according to the similarity; determining the functional domain category whose category probability exceeds a preset probability threshold as the target functional domain category, and determining the position corresponding to the target functional domain category as the target position.

[0105] In a possible embodiment, assuming that the length of the protein sequence is N, the dimension of the latent space is D, and the functional domain category corresponding to the protein sequence is M, the sequence representation h of the protein sequence after the pre-trained protein large language model layer is:

[0106] ;

[0107] The initial representation of the protein class p is:

[0108] ;

[0109] The protein functional domain category characterization q after the functional domain characterization layer is:

[0110] ;

[0111] Among them, R represents the set of real numbers in the latent space, the length of the protein sequence is N, the dimension of the latent space is D, and the functional domain category corresponding to the protein sequence is M.

[0112] Assuming that the functional domain is located between the i-th amino acid and the j-th amino acid of the protein, the potential spatial position representation of the interval is for:

[0113] ;

[0114] in, represents the concatenation operator, and They represent the protein representations at the i-th position and the j-th position extracted by the pre-trained protein language model layer. FFN represents the feed-forward neural network layer.

[0115] Then, the interval representation is calculated using the following formula Functional domain category representation Similarity between :

[0116] ;

[0117] in, is the Sigmoid activation function, represents the probability that the interval between the ith amino acid and the jth amino acid of the protein is the functional domain category t, R represents the real number set in the latent space, represents interval latent space representation.

[0118] In some of the embodiments, the method for determining the target position further includes: acquiring multiple functional domain positions corresponding to the target functional domain category; and determining the position of a non-overlapping functional domain interval as the target position.

[0119] In this embodiment, after calculating the similarity between the latent space category representation and the latent space position representation, the positions whose category probabilities corresponding to the similarities are greater than the specified probability threshold and the corresponding functional domain categories are determined according to the activation function of the activation layer; then, the intervals corresponding to all positions are evaluated, and the non-overlapping intervals with the highest similarity are selected as the target intervals; and the target positions of the protein sequences to be annotated are determined according to the target intervals.

[0120] The present embodiment is described and illustrated by means of specific examples below.

[0121] Figure 3 is a flowchart of the protein function annotation method based on pre-trained large language model and named entity recognition provided by this specific embodiment; Figure 3, input the complete protein sequence into the protein annotation model, wherein the protein annotation model includes a functional domain classification model and a functional domain recognition model; wherein the functional domain classification model is the functional domain classification model of the target protein in the aforementioned embodiment, the functional domain recognition model is the functional domain recognition model of the target protein in the aforementioned embodiment, and the protein sequence here is the protein sequence to be annotated in the aforementioned embodiment. The method of inputting the protein sequence into the protein annotation model includes: first, inputting the protein sequence into the functional domain classification model to obtain the functional domain category. Thereafter, the functional domain category and the protein sequence are input into the functional domain recognition model to determine the functional domain category and the corresponding interval of the protein sequence.

[0122] Figure 4 is a flowchart of the initial characterization calculation of protein functional domains provided in this specific embodiment; Figure 4 First, obtain the Pfam dataset, determine the categories corresponding to the Pfam functional domains, and determine the sequence fragments corresponding to all functional domain categories in the Pfam dataset. Input the sequence fragments corresponding to all functional domain categories in the Pram dataset into the protein pre-trained large language model to obtain the protein representations corresponding to all sequence fragments, and calculate the initial value of the protein representation to obtain the initial protein representation in the Pfam domain dataset. Obtain the protein representation of all functional domain sequence fragments through the pre-trained protein large language model, and then calculate the average value of the protein representation corresponding to each type of functional domain label, so as to obtain the initial representation of each type of protein functional domain.

[0123] Figure 5 : is an architecture diagram of the pre-trained protein large language model provided in this specific embodiment; the core architecture of the pre-trained protein large language model is based on the Transformer encoder, and a large number of protein sequences are used for unsupervised pre-training, wherein the Transformer encoder is a neural network model that adopts a self-attention mechanism. The parameters of the pre-trained protein large language model are obtained from a public website. The protein large language model includes an embedding layer, a position embedding layer, a language model module, and a linear normalization layer. Furthermore, the language model module includes a multi-head attention module and a feedforward neural network module. Figure 5As shown, the protein sequence is input into the pre-trained protein language model, and the embedding layer converts the amino acid residues in the protein sequence into an embedding vector to obtain the input representation. Subsequently, the position information is captured based on the input representation through position encoding. Among them, according to the multi-head attention module in the language model module, the relationship between each amino acid and all other amino acids in the protein sequence is calculated, and a residual connection operation is performed, so that the pre-trained protein language model can pay attention to different parts of the sequence and dynamically adjust the weights. The feedforward neural network module further processes the intermediate representation to enhance the representation ability, and then updates the language model module. According to the final updated language model module, the protein representation corresponding to the protein sequence is calculated, and then the linear normalization layer of the pre-trained protein language model is passed to obtain the amino acid probability distribution of the protein sequence, that is, the initial representation of the functional domain of the protein sequence.

[0124] Specifically, all protein functional domain sequence fragments in the Pfam-A seed dataset are input into the protein language model, and then go through the embedding layer, position embedding layer, and language model module in sequence to finally obtain protein representation. Then, the average value of the protein representation corresponding to each type of functional domain label is calculated to obtain the initial representation of the protein functional domain.

[0125] Figure 6 This is a flowchart of a protein functional domain classification model training based on a pre-trained large language model provided in this specific embodiment; the training protein functional domain classification model is used to solve the protein functional domain classification problem. The protein functional domain classification model based on the pre-trained protein large language model is trained according to the complete protein sequence and the real protein functional domain category label. Figure 6 The protein functional domain classification model includes a pre-trained protein large language model layer and a classification layer. The protein sequence is input into the protein pre-trained large language model to obtain the protein representation, and then the classification result is obtained through the classification layer, and the loss value between the classification result and the true category label is calculated. The classification layer includes a linear layer and an activation layer. The final output layer is a linear layer plus a Sigmoid activation function, which limits the output value to between 0 and 1. It is used to predict the functional domain category label of the protein. The output value represents the probability that the protein sequence belongs to a certain functional category.

[0126] Since a protein sequence may contain multiple functional domains, the protein functional domain classification problem is a multi-label classification problem. The protein functional domain classification model training process specifically includes: setting the number of training rounds, obtaining protein representations through the pre-trained protein large language model layer for the protein sequences in the training set, and then predicting the functional domain category of the protein through the classification layer, calculating the binary cross entropy loss between the prediction result and the true category label, and using the Adam optimizer to update the model parameters according to the calculated cross entropy loss to obtain a trained protein functional domain classification model. Furthermore, the F1 score of the classification on the validation set is used as an indicator to verify the protein functional domain classification model of each round, and the parameter with the largest F1 value on the validation set is taken as the ideal model parameter to obtain the optimal protein functional domain classification model.

[0127] Figure 7 This is a flowchart of a protein functional domain recognition model training based on a pre-trained large language model and named entity recognition provided by this specific embodiment; wherein the protein functional domain recognition model is trained to solve the protein functional domain recognition problem. Figure 7 According to the protein sequence, the real protein functional domain category and interval label, a protein functional domain recognition model based on a pre-trained protein large language model and named entity recognition is trained. Among them, the protein functional domain recognition model is used to identify the position of a certain functional domain on the complete protein sequence, that is, the interval corresponding to the functional domain, and the input of its training stage is the complete protein sequence and the real protein functional domain category label and position label. The protein functional domain recognition model includes a pre-trained protein large language model layer, an interval representation layer and a functional domain representation layer. Input the protein sequence to the protein pre-trained large language model layer to obtain the protein representation of the protein sequence. Thereafter, the protein representation is input to the interval representation layer composed of two layers of feedforward neural networks, which is used to obtain the potential space interval representation corresponding to the interval protein from the protein representation, that is, the potential space position representation in the aforementioned embodiment. At the same time, the real Pfam functional domain classification label is selected according to the classification label to obtain the Pfam functional domain initial representation. Thereafter, the Pfam functional domain initial representation is input to the functional domain representation layer composed of two layers of feedforward neural networks to obtain the potential space functional domain category representation corresponding to the functional domain category. The goal of the interval representation layer and the functional domain representation layer is to map the functional domain category representation and the interval representation to the same latent space to obtain their similarity, and then calculate the loss value, and train the protein functional domain recognition model according to the loss value to obtain the target protein functional domain recognition model. In addition, the loss value is calculated according to the real Pfam functional domain interval label.

[0128] Figure 8This is a flowchart of protein function annotation reasoning based on a pre-trained large language model and named entity recognition provided by this specific embodiment; the trained protein function domain classification model and protein function domain recognition model are used to reason about protein function annotations, and the types and positions of functional domains contained in the protein sequence are inferred. Among them, the reasoning process of the protein function domain annotation model includes: inputting the protein sequence into the protein function domain classification model to obtain the predicted functional domain category, and then inputting the predicted functional domain category and the protein sequence into the functional domain recognition model to obtain the interval where the predicted functional domain category is located.

[0129] Specifically, refer to Figure 8 , the protein sequences to be annotated in the test set are input into two branches at the same time, where branch one obtains the classification result of the predicted label of the functional domain category through the trained protein functional domain classification model, and then selects the initial functional domain representation according to the predicted label in the classification result, and then obtains the functional domain category representation after passing through the functional domain representation layer. Branch two obtains protein representation through the trained protein large language model layer, and then the protein representation obtains interval representation after passing through the interval representation layer. Calculate the similarity between the functional domain category representation obtained by branch one and the interval representation obtained by branch two. After activation by the Sigmoid function, select the functional domain interval greater than the specified threshold and the corresponding functional domain category, and then greedily select the non-overlapping interval with the highest similarity, and repeat this operation until all candidate intervals are evaluated, so as to infer the interval where the protein functional domain is located, and then annotate the protein sequences to be annotated in the test set according to the finally inferred functional domain category and the corresponding functional domain interval.

[0130] By the protein annotation method based on the pre-trained large language model and named entity recognition in the above-mentioned embodiment, the protein function annotation task is decomposed into two subtasks of protein function domain classification and protein function domain recognition. The functional domain types contained in the protein sequence are first predicted by the protein function domain classification model, and then the position of the functional domain is identified by the protein function domain recognition model. The present application embodiment applies the pre-trained protein large language model in the protein function domain classification model and the protein function domain recognition model to obtain protein representation, and the named entity recognition technology is applied to the new field of protein function recognition for the first time. By learning complex sequence features and contextual relationships, these models can more accurately annotate the functional domains in proteins. In addition, by reusing the parameters of the pre-trained protein large language model, only a small amount of model parameters need to be trained to obtain good results. This method not only realizes innovation in technology, but also provides important tools and support for research in multiple fields. In genomics and metagenomics research, the protein annotation method of the present application embodiment can efficiently and accurately annotate the protein functional domains in new genomes and metagenomes, and quickly understand the functions of new genes and the ecological functions of microorganisms. In the field of drug development and protein engineering, accurate identification of protein functional domains can accelerate drug design and screening processes, improve protein modification and optimization efficiency, and develop new biocatalysts and engineered enzymes. In terms of disease diagnosis and biomarker discovery, it helps to identify protein functional domains associated with specific diseases, discover new biomarkers, and provide important clues for early diagnosis and disease monitoring.

[0131] In this embodiment, a protein function annotation device based on a pre-trained large language model is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and the descriptions that have been made will not be repeated. The terms "module", "unit", "subunit", etc. used below can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0132] Fig. 9 is a structural block diagram of a protein function annotation device based on a pre-trained large language model provided in an embodiment of the present application, such as Fig. 9 As shown, the device includes: a category determination module 10, a location identification module 20 and an annotation module 30.

[0133] The category determination module 10 is used to input the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated; the target protein functional domain classification model is obtained by training the protein functional domain classification model based on the pre-trained protein large language model according to the complete protein sequence in the preset training set.

[0134] The position recognition module 20 is used to input the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into the target protein functional domain recognition model to obtain the target position of the functional domain category of the protein sequence to be annotated; the target protein functional domain recognition model is obtained by training a protein functional domain recognition model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in a preset training set; the pre-trained protein large language model is a model for determining protein representations in protein sequences.

[0135] The annotation module 30 is used to perform functional annotation on the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located.

[0136] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0137] In this embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0138] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0139] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0140] S1, input the protein sequence to be annotated into the target protein functional domain classification model to obtain the functional domain category contained in the protein sequence to be annotated; the target protein functional domain classification model is obtained by training the protein functional domain classification model based on the pre-trained protein large language model according to the complete protein sequence in the preset training set.

[0141] S2, inputting the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into the target protein functional domain recognition model to obtain the target position of the functional domain category of the protein sequence to be annotated; the target protein functional domain recognition model is trained on the protein functional domain recognition model based on the pre-trained protein large language model and the named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in the preset training set; the pre-trained protein large language model is a model for determining the protein representation in the protein sequence.

[0142] S3, functional annotation of the protein sequence to be annotated is performed according to the target position where the functional domain category of the protein sequence to be annotated is located.

[0143] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0144] In addition, in combination with the protein function annotation method based on the pre-trained large language model provided in the above embodiment, a storage medium can also be provided in this embodiment to implement. The storage medium stores a computer program; when the computer program is executed by the processor, any one of the protein function annotation methods based on the pre-trained large language model in the above embodiment is implemented.

[0145] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.

[0146] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.

[0147] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.

[0148] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.

Claims

1. A protein function annotation method based on a pre-trained large language model, characterized in that: The method comprises: The protein sequence to be annotated is input into a target protein functional domain classification model, wherein the target protein functional domain classification model includes a pre-trained protein large language model layer and a classification layer; based on the pre-trained protein large language model layer in the target protein functional domain classification model, the protein representation of the protein sequence to be annotated is determined; according to the classification layer of the target protein functional domain classification model, the functional domain category contained in the protein sequence to be annotated is predicted based on the protein representation; the target protein functional domain classification model is obtained by training a protein functional domain classification model based on the pre-trained protein large language model according to the complete protein sequence in a preset training set; the core architecture of the pre-trained protein large language model is based on the encoder of Transformer, and is obtained by unsupervised pre-training using the complete protein sequence; The functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated are input into a target protein functional domain recognition model, wherein the target protein functional domain recognition model comprises: a pre-trained protein large language model layer, an interval representation layer and a functional domain representation layer; based on the pre-trained protein large language model layer, the protein representation of the protein sequence to be annotated is determined; based on the functional domain representation layer, the functional domain category contained in the protein sequence to be annotated is mapped to a preset latent space to obtain a latent space category representation; based on the interval representation layer, the protein representation is mapped to a preset latent space to obtain a latent space position representation; the latent space category representation and the latent space position representation are determined. similarity between the features; according to the similarity, determining the category probability of the functional domain categories corresponding to the multiple functional domain positions in the protein to be annotated; determining the functional domain category whose category probability exceeds the preset probability threshold as the target functional domain category, and determining the position corresponding to the target functional domain category as the target position where the functional domain category of the protein sequence to be annotated is located; the target protein functional domain recognition model is obtained by training a protein functional domain recognition model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in a preset training set; the pre-trained protein large language model is a model for determining protein representations in protein sequences; Performing functional annotation on the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located; The method further comprises: setting a training set and a validation set including a plurality of protein sequences; the training set and the validation set include a plurality of complete protein sequences that do not overlap each other; obtaining a first protein sequence in the training set, inputting the first protein sequence into the pre-trained protein large language model layer, and obtaining an initial category representation in the protein representation; predicting the initial category representation through the classification layer, and determining a predicted functional domain category label of the first protein sequence; obtaining a true category label of the first protein sequence in the training set, and determining a label loss value according to the predicted functional domain category label and the true category label; updating the protein functional domain classification model according to the label loss value; obtaining a second protein sequence in the validation set, and determining a validation value of the second protein sequence; validating the updated protein functional domain classification model according to the validation value, and obtaining the target protein functional domain classification model; The method also includes: obtaining a first protein sequence in the training set, inputting the first protein sequence into the pre-trained protein large language model layer, and obtaining an initial position representation and an initial category representation of the protein sequence; determining the latent space position representation corresponding to the functional domain of the first protein sequence from the initial position representation through an interval representation layer; obtaining a true category label of the first protein sequence in the training set, and determining the initial representation of the functional domain according to the true category label; based on the initial category representation of the protein sequence, determining the latent space category representation corresponding to the functional domain of the first protein sequence through a functional domain representation layer; determining the similarity between the latent space position representation and the latent space category representation; calculating a loss value according to the category probability corresponding to the similarity and a preset probability threshold, and training and updating the protein functional domain recognition model according to the loss value; obtaining a second protein sequence in the validation set, and determining a validation value of the second protein sequence; validating the updated protein functional domain recognition model according to the validation value to obtain the target protein functional domain recognition model.

2. The protein function annotation method based on a pre-trained large language model according to claim 1, characterized in that: The step of determining the position corresponding to the target functional domain category as the target position where the functional domain category of the protein sequence to be annotated is located includes: Acquire multiple functional domain positions corresponding to the target functional domain category; The non-overlapping functional domain positions are determined as target positions.

3. A protein function annotation device based on a pre-trained large language model, characterized in that: The device comprises: a category determination module, a location identification module and an annotation module; The category determination module is used to input the protein sequence to be annotated into a target protein functional domain classification model, wherein the target protein functional domain classification model includes a pre-trained protein large language model layer and a classification layer; based on the pre-trained protein large language model layer in the target protein functional domain classification model, the protein representation of the protein sequence to be annotated is determined; according to the classification layer of the target protein functional domain classification model, the functional domain category contained in the protein sequence to be annotated is predicted based on the protein representation; the target protein functional domain classification model is obtained by training a protein functional domain classification model based on the pre-trained protein large language model according to the complete protein sequence in a preset training set; the pre-trained protein large language model is used to train a protein functional domain classification model based on the pre-trained protein large language model. The core architecture of the language model is based on the Transformer encoder and is obtained by unsupervised pre-training using complete protein sequences; it is also used to set a training set and a validation set including multiple protein sequences; the training set and the validation set include multiple non-overlapping complete protein sequences; the first protein sequence in the training set is obtained, and the first protein sequence is input into the pre-trained protein large language model layer to obtain an initial category representation in the protein representation; the initial category representation is predicted by the classification layer to determine the predicted functional domain category label of the first protein sequence; the true category label of the first protein sequence in the training set is obtained, and the predicted functional domain category label and the true category label are determined according to the predicted functional domain category label and the true category label. The method comprises the steps of: determining a label loss value based on a distinguishing label; updating the protein functional domain classification model according to the label loss value; obtaining a second protein sequence in the verification set and determining a verification value of the second protein sequence; verifying the updated protein functional domain classification model according to the verification value to obtain the target protein functional domain classification model; the position recognition module is used to input the functional domain category contained in the protein sequence to be annotated and the protein sequence to be annotated into the target protein functional domain recognition model, wherein the target protein functional domain recognition model comprises: a pre-trained protein large language model layer, an interval representation layer and a functional domain representation layer; determining the protein sequence to be annotated based on the pre-trained protein large language model layer. The protein representation of the sequence; based on the functional domain representation layer, mapping the functional domain category contained in the protein sequence to be annotated to a preset latent space to obtain a latent space category representation; based on the interval representation layer, mapping the protein representation to a preset latent space to obtain a latent space position representation; determining the similarity between the latent space category representation and the latent space position representation; according to the similarity, determining the category probability of the functional domain category corresponding to the multiple functional domain positions in the protein to be annotated; determining the functional domain category whose category probability exceeds a preset probability threshold as the target functional domain category, and determining the position corresponding to the target functional domain category as the target position where the functional domain category of the protein sequence to be annotated is located;The target protein functional domain recognition model is obtained by training a protein functional domain recognition model based on a pre-trained protein large language model and a named entity recognition model according to the complete protein sequence and the corresponding protein functional domain category and position in a preset training set; the pre-trained protein large language model is a model for determining protein representations in a protein sequence; it is also used to obtain the first protein sequence in the training set, input the first protein sequence to the pre-trained protein large language model layer, and obtain the initial position representation and initial category representation of the protein sequence; determine the potential space position representation corresponding to the functional domain of the first protein sequence from the initial position representation through the interval representation layer; obtain the first protein sequence in the training set, and then input the first protein sequence to the pre-trained protein large language model layer to obtain the initial position representation and initial category representation of the protein sequence. The real category label of the first protein sequence is obtained, and the initial representation of the functional domain is determined according to the real category label; based on the initial category representation of the protein sequence, the latent space category representation corresponding to the functional domain of the first protein sequence is determined through the functional domain representation layer; the similarity between the latent space position representation and the latent space category representation is determined; the loss value is calculated according to the category probability corresponding to the similarity and the preset probability threshold, and the protein functional domain recognition model is trained and updated according to the loss value; the second protein sequence in the verification set is obtained, and the verification value of the second protein sequence is determined; the updated protein functional domain recognition model is verified according to the verification value to obtain the target protein functional domain recognition model; The annotation module is used to perform functional annotation on the protein sequence to be annotated according to the target position where the functional domain category of the protein sequence to be annotated is located.

4. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the protein function annotation method based on a pre-trained large language model according to any one of claims 1 to 2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the protein function annotation method based on a pre-trained large language model described in any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • DNA binding residue prediction method based on multi-modal protein language model

    CN119418777A

  • Identification method and system of short antibacterial peptide sequence, terminal and storage medium

    CN119541641A