Protein sequence prediction method and device, equipment and storage medium

By obtaining the attribute tags and sequence information of predicted proteins, and using the coding network and prediction network for feature encoding, the problem of insufficient accuracy of protein sequence prediction in the prior art is solved, and accurate prediction of complex functional protein sequences is achieved.

CN120260686APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410011001.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is inaccurate when predicting protein sequences, especially complex functional protein sequences, and cannot effectively utilize the inherent properties and amino acid sequence information described in natural language.

Method used

By obtaining the attribute tags and sequence information of predicted proteins, using coding networks and prediction networks for feature encoding, combining natural language and amino acid sequence constraints, and using protein sequence prediction models to predict, to achieve accurate prediction of protein sequences.

Benefits of technology

Accurate expression based on natural language is realized, protein sequences with complex functions can be predicted, and the accuracy and accuracy of prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260686A_ABST
    Figure CN120260686A_ABST
Patent Text Reader

Abstract

The invention discloses a protein sequence prediction method and device, equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring an attribute tag of a predicted protein, wherein the attribute tag indicates inherent properties of the predicted protein based on a natural language; obtaining sequence information of the predicted protein, wherein the sequence information is used for constraining an amino acid sequence in the predicted protein; and calling a protein sequence prediction model to perform sequence prediction on the attribute tag and the sequence information to obtain a protein sequence which is an amino acid sequence of the predicted protein. The attribute tag of the predicted protein accurately expresses the inherent property of the predicted protein based on a natural language, and the sequence information restrains the amino acid sequence which needs to be included in the predicted protein; the protein sequence of the predicted protein is predicted by calling the protein sequence prediction model, and the protein sequence with complex properties indicated by the attribute tag is predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, apparatus, device, and storage medium for predicting protein sequences. Background Art

[0002] There is a rich variety of proteins, with significant differences in the physicochemical properties of different proteins and different functions in organisms.

[0003] In the related art, by clustering types and amino acid sequences of known proteins, a predicted protein with physicochemical properties similar to those of the known protein and including an amino acid sequence is predicted to achieve the prediction of protein sequences.

[0004] However, the above prediction method is relatively single and cannot accurately predict protein sequences with complex functions. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for predicting protein sequences, and the technical solutions are as follows:

[0006] According to one aspect of this application, a method for predicting protein sequences is provided. The method includes:

[0007] Obtain an attribute label of a predicted protein, where the attribute label indicates an inherent property of the predicted protein based on natural language;

[0008] Obtain sequence information of the predicted protein, where the sequence information is used to constrain the amino acid sequence in the predicted protein;

[0009] Call a protein sequence prediction model to perform sequence prediction on the attribute label and the sequence information to obtain a protein sequence, where the protein sequence is the amino acid sequence of the predicted protein, the protein sequence includes the sequence information, and the predicted protein has the inherent property indicated by the attribute label.

[0010] According to another aspect of this application, a device for predicting protein sequences is provided. The device includes:

[0011] An obtaining module, configured to obtain an attribute label of a predicted protein, where the attribute label indicates an inherent property of the predicted protein based on natural language;

[0012] The obtaining module is further configured to obtain sequence information of the predicted protein, where the sequence information is used to constrain the amino acid sequence in the predicted protein;

[0013] A processing module, configured to call a protein sequence prediction model to perform sequence prediction on the attribute label and the sequence information, so as to obtain a protein sequence, where the protein sequence is the amino acid sequence of the predicted protein, the protein sequence includes the sequence information, and the predicted protein has the inherent property indicated by the attribute label.

[0014] In an alternative design of the present application, the obtaining module is further configured to perform at least one of the following:

[0015] Obtain a species label, where the species label is used to indicate the biological species to which the predicted protein belongs;

[0016] Obtain a function label, where the function label is used to indicate the protein function of the predicted protein in the biological species to which it belongs;

[0017] Obtain a structure label, where the structure label is used to indicate the geometric structure of the predicted protein in three-dimensional space;

[0018] Obtain a task label, where the task label is used to indicate to the protein sequence prediction model the prediction method of the protein sequence.

[0019] In an alternative design of the present application, the obtaining module is further configured to:

[0020] Determine the category name of at least one level of category in the biological classification of the biological species to which the predicted protein belongs as the species label;

[0021] Wherein, the biological classification includes at least one level of category among domain, kingdom, phylum, class, order, family, genus, and species.

[0022] In an alternative design of the present application, the obtaining module is further configured to:

[0023] Determine the function label according to the protein function of the predicted protein in at least one description dimension of cellular component, biological process, and molecular function.

[0024] In an alternative design of the present application, the obtaining module is further configured to:

[0025] Determine the category name of at least one level of category in the structure classification of the predicted protein as the structure label;

[0026] Wherein, the structure classification includes at least one level of category among class, fold, superfamily, and family.

[0027] In an alternative design of the present application, the obtaining module is further configured to:

[0028] Concatenate the reference sequence and the associated label to obtain the task label;

[0029] Among them, the reference sequence is the amino acid sequence of a known protein, and the associated tag is used to indicate whether there is an affinity between the predicted protein and the known protein.

[0030] In an alternative design of the present application, the protein sequence prediction model includes an encoding network and a prediction network; the processing module is further configured to:

[0031] Call the encoding network to perform feature encoding on the attribute tag and the sequence information respectively, to obtain a tag identifier and a sequence identifier, where the tag identifier is the unique identifier of the attribute tag, and the sequence identifier is the unique identifier of the amino acid sequence;

[0032] Call the prediction network to perform sequence prediction on the tag identifier and the sequence identifier, to obtain a protein sequence.

[0033] In an alternative design of the present application, the encoding network includes a first sub-network and a second sub-network arranged in parallel;

[0034] The processing module is further configured to:

[0035] Call the first sub-network to perform feature encoding on the attribute tag, to obtain the tag identifier;

[0036] Call the second sub-network to perform feature encoding on the sequence information, to obtain the sequence identifier;

[0037] Among them, the network parameters of the first sub-network and the second sub-network are different.

[0038] In an alternative design of the present application, the first sub-network is a dictionary encoder; the processing module is further configured to:

[0039] Call the dictionary encoder to find at least two word identifiers corresponding one-to-one to at least two phrases in the attribute tag, where the at least two phrases in the attribute tag are phrases divided according to semantics;

[0040] Based on the positions of the at least two phrases in the attribute tag, splice the at least two word identifiers to obtain the tag identifier.

[0041] In an alternative design of the present application, the second sub-network is a word segmenter; the processing module is further configured to:

[0042] Call the word segmenter to perform feature encoding on the sequence information, to obtain the sequence identifier, where the sequence identifier includes at least two sub-identifiers corresponding one-to-one to at least two sub-sequences, and the at least two sub-sequences are obtained by splitting the sequence information based on the word segmenter.

[0043] In an alternative design of the present application, the obtaining module is further configured to obtain the sample sequence identifier and the sample label identifier of the sample protein;

[0044] The processing module is further configured to sample from the sample sequence identifier and the sample label identifier to obtain a training identifier;

[0045] The processing module is further configured to call the initial prediction network to perform sequence prediction on the training identifier to obtain a first protein sequence;

[0046] The apparatus further includes:

[0047] A training module, configured to train the initial prediction network based on the difference between the first protein sequence and the sample sequence identifier to obtain the prediction network.

[0048] In an alternative design of the present application, the processing module is further configured to:

[0049] Obtain the sample sequence and the sample identifier of the sample protein;

[0050] Call the encoding network to perform feature encoding on the sample sequence and the sample identifier respectively to obtain a sample sequence identifier and a sample label identifier;

[0051] The training module is further configured to:

[0052] Calculate the loss function between the first protein sequence and the sample sequence, and train the initial prediction network based on the loss function to obtain the prediction network;

[0053] Wherein, the protein sequence prediction model is obtained by splicing the encoding network and the trained prediction network.

[0054] In an alternative design of the present application, the processing module is further configured to call the prediction network to perform label prediction on the sample sequence identifier to obtain a predicted label;

[0055] The training module is further configured to supplement and train the prediction network based on the difference between the predicted label and the sample label identifier;

[0056] Wherein, the protein sequence prediction model is obtained by splicing the encoding network and the prediction network after supplementary training.

[0057] According to another aspect of the present application, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the protein sequence prediction method described in the above aspect.

[0058] According to another aspect of the present application, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the readable storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the protein sequence prediction method described in the above aspect.

[0059] According to another aspect of the present application, a computer program product is provided. The computer program product includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the protein sequence prediction method described in the above aspect.

[0060] The beneficial effects brought by the technical solution provided by the present application at least include:

[0061] Obtain the attribute label of the predicted protein, which accurately expresses the inherent properties of the predicted protein based on natural language, and the sequence information constrains the amino acid sequence to be included in the predicted protein; by invoking the protein sequence prediction model, the protein sequence of the predicted protein including the sequence information and having the inherent properties indicated by the attribute label is predicted based on the attribute label and the sequence information, realizing the prediction of the protein sequence with the complex properties indicated by the attribute label. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 is a schematic diagram of a computer system provided by an exemplary embodiment of the present application;

[0064] Figure 2 is a schematic diagram of a protein sequence prediction method provided by an exemplary embodiment of the present application;

[0065] Figure 3 is a flowchart of a protein sequence prediction method provided by an exemplary embodiment of the present application;

[0066] Figure 4 is a flowchart of a method for predicting a protein sequence provided by an exemplary embodiment of the present application;

[0067] Figure 5 is a schematic diagram of biological classification provided by an exemplary embodiment of the present application;

[0068] Figure 6 is a schematic diagram of structural classification provided by an exemplary embodiment of the present application;

[0069] Figure 7 is a schematic diagram of the spatial structure of triosephosphate isomerase provided by an exemplary embodiment of the present application;

[0070] Figure 8 is a flowchart of a method for predicting a protein sequence provided by an exemplary embodiment of the present application;

[0071] Figure 9 is a flowchart of a method for predicting a protein sequence provided by an exemplary embodiment of the present application;

[0072] Figure 10 is a structural block diagram of a device for predicting a protein sequence provided by an exemplary embodiment of the present application;

[0073] Figure 11 is a structural block diagram of a server provided by an exemplary embodiment of the present application.

[0074] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Detailed Description of the Embodiments

[0075] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0078] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, information such as the input attribute tags of predicted proteins involved in this application is obtained under full authorization.

[0079] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0080] First, some terms in this application are introduced.

[0081] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0082] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0083] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0084] The solution provided in the embodiments of this application relates to technologies such as the prediction of protein sequences in artificial intelligence, and will be specifically described through the following embodiments.

[0085] Figure 1 The figure shows a schematic diagram of a computer system provided by an embodiment of this application. This computer system can be implemented as the system architecture of a protein sequence prediction method. This computer system can include: a terminal 100 and a server 200.

[0086] The terminal 100 can be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a PC (Personal Computer), etc. A client for running a target application program can be installed in the terminal 100. The target application program can be a protein sequence prediction application program or other application programs that provide the function of predicting protein sequences. This application does not make any limitations in this regard. In addition, this application does not make any limitations on the form of the target application program, including but not limited to Apps (Application programs) installed in the terminal 100, applets, etc., and can also be in the form of a web page.

[0087] The server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The server 200 may be the background server of the above target application, and is used to provide background services for the client of the target application.

[0088] In the protein sequence prediction method provided by the embodiments of the present application, the execution subject of each step may be a computer device, and the computer device refers to an electronic device with data calculation, processing and storage capabilities. Taking Figure 1 the solution implementation environment shown as an example, the protein sequence prediction method may be executed by the terminal 100 (for example, the client of the target application installed and running in the terminal 100 executes the protein sequence prediction method), or may be executed by the server 200, or may be executed by the interaction and cooperation of the terminal 100 and the server 200. The present application does not make any limitations in this regard.

[0089] In addition, the technical solution of the present application may be combined with blockchain technology. For example, in the protein sequence prediction method disclosed in the present application, some of the data involved (such as the direction information of the first road and the second road) may be saved on the blockchain. The terminal 100 and the server 200 may communicate with each other through a network, such as a wired or wireless network.

[0090] Next, the image translation model in the present application will be introduced:

[0091] Figure 2 FIG. shows a schematic diagram of a protein sequence prediction method provided by an embodiment of the present application.

[0092] The protein sequence prediction model includes an encoding network 410 and a prediction network 420. The protein sequence prediction model has the ability to obtain the protein sequence of the predicted protein according to the attribute label 302 and the sequence information 304.

[0093] Exemplarily, the attribute label 302 and the sequence information 304 are obtained. The attribute label indicates the inherent properties of the predicted protein based on natural language, and the sequence information is used to indicate the amino acid sequence included in the predicted protein.

[0094] Specifically, the attribute label includes at least one of the following four labels:

[0095] · The species label 302a is used to indicate the biological species to which the predicted protein belongs, such as at least one level of category in the domain, kingdom, phylum, class, order, family, genus, and species in the biological classification of the biological species;

[0096] · The functional label 302b is used to indicate the protein function of the predicted protein in the organism to which it belongs, such as the protein function of the predicted protein in at least one of the description dimensions of cellular components, biological processes, and molecular functions;

[0097] · The structural label 302c is used to indicate the geometric structure of the predicted protein in three-dimensional space, such as at least one level of classification in the structure classification of the predicted protein, including class, fold, superfamily, and family;

[0098] · The task label 302d is used to indicate the prediction method of the protein sequence to the protein sequence prediction model. For example, the task label 302d indicates to the protein sequence prediction model the affinity between the predicted protein and the known protein.

[0099] The encoding network 410 includes a dictionary encoder 412 and a tokenizer 414;

[0100] Call the dictionary encoder 412 to find at least two word identifiers corresponding one-to-one to at least two phrases in the attribute label 302. The at least two phrases in the attribute label 302 are phrases divided according to the semantics of natural language;

[0101] Based on the positions of the at least two phrases in the attribute label 302, splice the at least two word identifiers to obtain a label identifier 312. The label identifier 312 is the unique identifier of the attribute label 302. The label identifier 312 includes one or more characters, but the label identifier 312 usually does not have the semantics of natural language.

[0102] Call the tokenizer 414 to perform feature encoding on the sequence information 304 to obtain a sequence identifier 314. The sequence identifier 314 includes at least two sub-identifiers corresponding one-to-one to at least two subsequences. The at least two subsequences are obtained by splitting the sequence information 304 based on the tokenizer 414. The subsequence includes a sequence of at least one amino acid.

[0103] Call the prediction network 420 to perform sequence prediction on the label identifier 312 and the sequence identifier 314 to obtain a protein sequence 320. The protein sequence 320 is the amino acid sequence of the predicted protein. The protein sequence 320 includes the sequence information 304, and the predicted protein has the functional characteristics indicated by the attribute label 302.

[0104] Input the protein sequence 320 output by the prediction network 420 into the structure prediction model 430, predict the corresponding three-dimensional structure according to the protein sequence 320, and obtain the protein structure 330 of the predicted protein. The protein structure 330 is the three-dimensional solid structure of the predicted protein in three-dimensional space.

[0105] Figure 3The flowchart of the prediction method for the protein sequence provided by an exemplary embodiment of the present application is shown. This method can be executed by a computer device. The method includes:

[0106] Step 510: Obtain the attribute tags of the predicted protein;

[0107] The attribute tags constrain the predicted protein in the dimension of protein attributes, and the predicted protein is the protein expected to be obtained. Exemplarily, the attribute tags indicate the inherent properties of the predicted protein based on natural language.

[0108] The inherent properties indicated by the attribute tags can be the properties of the predicted protein itself, or can indicate the association relationship between the predicted protein and other organic substances (such as the organism to which the predicted protein belongs, other proteins that are affinity with the predicted protein, etc.), or can also indicate the inherent properties of other organic substances that have an association relationship with the predicted protein. Specific examples will be given below.

[0109] Exemplarily, the attribute tags based on natural language carry semantic information, which realizes the accurate description of the inherent properties of the predicted protein. Further, in the case where it is desired that the predicted protein has complex functions, multiple attribute tags can be used to describe it independently from different dimensions, and it can be realized that the inherent properties of the predicted protein indicated by the attribute tags are different from those of known proteins.

[0110] It can be seen that the number of attribute tags can be one or more; the acquisition method of the attribute tags is usually obtained based on human-computer interaction operations, but other acquisition methods are not excluded.

[0111] Step 520: Obtain the sequence information of the predicted protein;

[0112] The sequence information is used to constrain the amino acid sequence in the predicted protein; the amino acid sequence in the sequence information is the amino acid sequence that needs to be included in the predicted protein. The sequence information includes at least one amino acid, and the sequence information is obtained by arranging and combining at least one kind of amino acid. In one example, the sequence information is a sequence including at least two amino acids, and the sequence information is used to constrain that there is a sub-part in the protein sequence of the predicted protein that is the same as the sequence information.

[0113] Exemplarily, step 510 can be executed before, after or simultaneously with step 520, and the present application does not limit the execution timing between step 510 and step 520.

[0114] Step 530: Invoke the protein sequence prediction model to perform sequence prediction on the attribute tags and the sequence information to obtain the protein sequence;

[0115] Using the attribute label and sequence information as input parameters of the protein sequence prediction model, performing sequence prediction based on the protein sequence prediction model, and outputting a protein sequence. By constraining the prediction with the attribute label for the inherent properties of the protein and with the sequence information for a partial amino acid sequence in the protein, it indicates to the protein sequence prediction model the way to predict the protein sequence.

[0116] The protein sequence prediction model is a computational model with the ability to predict protein sequences; in one example, the protein sequence prediction model is an artificial neural network (ANN) model based on artificial intelligence. The present application does not limit the model structure of the protein sequence prediction model, and the protein sequence prediction model may include at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), and a generative adversarial network (GAN).

[0117] The protein sequence is the amino acid sequence of the predicted protein. The protein sequence obtained by the protein sequence prediction model includes sequence information, and the predicted protein corresponding to the protein sequence has the inherent properties indicated by the attribute label.

[0118] In summary, the method provided in this embodiment obtains the attribute label of the predicted protein, accurately expresses the inherent properties of the predicted protein based on natural language, and the sequence information constrains the amino acid sequence to be included in the predicted protein; by invoking the protein sequence prediction model, based on the attribute label and sequence information, a protein sequence of the predicted protein including sequence information and having the inherent properties indicated by the attribute label is predicted, realizing the prediction of a protein sequence with complex properties indicated by the attribute label.

[0119] Figure 4 The flowchart of the method for predicting a protein sequence provided by an exemplary embodiment of the present application is shown. This method can be executed by a computer device. That is, in Figure 3 the embodiment shown, step 510 can be implemented as at least one of step 512, step 514, step 516, and step 518:

[0120] Step 512: Obtain a species label;

[0121] In this step, the species label is determined as the attribute label of the predicted protein, and the species label is used to indicate the inherent properties of the organism associated with the predicted protein; specifically, the species label is used to indicate the biological species to which the predicted protein belongs; the above biological species in nature includes the predicted protein, for example, the biological species has the ability to synthesize the predicted protein.

[0122] In an optional example, step 512 in this embodiment can be implemented as:

[0123] Determine the category name of at least one level of the biological classification of the biological species to which the predicted protein belongs as the species label;

[0124] Exemplarily, the biological classification includes at least one level of categories among Domain, Kingdom, Phylum, Class, Order, Family, Genus, and Species.

[0125] Taking the biological species American black bear as an example, Figure 5 shows a schematic diagram of biological classification provided by an exemplary embodiment of the present application. The category names of the American black bear in the first level to the eighth level of categories (Domain, Kingdom, Phylum, Class, Order, Family, Genus, Species) are respectively: Eukarya 601, Animalia 602, Chordata 603, Mammalia 604, Carnivora 605, Ursidae 606, Ursus 607, Ursus Americanus 608. It can be seen that as the number of category levels of the biological classification increases, the number of species belonging to the same category decreases, and the classification information of the biological species can be represented more accurately.

[0126] In an example, obtain the category name of the i-th level of the biological species; based on the biological classification library, find at least one level of the high-level category names of the biological classification in the first level to the i-1 level to which the category name of the i-th level belongs, and splice the category name of the i-th level and at least one level of the high-level category names to obtain the species label. i is an integer greater than 1.

[0127] Among them, the biological classification library carries the tree-like biological classification information of at least two levels among the above first level to eighth level of categories of the biological classification. For example, Figure 5For example, the class name of the i-th level of a biological species is the class Mammalia, and the high-level class names of the class name include at least one of the domain Eukaryota, the kingdom Animalia, and the phylum Chordata. The class name of the i-th level (Mammalia) and the high-level class name are concatenated to obtain the species label of the biological species to which the predicted protein belongs.

[0128] Step 514: Obtain a functional label;

[0129] In this step, the functional label is determined as the attribute label of the predicted protein, and the functional label is used to indicate the inherent properties of the predicted protein itself; specifically, the functional label is used to indicate the protein function of the predicted protein in the organism to which it belongs.

[0130] In an optional example, step 514 in this embodiment can be implemented as:

[0131] Determine the functional label according to the protein function of the predicted protein in at least one description dimension of cellular component, biological process, and molecular function;

[0132] Exemplarily, Cellular Component, Biological Process, and Molecular Function indicate the function of the predicted protein in the organism to which it belongs from three different description dimensions.

[0133] Specifically, the protein function in the description dimension of cellular component is used to indicate that the predicted protein is a component of the cellular component. A cellular component is the basic unit that constitutes a cell; for example, cellular components include at least one of extracellular exosome, focal adhesion, perinuclear region of cytoplasm, melanosome, cadherin binding, cytoplasm, membrane, vacuolar membrane, and cytosol.

[0134] The protein functions in the biological process description dimension are used to indicate the biological processes that the predicted proteins are involved in. Biological processes are a series of molecular events that occur in an organism with the goal of completing life activities; for example, biological processes include at least one of signal transduction, protein targeting, negative regulation of G protein-coupled receptor signaling pathway, histone deacetylase binding, protein kinase inhibitor activity, cytoplasmic sequestering of protein, negative regulation of protein dephosphorylation, and positive regulation of catalytic activity.

[0135] The protein functions in the molecular function description dimension are used to indicate the functions or activities that the predicted proteins undertake in cells or organisms; for example, at least one of enzyme binding, identical protein binding, protein domain specific binding, phosphoserine residue binding, phosphoprotein binding, histone deacetylase binding, and cadherin binding.

[0136] Exemplarily, the protein functions in the above three description dimensions of cellular components, biological processes, and molecular functions can be independent of each other. In one example, a predicted protein has multiple functional tags, which constrain the protein function of the predicted protein from different description dimensions. For example, the predicted protein is a component of the membrane in the cellular component dimension, participates in signal transduction in the biological process dimension, and the specific molecular function is to bind to proteins, and the biological process of signal transduction is achieved through protein binding. Further, the predicted protein can include multiple tags in the same description dimension to constrain the protein function of the predicted protein. For example, in the cellular component description dimension, the predicted protein is both a component of the membrane and a component of the vacuolar membrane.

[0137] Step 516: Obtain a structure tag;

[0138] In this step, the structure tag is determined as the attribute tag of the predicted protein, and the structure tag is used to indicate the inherent properties of the predicted protein itself; specifically, the structure tag is used to indicate the geometric structure of the predicted protein in three-dimensional space.

[0139] In an alternative example, step 516 in this embodiment can be implemented as:

[0140] Determine the category name of at least one level of category in the structure classification of the predicted protein as the structure tag;

[0141] Exemplarily, the structure classification includes at least one level of category among Class, Fold, Superfamily, and Family.

[0142] Figure 6 Shows a schematic diagram of the structure classification provided by an exemplary embodiment of the present application. The protein structure classification includes the first-level category to the fourth-level category (Class, Fold, Superfamily, Family).

[0143] The Structural Classification of Proteins (SCOP) is the root node 610, and the protein structure classification includes four classes (Class) 611: α class, β class, α / β class, and α+β class.

[0144] Taking the α / β class as an example for illustration, the α / β class includes three folds (Fold) 612: Rossmann fold, Flavodoxin-like, and α / β Barrel.

[0145] Taking the α / β barrel folding mode in the α / β class as an example for illustration, the α / β barrel folding mode includes four superfamilies 613: triosephosphate isomerase (TIM), tryptophan biosynthesis, glycosyltransferase, ribulose-1,5-bisphosphate carboxylase / oxygenase (RuBisCo).

[0146] Taking the glycosyltransferase superfamily of the α / β barrel folding mode in the α / β class as an example for illustration, the glycosyltransferase superfamily includes four families 614: β-galactosidase, β-glucanase, α-amylase, β-amylase.

[0147] It can be seen that as the level of the category in the structure classification increases, the granularity of the structure classification of the predicted protein becomes finer, and the geometric structure of the predicted protein in three-dimensional space can be represented more accurately. Figure 7 The schematic diagram of the spatial structure of triosephosphate isomerase provided by an exemplary embodiment of the present application is shown. Figure 7 A spatial structure of triosephosphate isomerase belonging to the α / β barrel folding mode in the α / β class is shown. It should be noted that triosephosphate isomerase is a superfamily in the α / β barrel folding mode. Figure 7 Only one of the spatial structures is shown, but it does not exclude that the triosephosphate isomerase superfamily has other spatial structures.

[0148] In one example, obtain the category name of the j-th level category of the structure classification; based on the structure classification library, find at least one high-level category name of the structure classification in the first level to the j-1 level categories to which the category name of the j-th level category belongs, and splice the category name of the j-th level category and at least one high-level category name to obtain a structure label. j is an integer greater than 1.

[0149] Among them, the structure classification library carries the tree-like structure classification information of at least two levels of categories from the first level category to the fourth level category of the structure classification. For example, Figure 6 Taking it as an example, four folding modes under the α / β class, four superfamilies under the α / β barrel folding mode, and four families under the glycosyltransferase superfamily are shown.

[0150] Step 518: Obtain a task label;

[0151] In this step, the task label is determined as the attribute label of the predicted protein, and the task label is used to indicate the association relationship between the predicted protein and other organic substances; specifically, the task label is used to indicate to the protein sequence prediction model the prediction method of the protein sequence.

[0152] In an alternative example, step 518 in this embodiment can be implemented as:

[0153] Concatenate the reference sequence and the association label to obtain the task label;

[0154] Exemplarily, the reference sequence is the amino acid sequence of a known protein, and the association label is used to indicate whether the predicted protein and the known protein are affinity. In one example, the task label includes the following characters: <protein>KRWIILGLNK <bind><:>. Correspondingly, the natural language semantics corresponding to the task label is: Given the amino acid sequence of a known protein, what is the protein that binds to it. Among them, <protein>For indicating that the amino acid sequence of the protein is KRWIILGLNK; specifically, KRWIILGLNK is the amino acid sequence of a known protein, <bind>For indicating the affinity between a predicted protein and a known protein. In another example, the associated tag is <concat>, the associated tag is used to indicate that the predicted protein is obtained by combining at least two known proteins.

[0155] In summary, for the method provided in this embodiment, the attribute tags include at least one of species tags, function tags, structure tags, and task tags; a way of describing the inherent attributes of the predicted protein from different dimensions based on natural language is provided; it is possible to expect to predict a predicted protein with complex properties through combinations of tags in different dimensions; compared with the clustering tags in the related art, the attribute tags do not depend on known proteins and restrict the predicted protein from the granularity of the inherent properties; by invoking the protein sequence prediction model, a protein sequence of the predicted protein including sequence information and having the inherent properties indicated by the attribute tags is predicted based on the attribute tags and sequence information, realizing the prediction of a protein sequence with complex properties indicated by the attribute tags.

[0156] Next, the protein sequence prediction model in step 530 will be introduced.

[0157] Figure 8 FIG. shows a flowchart of a method for predicting a protein sequence provided by an exemplary embodiment of the present application. This method can be executed by a computer device. That is, in Figure 3 the embodiment shown, step 530 can be implemented as step 532 and step 534:

[0158] Step 532: Invoke an encoding network to perform feature encoding on the attribute tags and sequence information respectively to obtain a tag identifier and a sequence identifier;

[0159] In this embodiment, the protein sequence prediction model includes an encoding network and a prediction network, which will be introduced separately in two steps of this embodiment.

[0160] The encoding network is used to perform feature encoding on the attribute tags and sequence information respectively. The encoding methods of the encoding network for the above two types of information can be the same or different. The tag identifier obtained by performing feature encoding is the unique identifier of the attribute tag, and the sequence identifier is the unique identifier of the amino acid sequence.

[0161] Taking the attribute tag as an example, the encoding network can encode the attribute tag based on a convolutional method and determine the hidden layer feature of the encoded attribute tag as the unique identifier of the attribute tag; it can also encode the attribute tag by invoking the encoding network to find the corresponding unique identifier of the attribute tag.

[0162] Exemplarily, the feature encoding performed by the encoding network on the attribute tags and sequence information respectively is independent; that is to say, the encoding process of the attribute tag has nothing to do with the sequence information, but it does not exclude that the attribute tag and sequence information perform feature encoding simultaneously in terms of time sequence.

[0163] In an alternative implementation, the encoding network includes a first sub-network and a second sub-network arranged in parallel. Step 530 in this embodiment can be implemented as the following two sub-steps:

[0164] Sub-step 1: Invoke the first sub-network to perform feature encoding on the attribute label to obtain a label identifier;

[0165] Exemplarily, the first sub-network and the second sub-network are two independent sub-parts in the encoding network; the network parameters of the first sub-network and the second sub-network are different. Feature encoding is performed on the attribute label and the sequence information respectively based on different sub-networks, fully considering that the attribute label is an inherent property for predicting proteins based on natural language descriptions, while the sequence information is an amino acid sequence and does not have the specific semantics in natural language. Different types of input information are feature-encoded using different network parameters.

[0166] Furthermore, the first sub-network is a dictionary encoder; sub-step one can be implemented as:

[0167] Invoke the dictionary encoder to find at least two word identifiers corresponding to at least two phrases in the attribute label; based on the positions of the at least two phrases in the attribute label, concatenate the at least two word identifiers to obtain a label identifier;

[0168] Exemplarily, at least two phrases in the attribute label are phrases divided according to semantics; the dictionary encoder carries the word identifiers of the above at least two phrases, and by looking up the at least two phrases in the dictionary encoder, at least two word identifiers corresponding to the at least two phrases are obtained.

[0169] Exemplarily, a phrase includes at least two characters, and the word identifier corresponding to the phrase carries the context semantic information between multiple characters, and the overall semantic information of the phrase can be obtained in combination with the context. Exemplarily, according to the positions of the at least two phrases in the attribute label, the at least two word identifiers are concatenated; ensuring the correct word order between the phrases.

[0170] Compared with the encoding method of each character, the dictionary encoder provides an encoding method with the phrase as the granularity, retaining the context information in the natural semantics. It fully considers that the attribute label usually consists of biological terms and there are common proper nouns, such as phrases like "glucose", "amylase", etc. By forming phrases from multiple characters, the dictionary encoder carries the corresponding relationship between the phrases and the word identifiers, reducing the number of corresponding relationships that the dictionary encoder needs to record.

[0171] Sub-step 2: Invoke the second sub-network to perform feature encoding on the sequence information to obtain a sequence identifier;

[0172] Exemplarily, the first sub-network and the second sub-network are two independent sub-parts in the encoding network; the network parameters of the first sub-network and the second sub-network are different.

[0173] Furthermore, the second sub-network is a tokenizer; sub-step two can be implemented as:

[0174] Call the tokenizer to perform feature encoding on the sequence information to obtain sequence identifiers;

[0175] Exemplarily, the sequence identifiers include at least two sub-identifiers corresponding to at least two sub-sequences one by one. The at least two sub-sequences are obtained by splitting the sequence information based on the tokenizer, and the tokenizer has the ability to split the sequence information. It should be noted that a sub-sequence is a sequence composed of at least two amino acids, and a sub-sequence is an amino acid arrangement sequence existing in a known protein.

[0176] It should be noted that there are twenty kinds of amino acids in nature. Proteins are organic compounds with large molecular weights and can be composed of dozens to thousands of amino acids. The sequence information constrains the amino acid sequence that needs to exist in the predicted protein; in some examples, the sequence information can be composed of dozens to hundreds of amino acids; encoding the sequence composed of multiple amino acids into one sub-identifier reduces the length of the sequence identifier compared to encoding each amino acid one by one.

[0177] Step 534: Call the prediction network to perform sequence prediction on the label identifier and the sequence identifier to obtain the protein sequence;

[0178] The prediction network predicts the protein sequence based on the input label identifier and sequence identifier; the prediction network jointly predicts the protein sequence from the label identifier and the sequence identifier. Among them, the sequence identifier provides a constraint on the information of the amino acid sequence when constructing the predicted protein, and a sub-part of the protein sequence to be predicted is the same as the sequence identifier. The label identifier provides a constraint on the inherent properties when constructing the predicted protein, and the predicted protein has the inherent properties indicated by the label identifier.

[0179] Exemplarily, the prediction network is usually an Artificial Neural Network (ANN), and the model structure of the prediction network is not limited in this application.

[0180] In summary, the method provided in this embodiment obtains the attribute tags of the predicted protein, accurately expresses the inherent properties of the predicted protein based on natural language, and the sequence information constrains the amino acid sequence to be included in the predicted protein; by calling the encoding network, the tag identifier and the sequence identifier are encoded, reducing the number of characters input to the prediction network and the computational complexity of the prediction network. The predicted protein sequence includes sequence information and has the inherent properties indicated by the attribute tags, realizing the prediction of a protein sequence with complex properties indicated by the attribute tags.

[0181] Next, the training process of the protein sequence prediction model will be introduced.

[0182] Figure 9 FIG. shows a flowchart of a method for predicting a protein sequence provided by an exemplary embodiment of the present application. This method can be executed by a computer device. That is, on the basis of the embodiment shown in Figure 8 the embodiment shown, it further includes steps 501, 502, and 503:

[0183] Step 501: Obtain the sample sequence identifier and the sample tag identifier of the sample protein, and sample to obtain a training identifier from the sample sequence identifier and the sample tag identifier;

[0184] The sample protein is a known protein. The sample sequence identifier is the unique identifier corresponding to the sequence of each amino acid in the sample protein, and the sample tag identifier is the unique identifier of the inherent properties of the sample protein.

[0185] The training identifier is sampled from the sample sequence identifier and the sample tag identifier. The training identifier indicates some characteristics of the sample protein. Using the training tag identifier as the input parameter of the initial prediction network can realize the prediction of the first protein sequence based on the initial prediction network.

[0186] In an optional implementation manner, the obtaining method of the sample sequence identifier and the sample tag identifier can be implemented as:

[0187] Obtain the sample sequence and the sample identifier of the sample protein;

[0188] Call the encoding network to perform feature encoding on the sample sequence and the sample identifier respectively to obtain the sample sequence identifier and the sample tag identifier;

[0189] Correspondingly, the training of the initial prediction network in this embodiment is implemented as:

[0190] Calculate the loss function between the first protein sequence and the sample sequence, and train the initial prediction network based on the loss function to obtain the prediction network;

[0191] The sample protein is a known protein. The sample sequence identifier and sample label identifier of the sample protein are encoded based on an encoding network for the sample sequence and sample identifier. During the training process of the protein sequence prediction model, only the network parameters of the initial prediction network are adjusted to train the initial prediction network. The prediction ability of the prediction network is trained based on the backward error propagation method, which helps the prediction network make full use of the constraint information provided by the sample sequence identifier and sample label identifier to predict the first protein sequence.

[0192] Correspondingly, the protein sequence prediction model is obtained by splicing an encoding network and a trained prediction network. By training the prediction network in the protein sequence prediction model, the accuracy of the protein sequence prediction model in predicting protein sequences is improved, and it can obtain protein sequences based on the input attribute labels and sequence information.

[0193] Step 502: Invoke the initial prediction network to perform sequence prediction on the training identifier to obtain the first protein sequence;

[0194] The training identifier is sampled from the sample sequence identifier and sample label identifier. The training identifier indicates some characteristics of the sample protein. Using the training label identifier as the input parameter of the initial prediction network can realize predicting the first protein sequence based on the initial prediction network.

[0195] Step 503: Train the initial prediction network based on the difference between the first protein sequence and the sample sequence identifier to obtain the prediction network;

[0196] Exemplarily, the network parameters in the initial prediction network are initial parameters. Based on the difference between the first protein sequence and the sample sequence identifier, the network parameters in the initial prediction network are adjusted to train the prediction network. Exemplarily, the process of training the initial prediction network is also called the machine learning process.

[0197] In an alternative implementation, after performing step 503 to train the initial prediction network to obtain the prediction network, additional training needs to be performed on the prediction network, that is, it further includes:

[0198] Invoke the prediction network to perform label prediction on the sample sequence identifier to obtain the predicted label;

[0199] Supplementarily train the prediction network based on the difference between the predicted label and the sample label identifier;

[0200] Correspondingly, the protein sequence prediction model is obtained by splicing an encoding network and the prediction network after supplementary training.

[0201] Exemplarily, through supplementary training, the prediction network is called to perform label prediction, that is, the predicted label of the sample protein is obtained according to the sample sequence. The network parameters of the prediction network can be further adjusted to abstract more protein-related knowledge into the model parameters from the perspective of mapping sequence information to property labels.

[0202] In summary, the method provided in this embodiment trains the initial prediction network based on the sample protein to obtain the prediction network, improves the prediction accuracy of the prediction network for protein sequences, and ensures that the predicted protein can meet the constraints of the attribute label and sequence information; only trains the prediction network, reduces the number of network parameters to be adjusted, and focuses on training the prediction ability of the prediction network on the basis that the coding network provides correct identification information, thereby improving the efficiency of model training.

[0203] Furthermore, the attribute label of the predicted protein is obtained, which accurately expresses the inherent properties of the predicted protein in natural language, and the sequence information constrains the amino acid sequence to be included in the predicted protein; by calling the protein sequence prediction model, the protein sequence of the predicted protein including the sequence information and having the inherent properties indicated by the attribute label is predicted based on the attribute label and the sequence information, realizing the prediction of the protein sequence with complex properties indicated by the attribute label.

[0204] In one example, in response to a selection operation on at least one candidate attribute label in the set of attribute labels, the attribute label is obtained;

[0205] The set of attribute labels is a set of at least two known candidate attribute labels; in one example, at least two candidate attribute labels in the set of attribute labels are displayed on the first interface; each candidate attribute label corresponds to a selection control one by one; the selection operation on at least one candidate attribute label is a triggering operation on the selection control corresponding to at least one candidate attribute label.

[0206] Similar to the attribute label, in response to a selection operation on at least one candidate sequence in the set of sequence information, the sequence information is obtained;

[0207] The set of sequence information is a set of at least two known candidate sequences; wherein, each candidate sequence includes at least two amino acids. Further, the set of sequence information also includes the known twenty kinds of amino acids, and the sequence information can be spliced by inputting one or more of the above twenty kinds of amino acids.

[0208] The protein sequence prediction model is called to perform sequence prediction on the attribute label and the sequence information to obtain the protein sequence;

[0209] Using the attribute label and sequence information as input parameters of the protein sequence prediction model, performing sequence prediction based on the protein sequence prediction model, and outputting the protein sequence. Constraining the inherent properties of the predicted protein through the attribute label, and constraining part of the amino acid sequence in the predicted protein through the sequence information, indicating the way to predict the protein sequence to the protein sequence prediction model.

[0210] Inputting the protein sequence output by the protein sequence prediction model into the structure prediction model, predicting the corresponding three-dimensional structure according to the protein sequence, and obtaining the three-dimensional spatial structure of the predicted protein.

[0211] In summary, the method provided in this embodiment obtains the attribute label of the predicted protein, accurately expresses the inherent properties of the predicted protein based on natural language, and the sequence information constrains the amino acid sequence to be included in the predicted protein; by calling the protein sequence prediction model, based on the attribute label and sequence information, the protein sequence of the predicted protein including the sequence information and having the inherent properties indicated by the attribute label is predicted, realizing the prediction of the protein sequence with complex properties indicated by the attribute label.

[0212] Those of ordinary skill in the art can understand that the above embodiments can be implemented independently, or the above embodiments can be freely combined to form new embodiments to implement the protein sequence prediction method of the present application.

[0213] Figure 10 The structural block diagram of the protein sequence prediction device provided by an exemplary embodiment of the present application is shown. The device includes:

[0214] An acquisition module 810, configured to acquire the attribute label of the predicted protein, where the attribute label indicates the inherent properties of the predicted protein based on natural language;

[0215] The acquisition module 810 is further configured to acquire the sequence information of the predicted protein, where the sequence information is used to constrain the amino acid sequence in the predicted protein;

[0216] A processing module 820, configured to call a protein sequence prediction model to perform sequence prediction on the attribute label and the sequence information to obtain a protein sequence, where the protein sequence is the amino acid sequence of the predicted protein, the protein sequence includes the sequence information, and the predicted protein has the inherent properties indicated by the attribute label.

[0217] In an alternative implementation manner of this embodiment, the acquisition module 810 is further configured to perform at least one of the following:

[0218] Acquire a species label, where the species label is used to indicate the biological species to which the predicted protein belongs;

[0219] Obtain a function tag, where the function tag is used to indicate the protein function of the predicted protein in the organism to which it belongs;

[0220] Obtain a structure tag, where the structure tag is used to indicate the geometric structure of the predicted protein in three-dimensional space;

[0221] Obtain a task tag, where the task tag is used to indicate to the protein sequence prediction model the prediction method of the protein sequence.

[0222] In an alternative implementation of this embodiment, the obtaining module 810 is further configured to:

[0223] Determine the category name of at least one level of category in the biological classification of the biological species to which the predicted protein belongs as the species tag;

[0224] Wherein, the biological classification includes at least one level of category among domain, kingdom, phylum, class, order, family, genus, and species.

[0225] In an alternative implementation of this embodiment, the obtaining module 810 is further configured to:

[0226] Determine the function tag according to the protein function of the predicted protein in at least one description dimension among cellular component, biological process, and molecular function.

[0227] In an alternative implementation of this embodiment, the obtaining module 810 is further configured to:

[0228] Determine the category name of at least one level of category in the structure classification of the predicted protein as the structure tag;

[0229] Wherein, the structure classification includes at least one level of category among class, fold, superfamily, and family.

[0230] In an alternative implementation of this embodiment, the obtaining module 810 is further configured to:

[0231] Concatenate a reference sequence and an associated tag to obtain the task tag;

[0232] Wherein, the reference sequence is the amino acid sequence of a known protein, and the associated tag is used to indicate whether the predicted protein and the known protein are affinity.

[0233] In an alternative implementation of this embodiment, the protein sequence prediction model includes an encoding network and a prediction network; the processing module 820 is further configured to:

[0234] Call the encoding network to perform feature encoding on the attribute label and the sequence information respectively to obtain a label identifier and a sequence identifier. The label identifier is the unique identifier of the attribute label, and the sequence identifier is the unique identifier of the amino acid sequence;

[0235] Call the prediction network to perform sequence prediction on the label identifier and the sequence identifier to obtain a protein sequence.

[0236] In an alternative implementation of this embodiment, the encoding network includes a first sub-network and a second sub-network arranged in parallel;

[0237] The processing module 820 is further configured to:

[0238] Call the first sub-network to perform feature encoding on the attribute label to obtain the label identifier;

[0239] Call the second sub-network to perform feature encoding on the sequence information to obtain the sequence identifier;

[0240] Wherein, the network parameters of the first sub-network and the second sub-network are different.

[0241] In an alternative implementation of this embodiment, the first sub-network is a dictionary encoder; the processing module 820 is further configured to:

[0242] Call the dictionary encoder to find at least two word identifiers corresponding to at least two phrases in the attribute label, where the at least two phrases in the attribute label are phrases divided according to semantics;

[0243] Based on the positions of the at least two phrases in the attribute label, splice the at least two word identifiers to obtain the label identifier.

[0244] In an alternative implementation of this embodiment, the second sub-network is a word segmenter; the processing module 820 is further configured to:

[0245] Call the word segmenter to perform feature encoding on the sequence information to obtain the sequence identifier, where the sequence identifier includes at least two sub-identifiers corresponding to at least two sub-sequences, and the at least two sub-sequences are obtained by splitting the sequence information based on the word segmenter.

[0246] In an alternative implementation of this embodiment, the acquisition module 810 is further configured to acquire a sample sequence identifier and a sample label identifier of a sample protein;

[0247] The processing module 820 is further configured to sample training identifiers from the sample sequence identifier and the sample label identifier;

[0248] The processing module 820 is further configured to call the initial prediction network to perform sequence prediction on the training identification to obtain a first protein sequence;

[0249] The apparatus further includes:

[0250] A training module 830, configured to train the initial prediction network based on the difference between the first protein sequence and the sample sequence identification to obtain the prediction network.

[0251] In an alternative implementation of this embodiment, the processing module 820 is further configured to:

[0252] Obtain a sample sequence and a sample identification of a sample protein;

[0253] Call the encoding network to perform feature encoding on the sample sequence and the sample identification respectively to obtain a sample sequence identification and a sample label identification;

[0254] The training module 830 is further configured to:

[0255] Calculate a loss function between the first protein sequence and the sample sequence, and train the initial prediction network based on the loss function to obtain the prediction network;

[0256] Wherein, the protein sequence prediction model is obtained by splicing the encoding network and the trained prediction network.

[0257] In an alternative implementation of this embodiment, the processing module 820 is further configured to call the prediction network to perform label prediction on the sample sequence identification to obtain a predicted label;

[0258] The training module 830 is further configured to supplement and train the prediction network based on the difference between the predicted label and the sample label identification;

[0259] Wherein, the protein sequence prediction model is obtained by splicing the encoding network and the prediction network after supplementary training.

[0260] It should be noted that when the apparatus provided in the above embodiment implements its functions, only the above-mentioned division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to actual needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0261] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method; the technical effects achieved by each module in performing operations are the same as those in the embodiments related to the method, and will not be elaborated herein.

[0262] An embodiment of the present application further provides a computer device, which includes: a processor and a memory, and a computer program is stored in the memory; the processor is configured to execute the computer program in the memory to implement the protein sequence prediction method provided in each of the above method embodiments.

[0263] Optionally, the computer device is a server. Exemplarily, Figure 11 is a structural block diagram of a server provided by an exemplary embodiment of the present application.

[0264] Generally, the server 2300 includes: a processor 2301 and a memory 2302.

[0265] The processor 2301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 2301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 2301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 2301 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 2301 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0266] The memory 2302 may include one or more computer-readable storage media, which may be non-transitory. The memory 2302 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2302 is used to store at least one instruction for being executed by the processor 2301 to implement the protein sequence prediction method provided in the method embodiments of this application.

[0267] In some embodiments, the server 2300 may further optionally include: an input interface 2303 and an output interface 2304. The processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 may be connected via a bus or signal lines. Each peripheral device may be connected to the input interface 2303 and the output interface 2304 via a bus, signal lines, or a circuit board. The input interface 2303 and the output interface 2304 may be used to connect at least one peripheral device related to input / output (I / O) to the processor 2301 and the memory 2302. In some embodiments, the processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2301, the memory 2302, the input interface 2303, and the output interface 2304 may be implemented on a separate chip or circuit board, and the embodiments of this application do not limit this.

[0268] Those skilled in the art can understand that the structure shown above does not constitute a limitation on the server 2300, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.

[0269] In an exemplary embodiment, a chip is further provided, and the chip includes programmable logic circuits and / or program instructions, which are used to implement the protein sequence prediction method described in the above aspects when the chip runs on a computer device.

[0270] In an exemplary embodiment, a computer program product is further provided, and the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor reads and executes the computer instructions from the computer-readable storage medium to implement the protein sequence prediction method provided in the above method embodiments.

[0271] In an exemplary embodiment, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the protein sequence prediction method provided in each of the above method embodiments.

[0272] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0273] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes a computer storage medium and a communication medium, where the communication medium includes any medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0274] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.< / concat> < / bind> < / protein> < / bind> < / protein>

Claims

1. A method for predicting a protein sequence, characterized in that, The method includes: Obtaining an attribute tag of a predicted protein, where the attribute tag indicates an inherent property of the predicted protein based on natural language; Obtaining sequence information of the predicted protein, where the sequence information is used to constrain the amino acid sequence in the predicted protein; Invoking a protein sequence prediction model to perform sequence prediction on the attribute tag and the sequence information to obtain a protein sequence, where the protein sequence is the amino acid sequence of the predicted protein, the protein sequence includes the sequence information, and the predicted protein has the inherent property indicated by the attribute tag.

2. The method according to claim 1, wherein The obtaining of the attribute tag of the predicted protein includes at least one of the following: Obtaining a species tag, where the species tag is used to indicate the biological species to which the predicted protein belongs; Obtaining a function tag, where the function tag is used to indicate the protein function of the predicted protein in the biological organism to which it belongs; Obtaining a structure tag, where the structure tag is used to indicate the geometric structure of the predicted protein in three-dimensional space; Obtaining a task tag, where the task tag is used to indicate to the protein sequence prediction model the prediction manner of the protein sequence.

3. The method according to claim 2, characterized in that, The obtaining of the species tag includes: Determining the category name of at least one level of category in the biological classification of the biological species to which the predicted protein belongs as the species tag; Wherein, the biological classification includes at least one level of category among domain, kingdom, phylum, class, order, family, genus, and species.

4. The method according to claim 2, wherein The obtaining of the function tag includes: Determining the function tag according to the protein function of the predicted protein in at least one description dimension among cellular component, biological process, and molecular function.

5. The method according to claim 2, characterized in that, The obtaining of the structure tag includes: Determining the category name of at least one level of category in the structure classification of the predicted protein as the structure tag; Wherein, the structure classification includes at least one level of category among class, fold, superfamily, and family.

6. The method according to claim 2, characterized in that, The obtaining of the task tag includes: Concatenating a reference sequence and an association tag to obtain the task tag; Wherein, the reference sequence is the amino acid sequence of a known protein, and the association tag is used to indicate whether the predicted protein and the known protein are affinity.

7. The method according to any one of claims 1 to 6, characterized in that, The protein sequence prediction model includes an encoding network and a prediction network; The invoking of the protein sequence prediction model to perform sequence prediction on the attribute tag and the sequence information to obtain a protein sequence includes: Invoking the encoding network to perform feature encoding on the attribute tag and the sequence information respectively to obtain a tag identifier and a sequence identifier, where the tag identifier is the unique identifier of the attribute tag, and the sequence identifier is the unique identifier of the amino acid sequence; Invoking the prediction network to perform sequence prediction on the tag identifier and the sequence identifier to obtain a protein sequence.

8. The method according to claim 7, characterized in that, The encoding network includes a first sub-network and a second sub-network arranged in parallel; The invoking of the encoding network to perform feature encoding on the attribute tag and the sequence information respectively to obtain a tag identifier and a sequence identifier includes: Invoking the first sub-network to perform feature encoding on the attribute tag to obtain the tag identifier; Call the second sub-network to perform feature encoding on the sequence information to obtain the sequence identifier; Among them, the network parameters of the first sub-network and the second sub-network are different.

9. The method according to claim 8, wherein The first sub-network is a dictionary encoder; the step of calling the first sub-network to perform feature encoding on the attribute label to obtain the label identifier includes: Call the dictionary encoder to find at least two word identifiers corresponding to at least two phrases in the attribute label, where the at least two phrases in the attribute label are phrases divided according to semantics; Based on the positions of the at least two phrases in the attribute label, splice the at least two word identifiers to obtain the label identifier.

10. The method according to claim 8, wherein The second sub-network is a tokenizer; the step of calling the second sub-network to perform feature encoding on the sequence information to obtain the sequence identifier includes: Call the tokenizer to perform feature encoding on the sequence information to obtain the sequence identifier, where the sequence identifier includes at least two sub-identifiers corresponding to at least two sub-sequences, and the at least two sub-sequences are obtained by splitting the sequence information based on the tokenizer.

11. The method according to claim 7, characterized in that The method further includes: Obtain the sample sequence identifier and sample label identifier of the sample protein, and sample the training identifier from the sample sequence identifier and sample label identifier; Call the initial prediction network to perform sequence prediction on the training identifier to obtain the first protein sequence; Based on the difference between the first protein sequence and the sample sequence identifier, train the initial prediction network to obtain the prediction network.

12. The method according to claim 11, wherein The step of obtaining the sample sequence identifier and sample label identifier of the sample protein includes: Obtain the sample sequence and sample identifier of the sample protein; Call the encoding network to perform feature encoding on the sample sequence and the sample identifier respectively to obtain the sample sequence identifier and sample label identifier; The step of training the initial prediction network based on the difference between the first protein sequence and the sample sequence identifier to obtain the prediction network includes: Calculate the loss function between the first protein sequence and the sample sequence, and train the initial prediction network based on the loss function to obtain the prediction network; Among them, the protein sequence prediction model is obtained by splicing the encoding network and the trained prediction network.

13. The method according to claim 11, wherein The method further includes: Call the prediction network to perform label prediction on the sample sequence identifier to obtain the predicted label; Based on the difference between the predicted label and the sample label identifier, supplement the training of the prediction network; Among them, the protein sequence prediction model is obtained by splicing the encoding network and the prediction network after supplementary training.

14. A prediction device for a protein sequence, characterized in that, The device includes: An acquisition module, configured to acquire an attribute label of a predicted protein, where the attribute label indicates an inherent property of the predicted protein based on natural language; The acquisition module is further configured to acquire the sequence information of the predicted protein, where the sequence information is used to constrain the amino acid sequence in the predicted protein; A processing module, configured to call a protein sequence prediction model to perform sequence prediction on the attribute tag and the sequence information, so as to obtain a protein sequence, where the protein sequence is the amino acid sequence of the predicted protein, the protein sequence includes the sequence information, and the predicted protein has the inherent property indicated by the attribute tag.

15. A computer device, characterized in that, The computer device includes: a processor and a memory, where at least one segment of program is stored in the memory; the processor is configured to execute the at least one segment of program in the memory to implement the protein sequence prediction method according to any one of claims 1 to 13 above.

16. A computer-readable storage medium, characterized in that, Executable instructions are stored in the readable storage medium, and the executable instructions are loaded and executed by a processor to implement the protein sequence prediction method according to any one of claims 1 to 13 above.

17. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the protein sequence prediction method according to any one of claims 1 to 13 above.