Protein structure prediction method and device

CN119948566APending Publication Date: 2025-05-06HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280098500.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing protein structure prediction methods are time-consuming and cannot effectively predict protein structures of orphan sequences in nature.

Method used

By obtaining the amino acid sequence, generating multiple sequence alignment data based on its contextual relationship, and using machine learning models for prediction, the 3D structure of the protein can be directly generated, and even orphan sequences can be accurately predicted.

Benefits of technology

It significantly shortens the time for protein structure prediction, makes it possible to predict orphan sequences, and improves the accuracy and efficiency of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948566A_ABST
    Figure CN119948566A_ABST
Patent Text Reader

Abstract

The invention discloses a protein structure prediction method and device. The method comprises the following steps: acquiring a first amino acid sequence; generating first multi-sequence alignment MSA data based on a context relationship of a plurality of amino acids in the first amino acid sequence in the first amino acid sequence; n times of prediction are carried out based on the first MSA data, N 3D structures of the protein are obtained, and N is a positive integer. Through the method and the device, the duration of a protein structure prediction process can be effectively shortened, the accuracy of the MSA data for protein structure prediction is improved, and the protein structure of an orphan sequence in nature can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Protein structure prediction method and device Technical Field

[0001] The present application relates to the field of artificial intelligence (AI) technology in big data, and in particular to a protein structure prediction method and device. Background Art

[0002] Proteins play an important role in the central dogma of molecular biology and are indispensable in various life processes. Protein folding is the process by which proteins acquire their functional structure and conformation. After obtaining the three-dimensional structure of a protein, its functional analysis can be performed based on the structure and subsequent drug development can be carried out based on the relevant analysis results. In other words, the structure of a protein determines its biological function. Therefore, determining the structure of a protein can help understand life processes (for example, including the mechanisms of many diseases) and design proteins (for example, as drugs or as enzymes for industrial processes). For example, which molecules (such as drugs) will bind to a protein (and where the binding occurs) depends on the structure of the protein. Since the effectiveness of drugs is affected by the extent to which they bind to proteins (for example, in the blood), determining the structure of different proteins is an important aspect of drug development. However, the process of determining protein structure using physical experiments (for example, by X-ray crystallography) is time-consuming and very expensive.

[0003] Existing methods for predicting protein structure mainly rely on searching amino acid sequences homologous to the protein's amino acid sequence in an amino acid sequence database, i.e., multiple sequence alignment (MSA) data, and then using the homologous amino acid sequences to predict the protein's structure.

[0004] However, the above method of searching for homologous sequences from an amino acid sequence database is time-consuming, and for orphan sequences in nature, homologous sequences cannot be searched from an amino acid sequence database, and thus protein structure cannot be predicted.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a protein structure prediction method and apparatus, which can shorten the protein structure prediction process and accurately predict the protein structure of orphan sequences in nature.

[0007] In a first aspect, the present application provides a protein structure prediction method, the method comprising: obtaining a first amino acid sequence; generating first multiple sequence alignment (MSA) data based on the contextual relationship of multiple amino acids in the first amino acid sequence; performing N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

[0008] The first MSA data includes at least one amino acid sequence, wherein each amino acid sequence in the at least one amino acid sequence includes the same number of amino acids as the first amino acid sequence, and the amino acids at the partially corresponding positions are identical.

[0009] The above-mentioned contextual relationship includes the positional relationship of multiple amino acids in the first amino acid sequence relative to other amino acids in the first amino acid sequence.

[0010] From a technical perspective, the present application can directly generate first MSA data based on the contextual relationship between multiple amino acids in a first amino acid sequence. Compared to the prior art method of searching from an amino acid sequence database, this method can effectively shorten the time required to obtain the first MSA data, thereby shortening the time required to predict protein structure. In addition, for orphan sequences in nature, the present application method can also be used to quickly generate MSA data directly based on the orphan sequence, and then perform protein structure prediction based on the MSA data. In other words, the present application makes protein structure prediction for orphan sequences feasible.

[0011] In a feasible embodiment, the generating of the first multiple sequence alignment MSA data based on the contextual relationship of multiple amino acids in the first amino acid sequence includes: extracting the contextual relationship of multiple amino acids in the first amino acid sequence based on a machine learning model, and generating the first MSA data based on the contextual relationship.

[0012] From a technical perspective, the present application can rapidly extract the contextual relationships between multiple amino acids in a first amino acid sequence using a pre-trained machine learning model, and then generate first MSA data based on this contextual relationship. It is readily understood that the process of generating the first MSA data from the model is much faster than searching a database in the prior art, and can also be applied to orphan sequences to generate first MSA data, something that is not possible with the prior art.

[0013] In a feasible embodiment, first multiple sequence alignment MSA data is generated based on the context relationship of multiple amino acids in the first amino acid sequence, including: based on the first amino acid sequence, searching from an amino acid sequence database to obtain second MSA data, each amino acid sequence in the second MSA data contains partially identical amino acids with the first amino acid sequence; based on a machine learning model, extracting the context relationship of multiple amino acids in the first amino acid sequence, and using the context relationship and the second MSA data to generate third MSA data; using the third MSA data or the fourth MSA data as the first MSA data; wherein, the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

[0014] From the technical effect point of view, the present application can search for the second MSA data of the first amino acid sequence from the amino acid database, and then input the second MSA data and the first amino acid sequence into the machine learning model together, and use the searched second MSA data to guide the machine learning model to generate the third MSA data based on the first amino acid sequence. In this way, the machine learning model will use the searched homologous sequence (i.e., the second MSA data) to make the generated third MSA data more accurate; at the same time, it can also effectively predict orphan sequences. Finally, the second MSA data and the third MSA data are spliced ​​to obtain more comprehensive MSA data, which can then enable the subsequent use of the first MSA data to predict accurate protein structures.

[0015] In a feasible embodiment, the training data of the machine learning model includes a second amino acid sequence and a label, and the label is fourth MSA data obtained by searching an amino acid sequence database based on the second amino acid sequence.

[0016] From a technical perspective, the machine learning model's extraction of the contextual relationships between multiple amino acids in an amino acid sequence and the generation of corresponding MSA data are trained based on data searched in a database, and the MSA data generated by the model has high accuracy.

[0017] In a feasible embodiment, the method of performing N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer, includes: performing N predictions on the first MSA data to obtain N 3D structures of the protein; or, performing N corrections on the first MSA data to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; and performing predictions on the N corrected first MSA data to obtain N 3D structures of the protein.

[0018] From a technical perspective, this application can generate multiple predicted protein structures through separate predictions, providing greater selectivity in subsequent protein-based biological function analysis and drug development. Furthermore, through multiple corrections, multiple corrected first MSA data are obtained, and predictions using these corrected first MSA data are made, resulting in more accurate predicted protein 3D structures.

[0019] In a feasible embodiment, the first MSA data includes M amino acid sequences, the M amino acid sequences include a third amino acid sequence, M is a positive integer, and the first MSA data is corrected N times to obtain N corrected first MSA data, including: filling in the missing amino acids in the third amino acid sequence; and / or replacing the amino acids at P positions in the third amino acid sequence, P is a positive integer; wherein the P positions include the first position, the first position of the third amino acid sequence is the first amino acid, and each of the Q amino acid sequences contained in the M amino acid sequences is the second amino acid at the first position, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0020] From a technical perspective, this application fills in or replaces some of the amino acid sequence sites in the first MSA data, thereby making the first MSA data subsequently used for protein 3D structure prediction more accurate, thereby enabling the prediction process to obtain more accurate results.

[0021] In a feasible embodiment, the N 3D structures correspond to N prediction quality scores respectively, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

[0022] In a feasible embodiment, the method further includes: determining the K 3D structures from the N 3D structures based on the predicted quality scores, where K is a positive integer less than or equal to N; wherein the K 3D structures are the K 3D structures with the highest predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the true 3D structure of the protein.

[0023] From a technical perspective, the prediction quality score is used to characterize the difference between the predicted protein 3D structure and the true protein 3D structure, so that multiple 3D structures (i.e., K 3D structures) with relatively accurate prediction results can be screened out based on the prediction quality score.

[0024] In a feasible embodiment, the N corrected first MSA data are predicted separately to obtain N 3D structures of the protein, including: in the prediction process of each corrected first MSA data, predicting the 3D structure of the protein based on the corrected first MSA data; or, searching for a template protein structure in a protein database, the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; in the prediction process of each corrected first MSA data, predicting the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0025] From a technical perspective, after quickly generating the first MSA data based on the first amino acid sequence, the present application can predict the structure of the protein by adding a template protein or not using a template protein, which is faster and can solve the problem that the existing technology cannot predict the protein structure of orphan sequences.

[0026] In a second aspect, the present application provides a protein structure prediction method, the method comprising: obtaining a first amino acid sequence; searching for first MSA data from an amino acid sequence database based on the first amino acid sequence, the first MSA data including a third amino acid sequence; filling in missing amino acids in the third amino acid sequence, and / or replacing amino acids at P positions in the third amino acid sequence to obtain corrected first MSA data, where P is a positive integer; predicting the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence.

[0027] It should be understood that the third amino acid sequence is any amino acid sequence in the first MSA data.

[0028] From a technical perspective, after searching for the first MSA data for protein 3D structure prediction, the present application improves the accuracy of the corrected first MSA data by correcting the amino acid sequence contained in the first MSA data, that is, filling in the missing amino acid positions or replacing the incorrect positions, thereby enabling accurate prediction results to be obtained when the corrected first MSA data is subsequently used for protein 3D structure prediction.

[0029] In a feasible embodiment, the P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0030] In a feasible embodiment, predicting the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence includes: predicting the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or searching for a template protein structure in a protein database, wherein the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; and predicting the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0031] Specifically, the technical effects of the above embodiments can refer to the description of the corresponding embodiments in the above-mentioned first aspect, and will not be repeated here.

[0032] In a third aspect, an embodiment of the present application provides a protein structure prediction device, comprising: an acquisition unit for acquiring a first amino acid sequence; a processing unit for generating first multiple sequence alignment MSA data based on the contextual relationship of multiple amino acids in the first amino acid sequence; and a prediction unit for performing N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

[0033] In a feasible embodiment, in the aspect of generating the first multiple sequence alignment MSA data based on the contextual relationship of multiple amino acids in the first amino acid sequence, the processing unit is specifically used to: extract the contextual relationship of multiple amino acids in the first amino acid sequence based on a machine learning model, and generate the first MSA data based on the contextual relationship.

[0034] In a feasible embodiment, in the aspect of generating first multiple sequence alignment MSA data based on the context relationship of multiple amino acids in the first amino acid sequence, the processing unit is specifically used to: search and obtain second MSA data from an amino acid sequence database based on the first amino acid sequence, each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence; extract the context relationship of multiple amino acids in the first amino acid sequence based on a machine learning model, and generate third MSA data using the context relationship and the second MSA data; use the third MSA data or the fourth MSA data as the first MSA data; wherein the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

[0035] In a feasible embodiment, the prediction unit is specifically used to: perform N predictions on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence; or perform N corrections on the first MSA data to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; and perform predictions on the N corrected first MSA data to obtain N 3D structures of the protein.

[0036] In a feasible embodiment, the first MSA data includes M amino acid sequences, and the M amino acid sequences include a third amino acid sequence, where M is a positive integer; in the aspect of performing N corrections on the first MSA data to obtain N corrected first MSA data, the prediction unit is specifically used to: fill in the missing amino acids in the third amino acid sequence; and / or replace the amino acids at P positions in the third amino acid sequence, where P is a positive integer; wherein the P positions include a first position, the first position of the third amino acid sequence is a first amino acid, and each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first position, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0037] In a feasible embodiment, the N 3D structures correspond to N prediction quality scores respectively, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

[0038] In a feasible embodiment, the device also includes: a determination unit, configured to determine the K 3D structures from the N 3D structures based on the predicted quality scores, where K is a positive integer less than or equal to N; wherein the K 3D structures are the K 3D structures ranked higher in the predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the actual 3D structure of the protein.

[0039] In a feasible embodiment, in the aspect of predicting the N corrected first MSA data respectively to obtain N 3D structures of the protein, the processing unit is specifically used to: in the prediction process of each corrected first MSA data, predict the 3D structure of the protein based on the corrected first MSA data; or search for a template protein structure in a protein database, the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; in the prediction process of each corrected first MSA data, predict the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0040] In a fourth aspect, an embodiment of the present application provides a protein structure prediction device, comprising: an acquisition unit for acquiring a first amino acid sequence; a processing unit for searching for first MSA data from an amino acid sequence database based on the first amino acid sequence, wherein the first MSA data includes a third amino acid sequence; a correction unit for filling in missing amino acids in the third amino acid sequence, and / or replacing amino acids at P positions in the third amino acid sequence to obtain corrected first MSA data, where P is a positive integer; and a prediction unit for predicting the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence.

[0041] In a feasible embodiment, the P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0042] In a feasible embodiment, the prediction unit is specifically used to: predict the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or search for a template protein structure in a protein database, wherein the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; and predict the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0043] In the fifth aspect, an embodiment of the present application provides a chip system, which includes at least one processor, a memory and an interface circuit, wherein the memory, the interface circuit and the at least one processor are interconnected through lines, and instructions are stored in the at least one memory; when the instructions are executed by the processor, the method described in any one of the first aspect and / or the second aspect above is implemented.

[0044] In a sixth aspect, an embodiment of the present application provides a computer device, comprising at least one processor, a memory and an interface circuit, wherein the memory, the interface circuit and the at least one processor are interconnected via lines, and instructions are stored in the at least one memory; when the instructions are executed by the processor, the method described in any one of the first and / or second aspects above is implemented.

[0045] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the method described in any one of the first aspect and / or second aspect above is implemented.

[0046] In an eighth aspect, an embodiment of the present application provides a computer program, wherein the computer program product includes instructions. When the instructions are executed, the method described in any one of the first aspect and / or second aspect above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The following is an introduction to the drawings used in the embodiments of this application.

[0048] FIG1 is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0049] FIG2 is a schematic diagram of another system architecture provided in an embodiment of the present application;

[0050] FIG3 is a schematic diagram of a process for predicting protein structure according to an embodiment of the present application;

[0051] FIG4 is a schematic diagram of the structure of a machine learning model provided in an embodiment of the present application;

[0052] FIG5 is a schematic diagram of a chip hardware structure provided in an embodiment of the present application;

[0053] FIG6 is another protein structure prediction method provided in an embodiment of the present application;

[0054] FIG7 is a schematic diagram of a protein structure prediction process provided in an embodiment of the present application;

[0055] FIG8 is a schematic diagram of another protein structure prediction process provided in an embodiment of the present application;

[0056] FIG9 is a schematic diagram of another protein structure prediction process provided in an embodiment of the present application;

[0057] FIG10 is a schematic diagram of the structure of a protein structure prediction device provided in an embodiment of the present application;

[0058] FIG11 is a schematic diagram of the structure of another protein structure prediction device provided in an embodiment of the present application;

[0059] FIG12 is a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] The following describes the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents or, for example, A / B can represent A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" refers to two or more than two.

[0061] The terms "first", "second", "third" and "fourth" in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices. Reference to "embodiments" herein means that the specific features, structures or characteristics described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0062] The following first explains the relevant terms in this application.

[0063] (1) Amino acid sequence: This is the primary structure of a protein. It is the order in which amino acids are linked to form a peptide chain (or polypeptide), and is also the most basic structure of a protein. A peptide bond is formed between the amino group of one amino acid and the carboxyl group of another amino acid, and multiple amino acids are linked in sequence to form a peptide chain. The amino acid sequence is determined by the order of the genetic code in the gene. Various amino acids are linked together by peptide bonds in the order of the genetic code to form a polypeptide chain, so the peptide bond is the main bond in the protein structure.

[0064] (2) Protein structure: Also known as protein three-dimensional (3D) structure or protein tertiary structure. It refers to the relative spatial positions of all amino acids in the entire peptide chain, that is, the three-dimensional spatial structure of the entire peptide chain. The formation and stability of the tertiary structure mainly rely on chemical bonds such as hydrophobic bonds, salt bonds, disulfide bonds, and hydrogen bonds.

[0065] (3) Multiple Sequence Alignment (MSA): It is used to describe whether the same characters are present at the same points / positions in the multiple sequences involved in the alignment. In protein structure prediction, it is used to describe whether the amino acid types are the same at the corresponding same points / positions in multiple amino acid sequences (usually more than three). Specifically: without changing the order of amino acids within each amino acid sequence, the same amino acids in different amino acid sequences are arranged in the same column as much as possible, and it is believed that the sequences in the same column are homologous in evolution and have a common ancestor. The MSA data in this application refers to multiple amino acid sequences sorted according to the above-mentioned multiple sequence alignment rules.

[0066] Please refer to Figure 1, which is a schematic diagram of a system architecture provided in an embodiment of the present application, and is used to describe the system architecture of a computer device 100 for executing the protein structure prediction method of the present application. As shown in Figure 1, the system architecture of the computer device may include an application layer 110, an operating system 120, and a device layer 130.

[0067] Optionally, the computer device 100 may be a personal computer or a server, etc., which is not limited in this application.

[0068] Optionally, the application layer 110 may include a deep learning framework 112 and a machine learning model / structure prediction model / scoring model 111 generated based on the deep learning framework. The machine learning model / structure prediction model / scoring model 111 generated by the deep learning framework is used to execute the protein structure prediction method of the present application.

[0069] Optionally, the operating system 120 may include a file system 121, a block layer 122, and a device driver 123. The file system 121 manages and schedules file storage space, providing the logical and physical structures and storage methods for files. The block layer 122 is the interface through which the file system 121 accesses the device layer 130, connecting the file system 121 and the device driver 123. The block layer 122 is used to encapsulate and decapsulate related requests. The device driver 123 may include a display driver, a camera driver, an audio driver, a sensor driver, and the like.

[0070] Optionally, the device layer 130 may include a memory 131 , a network card 132 , and a processor 133 .

[0071] In some feasible implementations, program execution can be divided into user mode and kernel mode. When a program runs in user mode, the processor can only access data in a portion of the memory and is not allowed to access peripheral devices such as the hard disk and network interface card. When a program runs in kernel mode, the processor can access all data in the memory, including peripheral devices such as the hard disk and network interface card. At the same time, the processor can also switch itself from one program to another. Typically, application programs in the application layer 110 run in user mode, and the operating system 120 runs in kernel mode.

[0072] During the process of the computer device 100 running the above-mentioned machine learning model / structural prediction model / scoring model 111, the application layer 110 first initiates a read / write request; the file system 121 determines the logical address of the data corresponding to the read / write request in the memory 131; the block layer 122 is used to distribute the read / write request to the device driver 123; the device driver 123 is used to encapsulate the read / write request, and send the encapsulated read / write request and the logical address corresponding to the read / write request to the memory 131; the memory 131 unpacks the encapsulated read / write request, and converts the logical address corresponding to the read / write request into the corresponding physical address in the memory 131; and then writes data to or reads data from the physical address based on the read / write request.

[0073] Optionally, the memory 100 can be any one of random access memory (RAM), read-only memory (ROM) or flash memory; wherein RAM includes static random access memory (SRAM) and dynamic random access memory (DRAM), and ROM includes erasable programmable ROM (EPROM) and electrically erasable programmable read-only memory (EEPROM).

[0074] Please refer to Figure 2, which is a schematic diagram of another system architecture provided in an embodiment of the present application.

[0075] As shown in FIG2 , the data acquisition device 260 is used to collect amino acid sequence and protein structure data and store them in the database 230. The training device 220 generates a machine learning model / structure prediction model / scoring model 111 based on the amino acid sequence and protein structure data maintained in the database 230. The method embodiment shown in FIG3 will now describe in detail how the training device 220 obtains the machine learning model / structure prediction model / scoring model 111 based on the training data. The machine learning model generates first MSA data based on the input amino acid sequence (i.e., the first amino acid sequence in the embodiment below). The structure prediction model predicts based on the first MSA data to obtain the 3D structure of the protein to be predicted. Finally, the scoring model scores the predicted protein 3D structure to obtain a prediction quality score. The first amino acid sequence is generated based on a user request sent by the client device 240, wherein the user request can be text information input by the user.

[0076] Figure 2 can also be a functional module diagram of the protein structure prediction process. Specifically, the execution device 210 and the data storage system 250 can be integrated into the user device 240 when the user device 240 has relatively strong data processing capabilities. In some embodiments, the execution device 210 can be a device independent of the user device 240. The data storage system 250, database 230, training device 220, and data acquisition device 260 can be integrated into the execution device 210 or located on other servers in the cloud or on the network, which is not limited in this application.

[0077] The data acquisition device 260 may be a terminal device or an input / output interface of a server or cloud, and may be an interactive layer (interface) for obtaining query statements and returning reply statements.

[0078] The following is a brief introduction to the training and reasoning principles of the machine learning model in this application.

[0079] The architecture of a machine learning model can be a deep neural network. The work of each layer in a deep neural network can be expressed mathematically as To describe: From a physical perspective, the work of each layer in a deep neural network can be understood as completing the transformation from input space to output space (i.e., from the row space to the column space of a matrix) through operations on five input spaces (a set of input vectors). These five operations include: 1. Dimensionality increase / decrease; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are represented by Completed, operation 4 is completed by +b, and operation 5 is implemented by a(). The word "space" is used here because the object being classified is not a single thing, but a class of things, and space refers to the collection of all individuals of this class of things. Among them, W is a weight vector, and each value in the vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the spatial transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training a deep neural network is to eventually obtain the weight matrix of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially to learn how to control spatial transformation, and more specifically to learn the weight matrix.

[0080] Because we want the output of a deep neural network to be as close as possible to the value we actually want to predict, we can compare the current network's predicted value with the desired target value, and then update the weight vector of each layer of the neural network based on the difference between the two (of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network). For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value, and this adjustment is continued until the neural network can predict the desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value." This is the loss function or objective function, which is an important equation used to measure the difference between the predicted value and the target value. Taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0081] In Figure 2 , the machine learning model / structure prediction model / scoring model 111 generated by the training device 220 can be applied to various systems or devices. The execution device 210 is configured with an I / O interface 212 for data exchange with external devices. A "user" can input data, i.e., a user request, into the I / O interface 212 via a client device 240, including the amino acid sequence to be predicted, i.e., text information.

[0082] The execution device 210 can call data, code, etc. in the data storage system 250 , and can also store data, instructions, etc. in the data storage system 250 .

[0083] The computing module 211 uses the machine learning model / structure prediction model / scoring model 111 to identify and process the input data (ie, user request) to generate a predicted protein 3D structure.

[0084] Finally, the I / O interface 212 returns the predicted protein 3D structure to the client device 240 and presents it to the user on the client device 240 .

[0085] In the scenario shown in Figure 2 , the user can manually specify data to be input into execution device 210, for example, by operating within the interface provided by I / O interface 212. Alternatively, client device 240 can automatically input data into I / O interface 212 and obtain results. If automatic data input by client device 240 requires user authorization, the user can set the corresponding permissions within client device 240. The user can view the results output by execution device 210 on client device 240, which can be presented in a display, sound, action, or other specific form. Client device 240 can also serve as a data acquisition terminal, storing the collected video and text data in database 230 for use in the training process.

[0086] It is worth noting that Figure 2 is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationship between the devices, components, modules, etc. shown in Figure 2 does not constitute any limitation. For example, in Figure 2, the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210.

[0087] When the execution device 210 and the user equipment 240 in FIG. 2 are independent devices, their system architecture may be the same as the system architecture shown in FIG. 1 .

[0088] Please refer to Figure 3, which is a flow chart of a protein structure prediction method provided in an embodiment of the present application. As shown in Figure 3, the method includes steps S310, S320 and S330.

[0089] Step S310: Obtain a first amino acid sequence.

[0090] The first amino acid sequence is the amino acid sequence of the protein for which three-dimensional structure prediction is required, that is, the primary structure of the protein.

[0091] Optionally, the above-mentioned first amino acid sequence can be obtained through experimental testing or other means, which is not limited in this application.

[0092] The first amino acid sequence is a sequence composed of multiple amino acids linked in sequence based on peptide bonds.

[0093] Step S320: generating first multiple sequence alignment (MSA) data based on the contextual relationship of multiple amino acids in the first amino acid sequence.

[0094] The first MSA data includes at least one amino acid sequence, wherein each amino acid sequence in the at least one amino acid sequence includes the same number of amino acids as the first amino acid sequence, and the types of amino acids at the partially corresponding identical sites / positions are the same.

[0095] Each amino acid sequence in the above-mentioned first MSA data can also be called a homologous sequence of the first amino acid sequence. When the first amino acid sequence and an amino acid sequence in the first MSA data contain more points / positions of the same amino acid, the amino acid sequence has a higher homology with the first amino acid sequence.

[0096] Optionally, the MSA data in this application may be represented by a matrix or other data structures, which is not limited in this application.

[0097] The above contextual relationship includes the positions of multiple amino acids in the first amino acid sequence relative to other amino acids in the first amino acid sequence.

[0098] Optionally, the above contextual relationship may be represented by a matrix, a vector or other forms, which is not limited in this application.

[0099] Specifically, the manners for generating the first MSA data in the above step S320 may include the following three.

[0100] (1) The first method

[0101] The above-mentioned method of generating the first multiple sequence alignment MSA data based on the context relationship of multiple amino acids in the first amino acid sequence includes: extracting the context relationship of multiple amino acids in the first amino acid sequence based on a first machine learning model, and generating the first MSA data based on the context relationship.

[0102] Specifically, the first amino acid sequence is input into a trained first machine learning model. The first machine learning model first extracts the contextual relationship of multiple amino acids in the first amino acid sequence (represented by a matrix or vector). Then, the first machine learning model generates first MSA data based on the contextual relationship of multiple amino acids in the first amino acid sequence.

[0103] (2) The second and third methods

[0104] The above-mentioned method of generating the first multiple sequence alignment MSA data based on the context relationship of multiple amino acids in the first amino acid sequence includes: based on the first amino acid sequence, searching for second MSA data from an amino acid sequence database, and each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence; based on a second machine learning model, extracting the context relationship of multiple amino acids in the first amino acid sequence in the first amino acid sequence, and processing the context relationship and the second MSA data through a third machine learning model to generate third MSA data; using the third MSA data or the fourth MSA data as the first MSA data; wherein, the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

[0105] The process of obtaining the first MSA data by the second method is: using the third MSA data generated based on the first amino acid sequence and the second MSA data as the first MSA data.

[0106] The process of obtaining the first MSA data in the third manner is: splicing the second MSA data and the third MSA data to generate the first MSA data.

[0107] Specifically, the process of searching for the second MSA data is as follows: based on the first amino acid sequence, searching an amino acid sequence database for at least one amino acid sequence homologous to the first amino acid sequence to obtain the second MSA data.

[0108] The homologous sequence indicates that each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence.

[0109] Optionally, the above search process may be obtained by searching using retrieval software such as MMSEQS2, Jackhmmer, or other feasible methods, which is not limited in this application.

[0110] The above-mentioned generation of the third MSA data using the contextual relationship and the second MSA data may specifically include: extracting the contextual relationship in the first amino acid sequence using the second machine learning model. Then, the contextual relationship between the second MSA data and the first amino acid sequence is input into the third machine learning model: first, model parameters in the third machine learning model are updated using the second MSA data; and then, the extracted contextual relationship is processed using the third machine learning model with the updated model parameters to generate the third MSA data.

[0111] From a technical perspective, the second MSA data searched from the database is used as a result to guide the third machine learning model to generate the third MSA data. This process takes into account both historical experience (i.e., the model parameters obtained through training) and actual results (using the second MSA data to update the model parameters), making the generated third MSA data more accurate.

[0112] Specifically, after the third MSA data is generated, the third MSA data and the second MSA data may be spliced ​​together to generate the first MSA data. The process may be:

[0113] For example, when the third MSA data is a 3*3 matrix and the second MSA data is a 2*3 matrix, the first MSA data obtained after splicing is a 5*3 matrix.

[0114] For another example, when the third MSA data is a four-element one-dimensional vector and the second MSA data is a six-element one-dimensional vector, the first MSA data obtained after splicing is a ten-element one-dimensional vector.

[0115] Optionally, the specific training process of the above machine learning model (which may be the first, second or third machine learning model) is as follows:

[0116] The training data for the machine learning model includes a second amino acid sequence and a label, wherein the label is fourth MSA data obtained by searching an amino acid sequence database based on the second amino acid sequence. It should be understood that the second amino acid sequence and the fourth MSA data obtained by searching the second amino acid sequence are only one set of training data for the machine learning model.

[0117] During the specific training process, the second amino acid sequence is input into the machine learning model, which performs feature extraction on the second amino acid sequence to obtain a contextual relationship between multiple amino acids in the second amino acid sequence. Based on the contextual relationship, fifth MSA data is generated. A loss function value is then calculated based on the difference between the fifth MSA data and the fourth MSA data, and model parameters of the machine learning model are adjusted based on the loss function value.

[0118] Repeat the above training process until the loss function value reaches the threshold value, that is, the model converges, and a trained machine learning model is obtained.

[0119] The tag of the second amino acid sequence may be a homologous sequence obtained by searching from an amino acid sequence database based on the search method in the aforementioned embodiment.

[0120] Optionally, the above-mentioned machine learning model can be a traditional machine learning model (such as decision tree, random forest, artificial neural network, Bayesian learning, etc.) or a deep learning model, which is not limited in this application.

[0121] Please refer to Figure 4, which is a structural diagram of a machine learning model provided in an embodiment of the present application, which is used to describe the specific structures of the aforementioned first machine learning model, second machine learning model, and third machine learning model. As shown in Figure 4, the machine learning model includes an encoding module and a decoding module. Among them,

[0122] The encoding module includes at least one conditional transformer (CT) and at least one encoding transformer (ET). The decoding module includes at least one latent transformer (LT) and a perturbation module, wherein the perturbation module is used to generate / update a perturbation factor.

[0123] After input data (amino acid sequence and / or MSA data) is input into the encoding module, the encoding module is used to extract features from the input data. For example, feature extraction can be performed on the input amino acid sequence to obtain contextual relationships between multiple amino acids in the amino acid sequence. Another example is feature extraction on the input MSA data, and model parameters can be updated based on the extracted features.

[0124] The decoding module is used to process the context relationship and generate a homologous sequence of the input amino acid sequence, namely, MSA data (for example, the first MSA data or the third MSA data in this application).

[0125] Optionally, the difference between the aforementioned first machine learning model, the second machine learning model and the third machine learning model may be a quantitative difference in at least one of the included CT, ET, LT and perturbation modules.

[0126] Step S330: performing N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

[0127] The above process includes: performing N predictions on the first MSA data to obtain N 3D structures of the protein; or, performing N corrections on the first MSA data to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; and performing predictions on the N corrected first MSA data to obtain N 3D structures of the protein.

[0128] Optionally, the first MSA data includes M amino acid sequences, the M amino acid sequences include a third amino acid sequence, and M is a positive integer. That is, the third amino acid sequence is any amino acid sequence in the first MSA data.

[0129] Specifically, the above-mentioned prediction using the first MSA data to obtain N 3D structures of the protein may include two ways: correcting the first MSA data before prediction and not correcting the first MSA data before prediction.

[0130] The following describes the specific correction process using the third amino acid sequence in the first MSA data as an example.

[0131] (1) Correction of the first MSA data before prediction

[0132] The first MSA data is corrected N times to obtain N corrected first MSA data, including: filling in the missing amino acids in the third amino acid sequence; and / or replacing the amino acids at P positions in the third amino acid sequence, where P is a positive integer; wherein the P positions include the first position, the first position of the third amino acid sequence is the first amino acid, and each of the Q amino acid sequences contained in the M amino acid sequences is the second amino acid at the first position, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0133] Specifically, there are three ways to modify the third amino acid sequence:

[0134] (1) Fill in the missing amino acids in the third amino acid sequence

[0135] Based on other amino acid sequences in the first MSA data, the missing amino acid positions on the third amino acid sequence are supplemented. Specifically: assuming that the amino acid at the second position on the third amino acid sequence is missing, at this time, among the M amino acid sequences contained in the first MSA data, if more than a preset number of amino acid sequences have the third amino acid at the second position, then the second position on the third amino acid sequence is supplemented with the third amino acid.

[0136] Among them, the above-mentioned preset number is a value pre-set based on a specific scenario, and this application does not limit this.

[0137] (2) Substitution of amino acids at some positions in the third amino acid sequence

[0138] The process of amino acid replacement is described with the first position among P positions as the object: for the first position, the amino acid at the first position of the third amino acid sequence is the first amino acid. When Q or more amino acid sequences among the M amino acid sequences do not have the first amino acid at their respective first positions, and all are the second amino acid, it means that there is an error in the amino acid at the first position of the third amino acid sequence, and it is corrected, that is, the first amino acid at the first position of the third amino acid sequence is replaced by the second amino acid.

[0139] Optionally, the condition for amino acid replacement at the first position is: first position - Q, that is, when there are Q amino acid sequences whose amino acids at the first position are different from the amino acid at the first position of the third amino acid sequence, replacement is performed.

[0140] Optionally, different replacement conditions can be set for the positions on the third amino acid sequence, which is not limited in this application.

[0141] The P sites mentioned above include co-evolution sites and / or conserved sites. The specific method for determining co-evolution sites and conserved sites is not elaborated in this application.

[0142] (3) Complement the missing amino acids in the third amino acid sequence and replace the amino acids at some positions

[0143] The third amino acid sequence is not only filled with missing amino acids, but also partially replaced. The specific process can be found in the detailed description of the first and second correction methods above, which will not be repeated here.

[0144] Specifically, the modification of the first MSA data may be implemented by inputting the first amino acid sequence and the first MSA data into a modification model to obtain the modified first MSA data.

[0145] The step of correcting the first MSA data N times to obtain N corrected first MSA data comprises: in each correction process, inputting the first amino acid sequence and the first MSA data into a correction model to correct the first MSA data to obtain corrected first MSA data.

[0146] Among them, the model structure of the above-mentioned correction model can be the same as or different from the model structure of the machine learning model in the aforementioned embodiment, and this application does not limit this.

[0147] Among them, there are two ways to predict the 3D structure of the protein based on the corrected first MSA data:

[0148] (1) No template protein is used

[0149] In this manner, during the prediction process for each corrected first MSA data, the 3D structure of the protein is predicted based on the corrected first MSA data.

[0150] Specifically, in each correction process, the first amino acid sequence and the corrected first MSA data are input into the structure prediction model to obtain the 3D coordinates of each atom (carbon, hydrogen, oxygen, etc.) in the protein, and the 3D structure of the protein is generated based on the 3D coordinates of each atom.

[0151] (2) Using template protein

[0152] In this way, a template protein structure is first searched in a protein database, wherein the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; in the prediction process of each corrected first MSA data, the 3D structure of the protein is predicted based on the template protein structure and the corrected first MSA data.

[0153] Among them, the protein database stores the three-dimensional structure of proteins and the amino acid sequence of each protein.

[0154] Specifically, the above-mentioned searching for the template protein structure in the protein database includes: searching for an amino acid sequence homologous to the first amino acid sequence from the protein database based on the first amino acid sequence, and then searching for the protein structure of the homologous amino acid sequence, that is, the template protein structure.

[0155] It is easy to understand that the amino acid sequence of the template protein structure contains the same number of amino acids as the first amino acid sequence, and the types of amino acids at the partially corresponding identical sites / positions are the same.

[0156] The above-mentioned generating the 3D structure based on the template protein structure and the corrected first MSA data includes: after searching for the template protein structure, inputting the template protein structure, the corrected first MSA data and the first amino acid sequence into a structure prediction model; for the first amino acid sequence and the homologous sequence fragments in the amino acid sequence of the template protein, the structure prediction model will refer to the 3D structure of the homologous fragment in the template protein structure to generate the protein 3D structure of the homologous fragment in the first amino acid sequence.

[0157] Optionally, the above-mentioned structure prediction model can be a structure prediction model composed of one or more of a convolutional neural network, a transformer, a recurrent neural network or other network structures, which is not limited in this application.

[0158] Furthermore, optionally, the training process of the structure prediction model is as follows:

[0159] The training data for the structure prediction model includes an amino acid sequence, MSA data obtained from searching an amino acid sequence database based on the amino acid sequence, and a label (the actual 3D structure of the protein corresponding to the amino acid sequence). During a single training run, the amino acid sequence and the MSA data obtained from searching an amino acid sequence database based on the amino acid sequence are fed into the structure prediction model, which predicts the 3D structure of the protein corresponding to the amino acid sequence. The predicted 3D structure is then compared and calculated with the actual 3D structure of the protein to obtain the value of the loss function, which is then used to adjust the model parameters.

[0160] The above training process is repeated until the loss function value reaches the threshold value, that is, the model converges, and a trained structure prediction model is obtained.

[0161] (2) No correction is made to the first MSA data before prediction

[0162] Specifically, the structure prediction model is directly used to perform N predictions on the first MSA data to obtain the 3D structures of N proteins.

[0163] The prediction process for each time may refer to the aforementioned process of predicting the corrected first MSA data, which will not be described in detail here.

[0164] Optionally, the N 3D structures correspond to N prediction quality scores respectively, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

[0165] Specifically, after N 3D structures of the protein are predicted, the prediction quality score of each of the N 3D structures is calculated, that is, a total of N prediction quality scores.

[0166] Further, optionally, the calculation process of each prediction quality score is as follows: a 3D structure of a protein is input into the scoring model, the scoring model extracts the 3D structural features of the protein, and predicts the prediction quality score of the input 3D structure based on the extracted features.

[0167] The prediction quality score is used to characterize the difference between the spatial position (spatial coordinates) of each atom in the predicted 3D structure of the protein and the spatial position of each atom in the true 3D structure of the protein.

[0168] Optionally, the above-mentioned scoring model can be a scoring model constructed by one or more of a convolutional neural network, a transformer, a recurrent neural network or other network structures, which is not limited in this application.

[0169] Furthermore, optionally, the training process of the scoring model is specifically as follows:

[0170] The scoring model is trained on a set of data consisting of predicted protein 3D structures and labels (scores calculated based on the predicted protein 3D structures and the protein's true 3D structures). During a single training run, the predicted protein 3D structure is fed into the scoring model, which generates a predicted score for the predicted protein 3D structure. The predicted score is then compared with the label (the true score) and calculated to determine the value of the loss function, based on which the model parameters are adjusted.

[0171] The above training process is repeated until the loss function value reaches the threshold value, that is, the model converges, and a trained scoring model is obtained.

[0172] From a technical perspective, the present application can directly generate first MSA data based on the contextual relationship between multiple amino acids in a first amino acid sequence. Compared to the prior art method of searching from an amino acid sequence database, this method can effectively shorten the time required to obtain the first MSA data, thereby shortening the time required to predict protein structure. In addition, for orphan sequences in nature, the present application method can also be used to quickly generate MSA data directly based on the orphan sequence, and then perform protein structure prediction based on the MSA data. In other words, the present application makes protein structure prediction for orphan sequences feasible.

[0173] After obtaining a predicted quality score for each of the N 3D structures, the K 3D structures are determined from the N 3D structures based on the predicted quality score, where K is a positive integer less than or equal to N; wherein the K 3D structures are the K 3D structures with the highest predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the true 3D structure of the protein.

[0174] Specifically, since the prediction quality score is used to describe the difference between the predicted 3D structure and the true 3D structure of the protein, the higher the prediction quality score, the smaller the difference between the predicted 3D structure and the true 3D structure; then, based on the prediction quality score, K 3D structures with the highest prediction quality scores are screened out from N 3D structures.

[0175] Please refer to Figure 6, which shows another protein structure prediction method provided by the embodiment of the present application. As shown in Figure 6, the method includes steps S610, S620, S630 and S640.

[0176] Step S610: Obtain a first amino acid sequence. Step S620: Search an amino acid sequence database for first MSA data based on the first amino acid sequence, wherein the first MSA data includes a third amino acid sequence. Step S630: Fill in missing amino acids in the third amino acid sequence and / or replace amino acids at P positions in the third amino acid sequence to obtain corrected first MSA data, where P is a positive integer. Step S640: Predict the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence.

[0177] Optionally, the P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0178] Optionally, predicting the corrected first MSA data to obtain the 3D structure of the protein includes: predicting the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or searching for a template protein structure in a protein database, wherein the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; and predicting the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0179] Specifically, the specific implementation of each step in the embodiment of FIG6 can refer to the corresponding description in the embodiment of FIG3 above, which will not be repeated here.

[0180] Please refer to Figure 5, which is a schematic diagram of the chip hardware structure provided by an embodiment of the present application. As shown in Figure 5, a neural network processing unit (NPU) 50 is mounted on the host CPU (Host CPU) as a coprocessor, and the Host CPU assigns tasks to execute the protein structure prediction method of the aforementioned embodiment or perform the training process of the related model in the aforementioned embodiment.

[0181] The core part of the NPU is the operation circuit 503, and the controller 504 controls the operation circuit 503 to extract data from the memory (weight memory or input memory) and perform operations.

[0182] In some implementations, the arithmetic circuit 503 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 503 is a two-dimensional systolic array. The arithmetic circuit 503 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 503 is a general-purpose matrix processor.

[0183] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 501 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 508.

[0184] The vector calculation unit 507 can further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 507 can be used for network calculations of non-convolutional / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.

[0185] In some implementations, the vector calculation unit 507 can store the processed output vector to the unified memory 506. For example, the vector calculation unit 507 can apply a nonlinear function to the output of the operation circuit 503, such as a vector of accumulated values, to generate an activation value. In some implementations, the vector calculation unit 507 generates a normalized value, a merged value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 503, for example, for use in a subsequent layer in a neural network.

[0186] The unified memory 506 is used to store input data and output data.

[0187] The memory unit access controller 505 (Direct Memory Access Controller, DMAC) moves the input data in the external memory to the input memory 501 and / or the unified memory 506, stores the weight data in the external memory into the weight memory 502, and stores the data in the unified memory 506 into the external memory.

[0188] The bus interface unit (BIU) 510 is used to implement interaction between the main CPU, DMAC and instruction fetch memory 509 through the bus.

[0189] An instruction fetch buffer 509 connected to the controller 504 is used to store instructions used by the controller 504 .

[0190] The controller 504 is used to call the instructions cached in the instruction fetch memory 509 to control the working process of the computing accelerator.

[0191] Generally, the unified memory 506, the input memory 501, the weight memory 502 and the instruction fetch memory 509 are all on-chip memories, and the external memory is a memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM) or other readable and writable memory.

[0192] Please refer to Figure 7, which is a schematic diagram of a protein structure prediction process provided in an embodiment of the present application.

[0193] In the method shown in Figure 7, step S710 is first executed: obtaining the amino acid sequence of the protein to be predicted. That is, the first amino acid sequence in the aforementioned embodiment, which can be represented by text. Step S720 is executed: inputting the amino acid sequence information into the machine learning model to obtain the first MSA data. Then step S730 is executed: correcting the obtained first MSA data to obtain the corrected MSA data, and step S740: predicting the 3D structure of the protein to be predicted based on the structure prediction model. Among them, step S730 and step S740 are executed N times, and N 3D structures are predicted. Step S750 is executed: scoring the predicted protein 3D structure based on the scoring model. Among them, step S760 is an optional step.

[0194] Specifically, the specific execution process of the above steps can be found in the description of the above embodiments, and will not be repeated here.

[0195] Please refer to Figure 8, which is a schematic diagram of another protein structure prediction process provided in an embodiment of the present application.

[0196] The method shown in Figure 8 differs from the method in Figure 7 in how to obtain the MSA data before correction (i.e., the first MSA data). In the method of Figure 8, after executing step S810, step S820 is executed: based on the amino acid sequence, the second MSA data is searched from the amino acid sequence database, and then step S830 is executed: the MSA data obtained by the search and the amino acid sequence are input into the machine learning model to obtain the third MSA data. That is, the second MSA data obtained by the search and the amino acid sequence are input into the machine learning model together to generate the third MSA data. Further, the second MSA data obtained by the search and the third MSA data generated by the machine learning model are spliced ​​to obtain the fourth MSA data. Finally, the fourth MSA data or the third MSA data is used as the first MSA data in the aforementioned embodiment.

[0197] The specific processes of the above steps and other steps in FIG8 can be referred to the description in the aforementioned embodiment and will not be repeated here.

[0198] It should be noted that step S870 in FIG. 8 is also an optional step.

[0199] Please refer to FIG9 , which is a schematic diagram of another protein structure prediction process provided in an embodiment of the present application.

[0200] The method shown in FIG9 differs from the method shown in FIG7 in how the pre-corrected MSA data (i.e., the first MSA data) is obtained. Specifically, in the method of FIG9 , the first MSA data is directly obtained by searching an amino acid sequence database based on the first amino acid sequence, i.e., step S920. The remaining steps are identical to those in FIG7 and are not further described here.

[0201] Please refer to Figure 10, which is a schematic diagram of the structure of a protein structure prediction device provided in an embodiment of the present application. As shown in Figure 10, the device includes an acquisition unit 1010, a processing unit 1020 and a prediction unit 1030.

[0202] An acquisition unit 1010 is configured to acquire a first amino acid sequence; a processing unit 1020 is configured to generate first multiple sequence alignment (MSA) data based on a contextual relationship between multiple amino acids in the first amino acid sequence; and a prediction unit 1030 is configured to perform N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

[0203] In a feasible embodiment, in the aspect of generating the first multiple sequence alignment MSA data based on the contextual relationship of multiple amino acids in the first amino acid sequence, the processing unit is specifically used to: extract the contextual relationship of multiple amino acids in the first amino acid sequence based on a machine learning model, and generate the first MSA data based on the contextual relationship.

[0204] In a feasible embodiment, in the aspect of generating first multiple sequence alignment MSA data based on the context relationship of multiple amino acids in the first amino acid sequence, the processing unit is specifically used to: search and obtain second MSA data from an amino acid sequence database based on the first amino acid sequence, each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence; extract the context relationship of multiple amino acids in the first amino acid sequence based on a machine learning model, and generate third MSA data using the context relationship and the second MSA data; use the third MSA data or the fourth MSA data as the first MSA data; wherein the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

[0205] In a feasible embodiment, the prediction unit is specifically used to: perform N predictions on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence; or perform N corrections on the first MSA data to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; and perform predictions on the N corrected first MSA data to obtain N 3D structures of the protein.

[0206] In a feasible embodiment, the first MSA data includes M amino acid sequences, and the M amino acid sequences include a third amino acid sequence, where M is a positive integer; in the aspect of performing N corrections on the first MSA data to obtain N corrected first MSA data, the prediction unit is specifically used to: fill in the missing amino acids in the third amino acid sequence; and / or replace the amino acids at P positions in the third amino acid sequence, where P is a positive integer; wherein the P positions include a first position, the first position of the third amino acid sequence is a first amino acid, and each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first position, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0207] In a feasible embodiment, the N 3D structures correspond to N prediction quality scores respectively, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

[0208] In a feasible embodiment, the device also includes: a determination unit, configured to determine the K 3D structures from the N 3D structures based on the predicted quality scores, where K is a positive integer less than or equal to N; wherein the K 3D structures are the K 3D structures ranked higher in the predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the actual 3D structure of the protein.

[0209] In a feasible embodiment, in the aspect of predicting the N corrected first MSA data respectively to obtain N 3D structures of the protein, the processing unit is specifically used to: in the prediction process of each corrected first MSA data, predict the 3D structure of the protein based on the corrected first MSA data; or search for a template protein structure in a protein database, the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; in the prediction process of each corrected first MSA data, predict the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0210] Specifically, the specific process of the protein prediction method executed by the protein prediction apparatus shown in FIG10 can be referred to the aforementioned method embodiment, which will not be described in detail here.

[0211] Please refer to Figure 11, which is a schematic diagram of the structure of another protein structure prediction device provided in an embodiment of the present application. As shown in Figure 11, the device includes an acquisition unit 1110, a processing unit 1120, a correction unit 1130 and a prediction unit 1140.

[0212] An acquisition unit 1110 is used to acquire a first amino acid sequence; a processing unit 1120 is used to search for first MSA data from an amino acid sequence database based on the first amino acid sequence, where the first MSA data includes a third amino acid sequence; a correction unit 1130 is used to fill in missing amino acids in the third amino acid sequence and / or replace amino acids at P positions in the third amino acid sequence to obtain corrected first MSA data, where P is a positive integer; and a prediction unit 1140 is used to predict the corrected first MSA data to obtain a 3D structure of a protein corresponding to the first amino acid sequence.

[0213] In a feasible embodiment, the P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

[0214] In a feasible embodiment, the prediction unit is specifically used to: predict the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or search for a template protein structure in a protein database, wherein the amino acid sequence in the template protein structure contains some identical amino acids with the first amino acid sequence; and predict the 3D structure of the protein based on the template protein structure and the corrected first MSA data.

[0215] Specifically, the specific process of the protein prediction method executed by the protein prediction apparatus shown in FIG11 can be referred to the aforementioned method embodiment, which will not be described in detail here.

[0216] Please refer to Figure 12, which is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be the computer device in the aforementioned method embodiment. As shown in Figure 12, the computer device includes a processor 1201, a memory 1202, an interface circuit 1203, and a bus 1204. The specific functions of each hardware structure of the computer device can be as follows:

[0217] Interface circuit 1203 is configured to obtain a first amino acid sequence. Processor 1201 is configured to generate first multiple sequence alignment (MSA) data based on the contextual relationships of multiple amino acids in the first amino acid sequence; and further configured to perform N predictions based on the first MSA data to obtain N 3D structures of the protein, where N is a positive integer. Memory 1202 is configured to store the first amino acid sequence and the N 3D structures of the protein.

[0218] It should be understood that the specific operation process of the processor and memory on the computer device in the embodiment of the present application can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0219] An embodiment of the present application provides a chip system, which includes at least one processor, a memory and an interface circuit. The memory, the interface circuit and the at least one processor are interconnected by lines, and instructions are stored in the at least one memory. When the instructions are executed by the processor, some or all of the steps of any one of the above method embodiments are implemented.

[0220] An embodiment of the present application provides a computer device, which includes at least one processor, a memory, and an interface circuit. The memory, the interface circuit, and the at least one processor are interconnected via lines, and instructions are stored in the at least one memory. When the instructions are executed by the processor, some or all of the steps of any one of the above method embodiments are implemented.

[0221] An embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed, part or all of the steps of any one of the above method embodiments are implemented.

[0222] An embodiment of the present application provides a computer program, which includes instructions. When the computer program is executed by a processor, some or all of the steps of any one of the above method embodiments are implemented.

[0223] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. It should be noted that for the aforementioned method embodiments, in order to simplify the description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the order of the actions described, because according to this application, some steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.

[0224] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0225] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0226] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A protein structure prediction method, characterized in that: The method comprises: obtaining a first amino acid sequence; generating a first multiple sequence alignment (MSA) data based on a contextual relationship between a plurality of amino acids in the first amino acid sequence; Based on the first MSA data, N predictions are performed respectively to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

2. The method according to claim 1, characterized in that The generating of first multiple sequence alignment (MSA) data based on the contextual relationship of the plurality of amino acids in the first amino acid sequence comprises: Based on a first machine learning model, contextual relationships among multiple amino acids in the first amino acid sequence are extracted, and the first MSA data is generated based on the contextual relationships.

3. The method according to claim 1, characterized in that The generating of first multiple sequence alignment (MSA) data based on the contextual relationship of the plurality of amino acids in the first amino acid sequence comprises: Based on the first amino acid sequence, searching an amino acid sequence database to obtain second MSA data, wherein each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence; extracting, based on the second machine learning model, a contextual relationship between a plurality of amino acids in the first amino acid sequence and the second MSA data, and processing the contextual relationship and the second MSA data through a third machine learning model to generate third MSA data; The third MSA data or the fourth MSA data is used as the first MSA data; wherein the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

4. The method according to any one of claims 1 to 3, characterized in that The step of performing N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer, includes: Performing N predictions on the first MSA data to obtain N 3D structures of the protein; or Correcting the first MSA data N times respectively to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; Predictions are performed on the N corrected first MSA data respectively to obtain N 3D structures of the protein.

5. The method according to claim 4, characterized in that The first MSA data includes M amino acid sequences, the M amino acid sequences include a third amino acid sequence, and M is a positive integer. The first MSA data is corrected N times to obtain N corrected first MSA data, including: Filling in the missing amino acids in the third amino acid sequence; and / or Replacing the amino acids at P positions in the third amino acid sequence, where P is a positive integer; Among them, the P sites include the first site, the first site of the third amino acid sequence is the first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is the second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

6. The method according to any one of claims 1 to 4, characterized in that The N 3D structures respectively correspond to N prediction quality scores, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

7. The method according to claim 6, characterized in that The method further comprises: Determining the K 3D structures from the N 3D structures according to the prediction quality scores, where K is a positive integer less than or equal to N; The K 3D structures are the K 3D structures with the highest predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the true 3D structure of the protein.

8. The method according to any one of claims 5 to 7, characterized in that The N corrected first MSA data are predicted respectively to obtain N 3D structures of the protein, including: In the prediction process for each corrected first MSA data, the 3D structure of the protein is predicted based on the corrected first MSA data; or Searching a protein database for a template protein structure, wherein the amino acid sequence of the template protein structure contains some identical amino acids to the first amino acid sequence; In the prediction process for each corrected first MSA data, the 3D structure of the protein is predicted based on the template protein structure and the corrected first MSA data.

9. A protein structure prediction method, characterized in that: The method comprises: obtaining a first amino acid sequence; Searching for first MSA data from an amino acid sequence database based on the first amino acid sequence, where the first MSA data includes a third amino acid sequence; Complementing the missing amino acids in the third amino acid sequence and / or replacing the amino acids at P positions in the third amino acid sequence to obtain a corrected first MSA data, where P is a positive integer; Prediction is performed on the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence.

10. The method according to claim 9, characterized in that The P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

11. The method according to claim 9 or 10, characterized in that Predicting the corrected first MSA data to obtain the 3D structure of the protein includes: Predicting the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or Searching a protein database for a template protein structure, wherein the amino acid sequence of the template protein structure contains some identical amino acids to the first amino acid sequence; Based on the template protein structure and the corrected first MSA data, the 3D structure of the protein is predicted.

12. A protein structure prediction device, characterized in that: The device comprises: an acquiring unit, configured to acquire a first amino acid sequence; a processing unit, configured to generate first multiple sequence alignment (MSA) data based on a contextual relationship of each amino acid in the first amino acid sequence; The prediction unit is configured to perform N predictions based on the first MSA data to obtain N 3D structures of the protein corresponding to the first amino acid sequence, where N is a positive integer.

13. The device according to claim 12, characterized in that In the aspect of generating first multiple sequence alignment (MSA) data based on a contextual relationship between a plurality of amino acids in the first amino acid sequence, the processing unit is specifically configured to: Based on a machine learning model, contextual relationships among multiple amino acids in the first amino acid sequence are extracted, and the first MSA data is generated based on the contextual relationships.

14. The device according to claim 12, characterized in that In the aspect of generating first multiple sequence alignment (MSA) data based on a contextual relationship between a plurality of amino acids in the first amino acid sequence, the processing unit is specifically configured to: Based on the first amino acid sequence, searching an amino acid sequence database to obtain second MSA data, wherein each amino acid sequence in the second MSA data contains some identical amino acids with the first amino acid sequence; extracting, based on a machine learning model, a contextual relationship between a plurality of amino acids in the first amino acid sequence and the second MSA data, and generating third MSA data using the contextual relationship and the second MSA data; The third MSA data or the fourth MSA data is used as the first MSA data; wherein the fourth MSA data is obtained by splicing the second MSA data and the third MSA data.

15. The device according to any one of claims 12 to 14, characterized in that The prediction unit is specifically configured to: Performing N predictions on the first MSA data to obtain N 3D structures of the protein; or Correcting the first MSA data N times respectively to obtain N corrected first MSA data, where N is a positive integer greater than or equal to K; Predictions are performed on the N corrected first MSA data respectively to obtain N 3D structures of the protein.

16. The device according to claim 15, characterized in that The first MSA data includes M amino acid sequences, the M amino acid sequences include a third amino acid sequence, and M is a positive integer. In the aspect of performing N corrections on the first MSA data to obtain N corrected first MSA data, the prediction unit is specifically configured to: Filling in the missing amino acids in the third amino acid sequence; and / or Replacing the amino acids at P positions in the third amino acid sequence, where P is a positive integer; Among them, the P sites include the first site, the first site of the third amino acid sequence is the first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is the second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

17. The device according to any one of claims 12 to 16, characterized in that The N 3D structures respectively correspond to N prediction quality scores, and the prediction quality scores are used to characterize the degree of difference between the 3D structure and the true 3D structure of the protein.

18. The device according to claim 17, characterized in that The device further comprises: a determining unit, configured to determine the K 3D structures from the N 3D structures according to the prediction quality scores, where K is a positive integer less than or equal to N; The K 3D structures are the K 3D structures with the highest predicted quality scores among the N 3D structures; the higher the predicted quality score, the higher the ranking, and the smaller the difference between the 3D structure corresponding to the predicted quality score and the true 3D structure of the protein.

19. The device according to any one of claims 15 to 17, characterized in that In the aspect of respectively predicting the N corrected first MSA data to obtain the N 3D structures of the protein, the processing unit is specifically configured to: In the prediction process for each corrected first MSA data, the 3D structure of the protein is predicted based on the corrected first MSA data; or Searching a protein database for a template protein structure, wherein the amino acid sequence of the template protein structure contains some identical amino acids to the first amino acid sequence; In the prediction process for each corrected first MSA data, the 3D structure of the protein is predicted based on the template protein structure and the corrected first MSA data.

20. A protein structure prediction device, characterized in that: The device comprises: an acquiring unit, configured to acquire a first amino acid sequence; a processing unit, configured to search an amino acid sequence database for first MSA data based on the first amino acid sequence, wherein the first MSA data includes a third amino acid sequence; A correction unit is used to fill in the missing amino acids in the third amino acid sequence and / or replace the amino acids at P positions in the third amino acid sequence to obtain the corrected first MSA data, where P is a positive integer The prediction unit is used to predict the corrected first MSA data to obtain the 3D structure of the protein corresponding to the first amino acid sequence.

21. The device according to claim 20, characterized in that The P sites include a first site, the first site of the third amino acid sequence is a first amino acid, each of the Q amino acid sequences contained in the M amino acid sequences is a second amino acid at the first site, and the type of the second amino acid is different from the type of the first amino acid, and Q is a positive integer less than or equal to M.

22. The device according to claim 20 or 21, characterized in that The prediction unit is specifically configured to: Predicting the corrected first MSA data based on a structure prediction model to obtain the 3D structure of the protein; or Searching a protein database for a template protein structure, wherein the amino acid sequence of the template protein structure contains some identical amino acids to the first amino acid sequence; Based on the template protein structure and the corrected first MSA data, the 3D structure of the protein is predicted.

23. A computer device, characterized in that: The computer device includes at least one processor, a memory and an interface circuit, the memory, the interface circuit and the at least one processor are interconnected via lines, and instructions are stored in the at least one memory; when the instructions are executed by the processor, the method described in any one of claims 1 to 11 is implemented.

24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 11 is implemented.

25. A computer program product, characterized in that The computer program product comprises instructions, and when the instructions are executed, the method according to any one of claims 1 to 11 is implemented.