Antibody sequence detection method, medium and device
By integrating the homologous sequence characteristics and language model characteristics of the fusion antibody sequence and combining with the deep learning model, the problem of three-dimensional structure detection of antibodies is solved, the structural detection accuracy of the CDR-H3 loop is improved, and the binding ability of the antibody to antigen is enhanced.
Patent Information
- Application Number
- CN202510026906.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
It is difficult for prior art to accurately detect the three-dimensional structure of an antibody, especially in the complementary determination region (CDR), which affects the affinity and specificity of the antibody to the antigen.
By obtaining the homologous sequence characteristics and language model characteristics of the antibody sequence and performing fusion processing, combining deep learning models such as AlphaFold2 and AntiBERTy, the three-dimensional structure of the antibody is detected.
The accuracy of antibody sequence structure detection, especially the accuracy of CDR-H3 loop, enhances the specificity and affinity of the antibody and antigen.
Smart Images

Figure CN119943129A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of antibody detection, and in particular to an antibody sequence detection method, medium and device. Background Art
[0002] Antibodies are immunoglobulins produced by B lymphocytes and are a core component of the adaptive immune system. Antibodies recognize and neutralize antigens through specific binding mechanisms. Therefore, in the development and application of antibody drugs, accurately detecting the three-dimensional structure of antibodies is of great significance for improving the affinity and specificity of antibodies to antigens. Among them, the sequence and structural diversity of the complementarity determining region (CDR) of antibodies is a key factor in determining the specificity and affinity between antibodies and antigens. However, due to the extremely diverse complementarity determining regions of antibodies and the relatively few antibody samples with known structures, the structural detection of the complementarity determining regions of antibodies is difficult, which affects the detection accuracy of the three-dimensional structure of antibodies. Summary of the invention
[0003] The embodiments of the present application provide an antibody sequence detection method, medium and device, which can improve the accuracy of antibody structure detection.
[0004] In a first aspect, an embodiment of the present application provides an antibody sequence detection method, the method comprising: obtaining a first antibody sequence, the first antibody sequence comprising a paired heavy chain and a light chain; obtaining a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and a homologous sequence, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate the dependency relationship between amino acids in the first antibody sequence; fusing the first feature and the third feature to obtain a fusion feature of the first antibody sequence; and detecting and obtaining a three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature.
[0005] As an example, the first feature of the first antibody sequence can be called a homologous sequence feature, that is, the evolutionary constraint information between the first antibody sequence and the homologous sequence, which can be used to detect the three-dimensional structure of amino acids in the first antibody sequence. As an example, the third feature of the first antibody sequence can also be called a language model feature, which can be used to detect the three-dimensional structure of amino acids in the first antibody sequence.
[0006] Therefore, the antibody sequence structure detection method provided by the present application can be achieved by integrating the homologous sequence features and language model features of the first antibody sequence to be detected, that is, not only the influence of the homologous sequence features of the antibody sequence on the structure is considered, but also the influence of the language model features of the antibody sequence on the structure is considered. Thus, by comprehensively considering the influence of the homologous sequence features and the language model on the structure of the first antibody sequence, the structural detection accuracy of the CDR-H3 loop of the complementary determining region (CDR) can be improved, thereby improving the structural detection accuracy of the antibody sequence as a whole.
[0007] In a possible implementation of the first aspect, the above-mentioned fusion processing of the first feature and the third feature to obtain the fusion feature of the first antibody sequence includes: converting the root mean square error information of the first antibody sequence into a scaling factor, the root mean square error information includes the root mean square error of each amino acid main chain atom in the first antibody sequence, and the dimension of the scaling factor is one-dimensional; converting the dimension of the third feature according to the scaling factor to obtain a fourth feature, wherein the fourth feature has the same dimension as the first feature; and summing the fourth feature and the first feature to obtain the fusion feature. In this way, the first feature and the third feature of the first antibody sequence can be fused through the first root mean square error information to obtain the fusion feature.
[0008] In a possible implementation of the first aspect, the first feature and the second feature are determined by: searching for homologous sequences of the first antibody sequence; obtaining a multiple sequence comparison representation between the first antibody sequence and the homologous sequence; connecting the heavy chain and the light chain in the first antibody sequence into a sequence to obtain a first protein sequence; pairing different amino acids in the first protein sequence to obtain a pairing representation of the first protein sequence; encoding the multiple sequence comparison representation and the pairing representation to obtain the first feature and the second feature, respectively. As an example, the first feature and the third feature can be determined based on the homologous sequence of the first antibody sequence through transfer learning using a deep learning model such as AlphaFold2 or AlphaFold-Multimer.
[0009] In a possible implementation of the first aspect, the third feature is determined by: connecting the heavy chain and the light chain in the first antibody sequence into a sequence through a gap to obtain a second protein sequence; inputting the second protein sequence into a pre-trained protein language model, outputting the first language model encoding and the first attention matrix; updating the first language model encoding and the first attention matrix based on the attention mechanism to obtain the third feature. As an example, the third feature can be obtained by using a protein language model such as AntiBERTy through transfer learning to capture the complex dependencies between amino acids in the first antibody sequence.
[0010] In a possible implementation of the first aspect, the above-mentioned detection of the three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature includes: updating the fusion feature and the second feature based on the attention mechanism; and detecting the three-dimensional structure of the first antibody sequence based on the updated second feature and the updated fusion feature. In this way, updating the second feature and the fusion feature based on the attention mechanism is conducive to improving the accuracy of the detected three-dimensional structure of the first antibody sequence.
[0011] In a possible implementation of the first aspect above, according to the updated second feature and the updated fusion feature, the three-dimensional structure of the first antibody sequence is detected, including: according to the updated second feature and the updated fusion feature, the mean and variance of the translation parameters of the amino acids in the first antibody sequence are detected, and the mean and variance of the rotation parameters of the amino acids in the first antibody sequence are detected; using a reparameterization method, according to the mean and variance of the translation parameters and the mean and variance of the rotation parameters, the target value of the translation parameters and the target value of the rotation parameters of each amino acid in the first antibody sequence are determined, wherein the translation parameter represents the position of the amino acid, the rotation parameter represents the angle of the amino acid, and the three-dimensional structure includes the position and angle of the amino acid. It can be understood that the structure of the CDR loop of the antibody is dynamic and may be affected by antibodies and environmental factors, resulting in structural changes. Therefore, when predicting the antibody structure, the present application does not directly predict the specific values of the translation and rotation of each amino acid, but predicts the mean and variance of each of the two, and then obtains the translation and rotation of each amino acid by reparameterization, which is conducive to improving the accuracy of structural detection.
[0012] In a possible implementation of the first aspect, different amino acids in the first antibody sequence have different position codes. It is understandable that in the process of performing structural detection on the first antibody sequence, the positions of different amino acids in the first antibody sequence can be identified based on the amino acid position codes.
[0013] In a possible implementation of the first aspect above, the three-dimensional structure is determined based on a structure detection model, the structure detection model includes a feature fusion module, the feature fusion module includes a confidence conversion submodule and a feature transformation submodule, the confidence conversion submodule includes three groups of network layers connected in sequence, each group of network layers includes a linear layer and an activation layer connected in sequence, and the feature transformation submodule includes an aggregation layer, a linear layer, an activation layer, a summation layer, a linear layer and an activation layer connected in sequence, and the fusion feature is determined based on the feature fusion module. For example, the aggregation layer in the feature transformation submodule 142 is a Multiplate layer, the activation layer is a Relu layer, and the linear layer can be a fully connected layer FC.
[0014] In a possible implementation of the first aspect above, the process of determining the fusion feature by the feature fusion module includes: inputting the root mean square error information into the first linear layer of the confidence conversion submodule, so that the last activation layer of the confidence conversion submodule outputs the scaling factor; inputting the scaling factor and the third feature into the aggregation layer in the feature conversion submodule, and inputting the first feature into the summation layer in the feature conversion submodule, so that the last activation layer in the feature conversion submodule outputs the fusion feature.
[0015] In a possible implementation of the first aspect, the three-dimensional structure is determined based on a structure detection module in the structure detection model, and the structure detection module includes a variational autoencoder. In this way, the dynamics of the antibody structure are modeled by a variational autoencoder, which is conducive to improving the accuracy of structure detection.
[0016] In a possible implementation of the first aspect above, the first feature is a tensor with a shape of N×384, the third feature is a tensor with a shape of N×64, and the second feature is a tensor with a shape of N×N×128, where N is the total length of amino acids in the heavy chain and the light chain in the first antibody sequence.
[0017] In a possible implementation of the first aspect, the method further includes: training the structure detection model to be trained using a teacher-student self-supervised learning method to obtain a trained structure detection model. The present application adopts a teacher-student self-supervised learning method, which can introduce a large number of unlabeled antibody sequences, train them in a self-supervised manner, and use them together with labeled antibody sequences for model training, thereby improving the prediction ability of the model.
[0018] In a possible implementation of the first aspect, a structure detection model to be trained is trained by using a teacher-student self-supervised learning method to obtain a trained structure detection model, including: based on a training data set, a structure detection model to be trained is trained by using a teacher-student self-supervised learning method to obtain a trained structure detection model, wherein the training data set includes at least one training data group, and the training data group includes an unlabeled antibody sequence and a labeled antibody sequence. In this way, a small amount of labeled structural information can be effectively propagated to a large number of unlabeled antibody sequences, thereby improving the model's ability to predict the structure of antibody sequences.
[0019] In a possible implementation of the first aspect, the training process of the structure detection model includes: obtaining a teacher model and a student model corresponding to the structure detection model to be trained; inputting the first unlabeled training antibody sequence in the training data group into the student model to obtain a first detection result, and inputting the first training antibody sequence into the teacher model to obtain a second detection result; calculating a first loss according to the first detection result and the second detection result; inputting the second labeled training antibody sequence in the training data group into the student model to obtain a third detection result; calculating a second loss according to the third detection result and the true label of the second training antibody sequence; calculating the sum of the first loss and the second loss to obtain a third loss; updating the parameters of the student model by back-propagation gradient according to the third loss; updating the parameters of the teacher model by performing exponential sliding average processing on the updated parameters of the student model. As an example, in the training process, the labeled antibody sequences and the unlabeled antibody sequences each account for half, the labeled antibody data uses the natural structure (i.e., the true label) as the supervisory signal, and the unlabeled training antibody sequence uses the three-dimensional structure predicted by the teacher model as the supervisory signal.
[0020] In a second aspect, an embodiment of the present application provides a readable medium having instructions stored thereon, which, when executed on an electronic device, causes the electronic device to execute the antibody sequence detection method in the first aspect and any possible implementation thereof.
[0021] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the antibody sequence detection method as in the first aspect and any possible implementation thereof.
[0022] The beneficial effects of the second to third aspects of the present application can be referred to the description of the first aspect and its various possible implementation methods, and will not be elaborated on here. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 According to some embodiments of the present application, a structural block diagram of an antibody structure detection model is shown;
[0024] Figure 2 According to some embodiments of the present application, a schematic diagram of a homologous sequence feature extraction process is shown;
[0025] Figure 3 According to some embodiments of the present application, a schematic diagram of a language model feature extraction process is shown;
[0026] Figure 4 According to some embodiments of the present application, a schematic diagram of a feature fusion process is shown;
[0027] Figure 5 According to some embodiments of the present application, a schematic flow chart of an antibody sequence detection method is shown;
[0028] Figure 6 According to some embodiments of the present application, a schematic flow chart of an antibody sequence detection method based on a structure detection model is shown;
[0029] Figure 7 According to some embodiments of the present application, a schematic flow chart of an antibody sequence detection method is shown;
[0030] Figure 8 According to some embodiments of the present application, a schematic diagram of a training process of an antibody structure detection model is shown;
[0031] Fig. 9 According to some embodiments of the present application, a schematic diagram of a network architecture for model training based on a teacher-student self-supervised learning method is shown;
[0032] Fig.10 According to some embodiments of the present application, a schematic diagram of a training process for a structure detection model based on a training data set is shown;
[0033] Fig.11A According to some embodiments of the present application, box plots of prediction accuracy of various antibody structure prediction methods are shown;
[0034] Fig. 11B According to some embodiments of the present application, a schematic diagram showing test results of various models on a test set is shown;
[0035] Fig. 11C According to some embodiments of the present application, a schematic diagram showing test results of various models on a test set is shown;
[0036] Fig. 12A According to some embodiments of the present application, a scatter plot comparing AbFold and AlphaFold2-Multimer is shown;
[0037] Fig. 12B According to some embodiments of the present application, a scatter plot comparing AbFold and IgFold is shown;
[0038] Fig. 12C According to some embodiments of the present application, a schematic diagram comparing the CDR-H3 structures predicted by the AbFold, AlphaFold2 and IgFold models with the native structure is shown;
[0039] Fig.13According to some embodiments of the present application, a schematic structural diagram of an electronic device is shown. DETAILED DESCRIPTION
[0040] Illustrative embodiments of the present application include, but are not limited to, antibody sequence structure detection methods, media, and electronic devices.
[0041] In order to more clearly describe the embodiments of the present application, some terms provided in the embodiments of the present application are first introduced below.
[0042] Antibodies, also known as immunoglobulins, are proteins with immune functions produced in the serum of humans and animals due to the invasion of pathogens or viruses. They are widely used in medicine, biology and other fields.
[0043] Antigen refers to a substance that can stimulate an organism to produce an immune response and can bind to lymphocytes such as antibodies in vivo and in vitro to produce an immune effect (specific reaction). The binding between antibodies and antigens has specificity and affinity.
[0044] Specificity refers to the high selectivity of the binding reaction between antigen and antibody, that is, an antibody can only bind to one or a specific type of antigen, and an antigen can only bind to one or a specific type of antibody. The specific binding of antigen and antibody is based on the complementary structure of their molecular surfaces.
[0045] Affinity refers to the binding strength between an antigen binding site on an antibody and a corresponding antigen epitope, which depends on the degree of complementarity of the spatial structures of the two.
[0046] The antibody sequence refers to the order of amino acids in the antibody, which determines the structure and function of the antibody. In the following embodiments, the structure of the antibody can also be described as the structure of the antibody sequence of the antibody.
[0047] The conformation of an antibody usually refers to the relative arrangement and orientation of atoms or groups in space within the molecule, which determines the shape and three-dimensional structure of the molecule. For antibodies, their conformation refers to the folding mode of the peptide chains (including heavy chains and light chains) within the antibody molecule, the connection mode of the disulfide bonds between the chains, and the relative positions of the variable and constant regions.
[0048] The three-dimensional structure of an antibody refers to the specific shape and arrangement of the molecule in three-dimensional space. For antibodies, their three-dimensional structure is directly observed and analyzed through modern biophysical techniques such as X-ray crystallography and nuclear magnetic resonance (NMR). The three-dimensional structure of an antibody shows its fine molecular conformation, including the folding of peptide chains, the interactions between and within chains, and the specific location and configuration of variable and constant regions. This structural information is of great significance for understanding the antigen binding mechanism, biological function, and design and application of antibody engineering.
[0049] Specifically, antibodies are large protein molecules produced by B lymphocytes, and their monomers can be a Y-shaped structure, consisting of two identical heavy chains and two identical light chains connected by disulfide bonds. As an example, antibodies are usually folded into a ring structure by heavy chains and light chains, for example, a Y-shaped antibody contains two symmetrical pairs of heavy and light chains. Accordingly, the antibody sequence can be a sequence composed of amino acids in paired heavy and light chains.
[0050] Among them, the key functional region of the antibody is located in the variable region (Fv) at the top of the structure in the ring structure, which is responsible for the specific recognition of the antibody. The variable region contains the complementarity determining region (CDR) and the framework region (Fr), which is the part of the variable region outside the complementarity determining region. The complementarity determining region is composed of six highly variable loops: three loops (CDR-H1, CDR-H2, CDR-H3) are located in the heavy chain, and three loops (CDR-L1, CDR-L2, CDR-L3) are located in the light chain. These loops are responsible for the recognition and binding of antigens. The sequence and structural diversity of the complementary determining region is a key factor in determining the specificity and affinity of the antibody. In particular, the structure of the CDR-H3 loop is extremely diverse. The CDR-H3 loop is located at the junction of the heavy chain and the light chain, and the structure is affected by the interaction between the two chains. Therefore, the structural detection of the CDR-H3 loop is more difficult.
[0051] Amino acids are the basic units that make up proteins. Each amino acid contains one or more amino groups (—NH2) and carboxyl groups (—COOH), as well as a specific side chain group (R group). During protein synthesis, amino acids are linked by peptide bonds to form polypeptide chains, which are then folded into proteins with specific functions, such as antibodies.
[0052] An amino acid pair refers to a combination of two adjacent or non-adjacent amino acids in a protein sequence. In the structure and function of a protein, the interaction between amino acid pairs is one of the key factors that determine the three-dimensional structure and function of a protein. These interactions include hydrogen bonds, hydrophobic interactions, ionic bonds, and van der Waals forces, which together maintain the stability and function of the protein.
[0053] Amino acid residue refers to the remaining structural part of the amino acids that make up the polypeptide when they combine with each other, because some of their groups (such as amino and carboxyl groups) participate in the formation of peptide bonds and lose a molecule of water. In other words, the amino acid residue is the part of the amino acid that remains after the peptide bond is formed. In the protein sequence, each amino acid residue occupies a specific position and is connected to other amino acid residues through peptide bonds.
[0054] As mentioned above, the structure of the complementary determining region, especially the CDR-H3 loop, is difficult to detect, which affects the detection accuracy of the three-dimensional structure of the antibody.
[0055] In some embodiments, deep learning models such as AlphaFold2 or AlphaFold-Multimer can determine homologous sequence features such as evolutionary constraint information of the antibody sequence to be detected based on the homologous sequence, and detect the three-dimensional structure (i.e., protein structure) of the antibody sequence based on the homologous sequence features. Among them, homologous sequences refer to amino acid sequences with a common evolutionary ancestor. The evolutionary trajectory of the amino acid sequence of a protein (such as an antibody) is limited by its function, and its evolutionary constraints can be inferred through a set of homologous sequences to detect the protein structure.
[0056] In other embodiments, deep learning models such as IgFold can determine the language model features of the antibody sequence to be detected based on pre-trained protein language models such as BERT, AntiBERTy, etc., and detect the three-dimensional structure of the antibody sequence based on the language model features. Among them, the above-mentioned language model features can characterize the complex dependencies between amino acids in the antibody sequence. It can be understood that the protein language model works by simulating the interdependencies between amino acids in proteins. Specifically, the model uses a neural network architecture called Transformer to learn sequence dependencies of arbitrary length contexts from data. When processing antibody sequences, the protein language model can analyze the position, type, and interaction relationship of each amino acid with other amino acids, thereby capturing the dependencies between amino acids.
[0057] However, for the same antibody, the above-mentioned detection method based on homologous sequences and the detection method based on language model features often have large differences in the detection of CDR loops in antibodies, such as CDR-H3 loop structures. Moreover, these two detection methods will appear in many antibodies. One method is accurate, while the other method has a large error. It can be understood that the detection method based on homologous sequences only considers the influence of the homologous sequence characteristics of the antibody sequence on the structure, while the detection method based on language model features only considers the influence of the language model characteristics of the antibody sequence on the structure. However, the structures of different antibody sequences are usually affected by homologous sequence characteristics and by language model characteristics. Then, when the structure of the antibody sequence is highly affected by the homologous sequence characteristics, the error of detecting the structure using only the language model characteristics is large. Conversely, when the structure of the antibody sequence is highly affected by the language model characteristics, the error of detecting the structure using only the homologous sequence characteristics is large.
[0058] Therefore, in order to improve the accuracy of structural detection of the CDR H3 loop in an antibody, an embodiment of the present application provides an antibody sequence detection method. Specifically, the method includes: obtaining a first antibody sequence, the first antibody sequence includes a paired heavy chain and a light chain, that is, the antibody sequence includes amino acids in three loops in the heavy chain and three loops in the light chain; obtaining a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate the evolutionary constraint information between the first antibody sequence and the homologous sequence, the second feature is used to indicate the relationship between the amino acid pairs in the first antibody sequence (i.e., the pairing relationship of the amino acids), and the third feature is used to indicate the dependency relationship between the amino acids in the first antibody sequence; the first feature and the third feature are fused to obtain the fusion feature of the first antibody sequence; and the three-dimensional structure of the first antibody sequence is detected according to the second feature and the fusion feature.
[0059] As an example, the first feature and the third feature can be determined based on the homologous sequence of the first antibody sequence by a deep learning model such as AlphaFold2 or AlphaFold-Multimer. At this time, the first feature of the first antibody sequence (also called the homologous sequence feature), that is, the evolutionary constraint information between the first antibody sequence and the homologous sequence, can be used to detect the three-dimensional structure of amino acids in the first antibody sequence.
[0060] As an example, the third feature can be obtained by capturing the complex dependencies between amino acids in the first antibody sequence using a protein language model such as AntiBERTy. At this time, the third feature of the first antibody sequence (also called a language model feature) can be used to detect the three-dimensional structure of amino acids in the first antibody sequence.
[0061] Thus, the antibody sequence structure detection method provided by the present application can improve the structure detection accuracy of the CDR H3 loop by fusing the homologous sequence features and language model features of the first antibody sequence to be detected, combining the advantages of the two methods, thereby improving the overall structure detection accuracy of the antibody sequence. Specifically, the present application can perform antibody structure detection based on the homologous sequence features and language model of the first antibody sequence, so as to comprehensively consider the influence of the homologous sequence features and the language model on the structure of the first antibody sequence, which is conducive to reducing detection errors.
[0062] In some embodiments of the present application, the subject that performs the antibody sequence detection method in the present application may be an electronic device or a device in an electronic device for performing antibody structure detection.
[0063] As an example, the electronic devices applicable to the present application include but are not limited to mobile phones, tablet computers, wearable electronic devices, vehicle-mounted electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. Of course, the above electronic devices are not limited to the above examples, and also include hardware servers, cloud servers and other devices, which are not limited in the present application.
[0064] In some embodiments, the electronic device in the present application can perform an antibody sequence detection method through an antibody structure detection model based on deep learning. For example, the antibody structure detection model provided in the present application can be called AbFold, or a detection network, etc. Among them, the electronic devices used by the antibody structure detection model of the present application in the model training phase and the model use phase are the same or different, and the present application does not specifically limit this.
[0065] Next, combine Figures 1 to 4 The structure and function of the pre-trained structure detection model of this application are introduced.
[0066] Reference Figure 1 , is a structural block diagram of an antibody structure detection model provided for the implementation of this application. Figure 1 As shown, the structure detection model 01 includes the following modules: a homology feature extraction module 11, a language model feature extraction module 12, a root mean square error determination module 13, a feature fusion module 14, a feature update module 15, an amino acid position encoding module 16 and a structure detection module 17.
[0067] The homology feature extraction module 11 is used to search for homologous sequences of an input antibody sequence such as a first antibody sequence, extract evolutionary constraint information between the input antibody sequence and the homologous sequence (i.e., the first feature), and extract the correlation between amino acid pairs in the input antibody sequence (i.e., the second feature).
[0068] The language model feature extraction module 12 is used to extract the dependency relationship between amino acids in the antibody sequence (ie, the third feature).
[0069] The root mean square error determination module 13 is used to calculate the root mean square error of each amino acid main chain atom (C α , C β , N, O), the root mean square error (RMSE) can be recorded as pRMSD. The root mean square error is also called root mean square deviation (RMSD), which is a commonly used measure of the difference between measurement values.
[0070] The feature fusion module 14 is used to fuse the first feature and the third feature of the input antibody sequence to obtain a fused feature. For example, the feature fusion module 14 can fuse the first feature and the third feature of the input antibody sequence based on the root mean square error of each amino acid main chain atom in the input antibody sequence. The specific process of the fusion process will be described below and will not be described in detail here.
[0071] The feature updating module 15 is used to update the fusion feature and the second feature of the input antibody sequence based on an attention mechanism such as a triangular attention mechanism.
[0072] The amino acid position coding module 16 is used to perform position coding on the amino acids in the input antibody sequence to obtain the amino acid position coding features of the input antibody sequence.
[0073] The structure detection model 17 is used to perform antibody structure detection on the input antibody sequence based on the third feature and the fusion feature of the input antibody sequence to obtain the three-dimensional structure of the antibody sequence. For example, the structure detection model 17 can perform antibody structure detection on the input antibody sequence based on the third feature and the fusion feature and the amino acid position coding feature of the input antibody sequence.
[0074] Understandably, Figure 1 The structure of the structure detection model 01 shown is only an example, and in some other embodiments, the structure detection model 01 may include more or fewer modules. For example, in some embodiments, the structure detection model 01 may not include modules such as the amino acid position encoding module 16.
[0075] In some embodiments, the structure detection model 01 in the present application, such as AbFold, can extract features such as evolutionary constraint information between the input antibody sequence and the homologous sequence, the correlation between amino acid pairs, and the dependency between amino acids through transfer learning, that is, extract the first feature, the second feature, and the third feature of the input antibody sequence through transfer learning.
[0076] As an example, the homologous feature extraction module 11 can use AlphaFold2 to extract the first feature and the second feature of the input antibody sequence through transfer learning, for example, using the Evoformer module (an encoder) in AlphaFold2 to extract the first feature and the second feature.
[0077] It can be understood that transfer learning can transfer the knowledge learned in the source domain to the target domain, thereby solving the problem of insufficient labeled data in the target domain. In the task of antibody structure prediction, due to the limited number of high-quality antibody structures, it is difficult for neural networks to fully learn the structural characteristics of antibodies from a small amount of labeled data. Therefore, this application uses general protein structure prediction as the source domain and antibody structure prediction as the target domain, and extracts the key features of antibodies through transfer learning, which can improve the accuracy of antibody structure prediction. Among them, AlphaFold2 is an advanced model in the field of general protein structure prediction. Its training process uses a large number of known structures and protein structures generated by self-distillation.
[0078] In some embodiments, the homologous sequence feature extraction module 11 can use AlphaFold2-Multimer to extract the first feature, the second feature and other features from the homologous sequence of the input antibody sequence through transfer learning. Among them, AlphaFold2-Multimer provides five sets of model parameters, and the structure detection model 01 in this application uses the first set of parameters to predict the structure and extract features. Since the homologous sequence search of AlphaFold2-Multimer is time-consuming (about one hour for each antibody), this application can adopt the acceleration solution in ColabFold (a deep learning model), that is, use the search tools in this solution such as MMseqs2 to search protein databases such as Uniref90, and shorten the homologous sequence search time for each antibody to minutes.
[0079] Reference Figure 2 , is a flow chart of homologous sequence feature extraction provided in an embodiment of the present application, and the extraction process can be executed by the above homologous feature extraction module 11. Specifically, Figure 2The process shown includes: the homology feature extraction module 11 obtains the input antibody sequence (input sequence) composed of the amino acid composition of the paired heavy chain and light chain, uses the search tool MMseqs2 to search the Uniref90 database, determines the homologous sequence of the input antibody sequence, and generates a multiple sequence alignment (MSA) representation between the input antibody sequence and the homologous sequence, Figure 2 The dark blocks in the MSA representation shown represent sequences with the same evolutionary constraint information. In addition, the homology feature extraction module 11 can connect the amino acids of the heavy chain and the light chain in the input antibody sequence into a sequence by using AlphaFold2-Multimer through transfer learning, and perform pairing processing on the amino acids in the sequence to obtain a pair representation (pair repr, i.e., pair representation) of the sequence. Figure 2 The pair representation includes N×N dimensional data blocks, each of which corresponds to two amino acids in the input antibody sequence.
[0080] In addition, the homologous feature extraction module 11 can encode the MSA representation and pair representation of the input antibody sequence using the Evoformer module (encoder) in AlphaFold2-Multimer through transfer learning, such as encoding the output tensors S1 and Z multiple times (such as 3 times). Among them, the shape of tensor S1 can be N×384, representing the multidimensional features of the input antibody sequence, specifically the evolutionary constraint information between the input antibody sequence and the homologous sequence. The shape of tensor z can be N×N×128, representing the pairing information between the two amino acids in the input antibody sequence, that is, the correlation between the amino acid pairs of the input antibody sequence. N is the sum of the lengths of the heavy and light chains in the input antibody sequence, that is, the sum of the numbers in the amino acids. At this time, tensor S1 represents the first feature of the input antibody sequence, and tensor z represents the second feature of the input antibody sequence.
[0081] In addition, as an example, the language model feature extraction module 12 can use the protein language model AntiBERTy and the feature encoding module in IgFold through transfer learning to extract the third feature of the input antibody sequence, namely the dependency between amino acids. It can be understood that AntiBERTy is a protein language model pre-trained based on a large number of natural antibody sequences.
[0082] Reference Figure 3 , is a flow chart of language model feature extraction provided in an embodiment of the present application, and the extraction process can be executed by the above-mentioned language model feature extraction module 12. Specifically, Figure 3The process shown includes: the language model feature extraction module 12 connects the heavy chain and the light chain of the input antibody sequence into a sequence through a space (or interval), inputs the sequence into the protein language model AntiBERTy, and generates the language model encoding and attention matrix of the sequence. Furthermore, the language model feature extraction module 12 can use the feature encoding module of IgFold through transfer learning to interact with the generated language model encoding and attention matrix, and further extract the tensor S2 of the input antibody sequence. For example, the dimension of the tensor S2 is N×64. At this time, the tensor S2 represents the third feature of the input antibody sequence, that is, the tensor S2 is used to indicate the dependency relationship between amino acids in the input antibody sequence.
[0083] In some embodiments, the root mean square error calculation module 13 can use IgFold to predict the root mean square error (pRMSD) of each amino acid main chain atom in the input antibody sequence through transfer learning. It can be understood that the root mean square error (RMSD) is usually used to measure the deviation between the structure predicted by the model and the true structure, thereby reflecting the accuracy of each part of the structure detection model to evaluate the credibility of the predicted structure. Studies have shown that the predicted pRMSD is highly correlated with the actual error of the structure detection, so it can be used as a basis for evaluating the credibility of the predicted structure. The present application can use IgFold to obtain the pRMSD of the predicted three-dimensional structure of the input antibody sequence to evaluate the credibility of the antibody sequence features extracted from the protein language model, such as the second feature, and use it in the subsequent feature fusion process. In addition, the dimension of the root mean square error (pRMSD) of each amino acid main chain atom in the input antibody sequence is N×4.
[0084] In some embodiments, the feature fusion module 14 includes a confidence conversion submodule and a feature transformation submodule. The confidence conversion submodule includes three groups of network layers connected in sequence, and each group of network layers includes a linear layer and an activation layer connected in sequence. For example, the linear layer in the confidence conversion submodule 141 is a fully connected layer (FC layer), and the activation layer is a Sigmoid layer. In addition, the feature transformation submodule 142 includes a polymerization layer, a linear layer, an activation layer, a summation layer, a linear layer and an activation layer connected in sequence. For example, the polymerization layer in the feature transformation submodule 142 is a Multiplate layer, the activation layer is a Relu layer, and the linear layer can be a fully connected layer FC.
[0085] Reference Figure 4 , is a feature fusion flow chart provided in an embodiment of the present application, and the process can be executed by the feature fusion module 14. Figure 4 As shown, the feature fusion module 14 includes a confidence conversion submodule 141 and a feature transformation submodule 142 .
[0086] As an example, Figure 4 The confidence conversion submodule 141 in the example includes a FC(4,32) layer, a Sigmoid layer, a FC(32,32) layer, a Sigmoid layer, a FC(32,1) layer, and a Sigmoid layer in sequence. Figure 4 The feature conversion submodule 142 in the example includes a Multiplate layer, a FC (64, 384) layer, a Relu layer, an Add layer, a FC (384, 384) layer, and a Sigmoid layer. Specifically, Figure 4 As shown, the root mean square error pRMSD with a dimension of N×4 is input into the confidence conversion submodule 141, and is converted into a 1-dimensional scaling factor through multiple layers of the confidence conversion submodule 141. Then, the scaling factor and the tensor S2 of the input antibody sequence are input into the Multiplate layer in the feature conversion submodule 142, and the tensor S1 of the input antibody sequence is input into the Add layer in the feature conversion submodule 142, so that the tensor S2 passes through the FC(64,384) layer and the Relu layer and is summed with the tensor S1 in the Add layer. Then, the summed tensor is converted through the FC(64,384) layer and the Sigmoid layer and outputs the tensor S. At this time, the tensor S is the fusion feature of the input antibody sequence, and the shape of the tensor S is N×384.
[0087] In some embodiments, the feature updating module 15 may update the fused features of the input antibody sequence, such as the tensor S, and the second features, such as the tensor z, based on an attention updating mechanism.
[0088] As an example, the feature updating module 15 can update the integrated features and the second features of the input antibody sequence based on a graph attention network (GAT) using a graph attention mechanism and a triangle update mechanism.
[0089] Among them, the graph attention network is a model based on graph neural network, which is used to process graph structure data. The graph attention network optimizes the update process of node features by introducing the attention mechanism. In the graph attention network, each node calculates the attention weight based on the features of its neighboring nodes. These weights are used to weight the features of the aggregated neighboring nodes to update the feature representation of the current node.
[0090] In some embodiments, the amino acid position coding module 16, when performing position coding on the input antibody sequence, starts the coding of the heavy chain and the light chain from 0, respectively, and increases the position coding by 1 for each amino acid. To distinguish between heavy chains and light chains, each position coding of the light chain is uniformly added with 256 after the initial coding is completed. The final dimension of the amino acid position coding feature of the input antibody sequence is N×1. It can be understood that the type and position of amino acids can uniquely determine an antibody sequence, and accurate amino acid position coding can significantly improve the accuracy of structure prediction. In addition, the amino acid position coding can also help the structure detection module 17 based on the Transformer architecture to identify the positions of different amino acids in the input antibody sequence.
[0091] In some embodiments, the structure detection module 17 can use a variational autoencoder to model the dynamics of the antibody structure of the input antibody sequence based on the tensor S of the fusion feature and the tensor z of the second feature after the input antibody sequence is updated. The structure of the input antibody sequence, i.e., the three-dimensional structure, includes the translation parameters and rotation parameters of each amino acid, the translation parameters are used to indicate the position of the amino acid, and the rotation parameters are used to indicate the position of the amino acid.
[0092] In addition, in some embodiments, the structure detection module 17 can use the invariant point attention mechanism to detect the antibody structure. It can be understood that the core idea of the invariant point attention mechanism is to maintain attention to certain key points when processing data with spatial or structural features, and to aggregate and reason about information based on these key points. In protein structure prediction, this mechanism is used to calculate each part of the protein separately to construct an accurate three-dimensional structure.
[0093] As an example, the implementation process of the invariant point attention mechanism includes key point selection, attention weight calculation, information aggregation, and three-dimensional structure construction. Specifically, in the key point selection process, the model determines which atoms or molecular structure points are key points. These key points are usually key residues in the protein structure or atoms with specific functions. In the attention weight calculation process, for each key point, the model calculates its attention weight with the surrounding atoms or structure points. These weights reflect the degree of association and importance between the key point and the surrounding points. In the information aggregation process, based on the attention weight, the model aggregates the information of the surrounding points to the key points, thereby forming a more accurate and robust structural representation. In the three-dimensional structure construction process: the model constructs the three-dimensional structure of the protein based on these key points and their aggregated information.
[0094] It is understandable that the study found that the structure of the CDR loop of the antibody is dynamic and may be affected by antibodies and environmental factors, resulting in structural changes. Therefore, the present application is not suitable for traditional deterministic structure prediction networks, but uses variational autoencoders to model the dynamics of antibody structure. Specifically, the present application represents the structure of the antibody by predicting the translation and rotation of amino acids in the antibody. When predicting the antibody structure, the present application does not directly predict the specific values of the translation and rotation of each amino acid, but predicts the respective means and variances of the two, and then obtains the translation and rotation of each amino acid by reparameterization.
[0095] As an example, the structure detection module 17 in the present application can use formula (1) to calculate the translation parameter and rotation parameter of each amino acid in the input antibody sequence.
[0096]
[0097] Where μ is the mean, σ is the variance, and ∈ is a random number sampled from a standard normal distribution.
[0098] For example, the structure detection module 17 can first predict the mean μ and variance σ of the translation parameters of the amino acids in the input antibody sequence. Then, for the translation parameter of an amino acid in the input antibody sequence, the mean μ and the variance σ and a random number ∈ can be used to calculate the translation parameter of the amino acid using formula (1). Similarly, the structure detection module 17 can determine the rotation parameters of each amino acid in the input antibody sequence by reparameterization.
[0099] Next, combine Figures 5 to 7 , the method flow of the model usage phase of the deep learning-based antibody structure detection model in this application is described in detail.
[0100] Reference Figure 5 As shown, it is a schematic diagram of a flow chart of an antibody sequence detection method provided for the implementation of the present application. The execution subject of the method can be an electronic device or a device in the electronic device.
[0101] Specifically, Figure 5 The process shown includes the following steps:
[0102] S501: Obtain a first antibody sequence, where the first antibody sequence includes a paired heavy chain and a paired light chain.
[0103] As an example, the first antibody sequence may include a pair of amino acids in a heavy chain and a light chain. As another example, the first antibody sequence may include multiple pairs of amino acids in a heavy chain and a light chain, such as an antibody sequence for a Y-shaped antibody including two pairs of heavy chains and light chains.
[0104] In some embodiments, the first antibody sequence can be input as an input antibody sequence into a pre-trained structure detection model 01 such as AbFold.
[0105] S502: Obtain a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and a homologous sequence, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate a dependency relationship between amino acids in the first antibody sequence.
[0106] In some embodiments, the first antibody sequence is input into the homology feature extraction module 11 in the pre-trained structure detection model 01, so as to extract the first feature and the second feature of the first antibody sequence through the homology feature extraction module 11. For example, the first feature is a tensor with a shape of N×384 (such as S1), and the second feature is a tensor with a shape of N×N×128 (such as tensor z).
[0107] In some embodiments, the first antibody sequence is input into the language model feature extraction module 12 in the pre-trained structure detection model 01 so as to extract the third feature of the first antibody sequence through the language model feature extraction module 12 .
[0108] S503: Fusing the first feature and the third feature to obtain a fusion feature of the first antibody sequence.
[0109] In some embodiments, the present application can convert the first feature and the third feature of the first antibody sequence into features of the same dimension, and then fuse the two features to obtain a fused feature.
[0110] In some embodiments, the present application may input the first feature and the third feature into the feature fusion module 13 in the pre-trained structure detection model 01, so as to obtain a fused feature by fusing the first feature and the third feature through the feature fusion module 13. For example, the fused feature is a tensor (such as tensor S) with a dimension of N×384.
[0111] S504: Detect the three-dimensional structure of the first antibody sequence according to the second feature and the fusion feature.
[0112] In some embodiments, the present application may input the fusion feature of the first antibody sequence and the second feature into the structure detection module 14 in the pre-trained structure detection model 01, so as to fuse the first feature and the third feature through the structure detection module 13 to obtain the fusion feature.
[0113] It can be understood that the first feature and the second feature are homologous sequence features related to the antibody structure, and the second feature is a language model feature related to the antibody structure. Therefore, the fusion feature of the above-mentioned first antibody sequence can combine the homologous sequence features and the language model features of the first antibody sequence. Thus, the antibody sequence detection method provided by the present application can improve the structural detection accuracy of the CDR H3 loop by fusing the homologous sequence features and the language model features of the first antibody sequence to be detected, combining the effects of the two features on antibody structure detection, thereby improving the overall structural detection accuracy of the antibody sequence.
[0114] In some embodiments, Figure 6 As shown, it is a schematic diagram of a flow chart of an antibody sequence detection method provided by the present application. The execution subject of the method can be an electronic device or a device in the electronic device. As an example, Figure 6 The process shown can be performed by the various modules in the pre-trained structure detection model 01 of this application, such as AbFold. Figure 6 and Figure 5 The same steps will not be repeated, the main difference is that Figure 6 The processing flow of each module in the structure detection model 01 is shown.
[0115] Specifically, Figure 6 The process shown includes the following steps:
[0116] S601: The structure detection model 01 inputs a first antibody sequence, where the first antibody sequence includes a pair of heavy chains and light chains.
[0117] S602: The homologous sequence feature extraction module 11 in the structure detection model 01 searches for homologous sequences of the first antibody sequence, and obtains a multiple sequence comparison representation between the first antibody sequence and the homologous sequence.
[0118] In some embodiments, the homologous sequence feature extraction module 11 uses MMseqs2 to search a protein database such as Uniref90 to obtain homologous sequences of the first antibody sequence, and the search time is relatively short.
[0119] In some embodiments, the homologous sequence feature extraction module 11 is based on the MSA representation between the first antibody sequence and the homologous sequence, and the MSA representation is used to generate information indicating the evolutionary constraint between the first antibody sequence and the homologous sequence.
[0120] S603: The homologous sequence feature extraction module 11 in the structure detection model 01 connects the heavy chain and the light chain in the first antibody sequence into one sequence to obtain a first protein sequence, and acquires a pairing representation of the first protein sequence.
[0121] In some embodiments, the homologous sequence feature extraction module 11 performs pairing processing on different amino acids in the first protein sequence to obtain a pairing representation of the first protein sequence.
[0122] For example, the homologous sequence feature extraction module 11 may perform pairing processing on the amino acids in pairs in the first protein sequence to obtain a pair representation of the first protein sequence, namely, a pair representation.
[0123] S604: The homologous sequence feature extraction module 11 in the structure detection model 01 encodes the multiple sequence comparison representation and the pairing representation to obtain a first feature and a second feature respectively.
[0124] In some embodiments, the homologous sequence feature extraction module 11 can encode the multiple sequence comparison representation and pairing representation of the first protein sequence through transfer learning using the Evoformer module in AlphaFold2-Multimer to obtain the first feature of the first antibody sequence such as tensor S1 and the second feature such as tensor S2.
[0125] S605: The amino acid position encoding module 16 in the structure detection model 01 encodes the position of each amino acid in the first protein sequence to obtain the position encoding of each amino acid in the first antibody sequence.
[0126] Among them, the position codes of amino acids at different positions in the first antibody sequence are different.
[0127] For example, the amino acid position encoding module 16 may sequentially position-code the first N / 2 amino acids belonging to the heavy chain in the first protein sequence as 0 to N / 2, and sequentially position-code the last N / 2 amino acids belonging to the light chain in the first protein sequence as 256 to N / 2+256. Thus, the position encoding of each amino acid can distinguish the position of the amino acid and whether the amino acid belongs to the heavy chain or the light chain.
[0128] S606: The language model feature extraction module 12 in the structure detection model 01 connects the heavy chain and the light chain in the first antibody sequence into a sequence through the gaps to obtain a second protein sequence, and obtains the first language model encoding and the first attention matrix of the second protein sequence.
[0129] In some embodiments, the language model feature extraction module 12 inputs the second protein sequence into a pre-trained protein language model and outputs a first language model encoding and a first attention matrix.
[0130] In some embodiments, the language model feature extraction module 12 can use the protein language model AntiBERTy to identify the second protein sequence through transfer learning, and determine the first language model encoding and the first attention matrix of the second protein sequence.
[0131] S607: The language model feature extraction module 12 in the structure detection model 01 updates the first language model encoding and the first attention matrix based on the attention mechanism to obtain a third feature.
[0132] In some embodiments, the language model feature extraction module 12 can use each amino acid of the second protein sequence as a node and construct edges with the first attention matrix to obtain graph structure data. Then, the graph structure data is iteratively updated based on the triangle update mechanism, such as iterating 3 times, to obtain the updated first language model code and the first attention matrix for update processing, and the updated first language model code is determined as the third feature.
[0133] S608: The root mean square error calculation module 14 in the structure detection model 01 obtains the root mean square error information of each amino acid main chain atom in the first antibody sequence.
[0134] In some embodiments, the root mean square error calculation module 14 may calculate the root mean square error of each amino acid backbone atom (C α , C β , N, O) (pRMSD), that is, calculating the main chain atom C of each amino acid α The root mean square error of the main chain atoms C β The root mean square error of the main chain atom N and the root mean square error of the main chain atom O. Then, the root mean square error information of the main chain atoms of each amino acid in the first antibody sequence is obtained as N×4 data.
[0135] S609: The feature fusion module 14 in the structure detection model 01 converts the root mean square error information into a scaling factor, performs dimension conversion on the third feature according to the scaling factor to obtain a fourth feature, and obtains a fused feature based on the fourth feature and the first feature.
[0136] The dimension of the scaling factor is one-dimensional, and the dimension of the fourth feature is the same as the dimension of the first feature.
[0137] In some embodiments, the feature fusion module 14 can input the N×4-dimensional root mean square error information pRMSD of the first antibody sequence into the first linear layer of the confidence conversion submodule 141 in the feature fusion module 14, namely FC(4,32), so that the last activation layer (Sigmoid layer) of the confidence conversion submodule 141 outputs the scaling factor. At this time, the N×4-dimensional root mean square error information pRMSD passes through the FC(4,32) layer, the Sigmoid layer, the FC(32,32) layer, the Sigmoid layer, the FC(32,1) layer, and the Sigmoid layer in the confidence conversion submodule 141, so that the confidence conversion submodule 141 outputs the scaling factor.
[0138] In some embodiments, the feature fusion module 14 inputs the scaling factor and the third feature into the aggregation layer in the feature conversion submodule 142, and inputs the first feature into the summation layer such as the Add layer in the feature conversion submodule 142, so that the last activation layer (Sigmoid) in the feature conversion submodule 142 outputs the fusion feature. That is, the scaling factor and the third feature of the first antibody sequence, that is, the tensor S2, are input into the Multiplate layer in the feature conversion submodule 142, and are converted into the fourth feature through the FC (64, 384) layer and the Relu layer. The dimension of the fourth feature is the same as the dimension of the first feature, which is N×384. Then, the tensor S1 corresponding to the first feature of the first antibody sequence is input into the Add layer in the feature conversion submodule 142, so that the tensor corresponding to the fourth feature is summed with the tensor S1 corresponding to the first feature in the Add layer. Then, the summed tensor is converted through the FC (64, 384) layer and the Sigmoid layer and the tensor S corresponding to the fusion feature is output.
[0139] In some embodiments, the feature fusion module 14 performs a summation process on the fourth feature and the first feature to obtain a fused feature.
[0140] Specifically, the feature conversion submodule 142 in the feature fusion module 14 inputs the tensor S1 corresponding to the first feature of the first antibody sequence into the Add layer in the feature conversion submodule 142, so that the tensor corresponding to the fourth feature is summed with the tensor S1 corresponding to the first feature in the Add layer. Then, the summed tensor is converted through the FC (64, 384) layer and the Sigmoid layer and outputs the tensor S corresponding to the fused feature.
[0141] S610: The feature updating module 15 in the structure detection model 01 updates the fused feature and the second feature based on the attention mechanism.
[0142] In some embodiments, the feature updating module 15 may update the fused feature and the second feature based on an attention mechanism such as a triangular attention mechanism.
[0143] S611: The structure detection module 17 in the structure detection model 01 detects the mean and variance of the translation parameters of the amino acids in the first antibody sequence based on the updated second feature and the updated fusion feature, and detects the mean and variance of the rotation parameters of the amino acids in the first antibody sequence.
[0144] In some embodiments, the structure detection module 17 can identify amino acids at different positions based on the position codes of each amino acid in the first antibody sequence during the process of detecting the mean and variance of the translation parameters of each amino acid. In this case, the structure detection module 17 can adopt a Transformer architecture.
[0145] S612: The structure detection module 17 in the structure detection model 01 uses a re-parameterization method to determine the target value of the translation parameter and the target value of the rotation parameter of each amino acid in the first antibody sequence according to the mean and variance of the translation parameter and the mean and variance of the rotation parameter.
[0146] As an example, the structure detection module 17 may use the above formula (1) to determine the target value of the translation parameter and the target value of the rotation parameter of each amino acid in the first antibody sequence to determine the three-dimensional structure of the first antibody sequence.
[0147] The translation parameter represents the position of the amino acid, the rotation parameter represents the angle of the amino acid, and the three-dimensional structure includes the position and angle of the amino acid. At this time, the position and angle of each amino acid in the first antibody sequence are used to characterize the three-dimensional structure of the first antibody sequence.
[0148] Thus, in the antibody sequence detection method provided in the embodiment of the present application, the pre-trained structural antibody detection model can use the search tool MMseqs2 to quickly search the protein sequence database, which is conducive to quickly obtaining the homologous sequence of the first antibody sequence. In addition, the structural antibody detection model can use transfer learning to extract the multidimensional features of the homologous sequence from AlphaFold2, that is, the first feature, and the interaction information between the amino acids in the sequence, that is, the second feature. At the same time, the structural detection model can also use the pre-trained protein language model AntiBERTy through transfer learning to extract the language model features of the first antibody sequence, and further use the feature encoding module in IgFold such as the Evoformer module through transfer learning to quickly generate multidimensional feature information, that is, the third feature. As a result, the pre-trained structural antibody detection model of the present application can fuse the homologous sequence features of the first antibody sequence with the language model features to perform antibody structure detection, thereby improving the overall structural detection accuracy of the antibody sequence.
[0149] In some embodiments, reference Figure 7 As shown, it is a flow chart of performing structure detection on an antibody sequence provided in the embodiment of the present application. Specifically, Figure 7The structure detection model 01 shown, i.e., the antibody structure detection process of AbFold, includes: in terms of input, AbFold receives the input antibody sequence (such as the first antibody sequence) composed of heavy chains and light chains, uses MMseqs2 to search the protein sequence database Uniref90, obtains the homologous sequence of the input antibody sequence, and generates a multiple sequence alignment (MSA) representation. Next, AbFold uses transfer learning to extract the multidimensional features of the homologous sequence (i.e., the first feature such as tensor S1) and the interaction information between the two amino acids in the sequence (i.e., the second feature such as tensor z) from AlphaFold2. At the same time, the input antibody sequence will be input into the pre-trained protein language model AntiBERTy to extract the language model features of the sequence, and further input into the IgFold feature encoding module to generate multidimensional feature information (i.e., the third feature such as tensor S2). Then, AbFold's deep neural network fuses the multidimensional feature information represented by tensor S1 and tensor s2 respectively to obtain a fused feature (such as tensor S). Through the triangular self-attention mechanism, the tensor S corresponding to the fused feature and the tensor z corresponding to the second feature are mutually updated. Finally, the updated tensors S and z are input into the structure prediction module, and the invariant point attention mechanism and variational autoencoder are used to dynamically predict the three-dimensional structure of the antibody. In this way, AbFold starts with the input of the amino acid sequence of the antibody sequence, not only extracting information from the homologous sequence, but also extracting sequence features with the help of the pre-trained antibody language model. By fusing homologous sequence features with language model features through a deep neural network, AbFold dynamically models the three-dimensional structure of the antibody, which is conducive to improving the accuracy of antibody three-dimensional structure detection.
[0150] Next, combine Figures 8 to 10 , the model training stage of the deep learning-based antibody structure detection model in this application is described in detail.
[0151] Reference Figure 8 As shown, it is a schematic diagram of the training process of an antibody structure detection model provided for the implementation of the present application. The execution subject of the process can be an electronic device or a device in the electronic device.
[0152] Specifically, Figure 8 The training process shown includes the following steps:
[0153] S801: Acquire a training data set, where the training data set includes at least one training data group, and each training data group includes an unlabeled antibody sequence and a labeled antibody sequence.
[0154] In some embodiments, the labeled antibody sequences and unlabeled antibody sequences in a training data set can be randomly selected from the training data set or pre-divided, and this application does not make any specific limitation on this.
[0155] As an example, to train and evaluate the performance of the structure detection model 01, namely the AbFold model, the present application screened 1872 pairs of high-quality antibody structures from the structural antibody database SAbDab, with a maximum sequence similarity of 99% and a minimum resolution of 4.0 angstroms. Among them, in order to make an effective comparison with the existing antibody structure prediction model (such as IgFold), this application divides the data set corresponding to the SAbDab database into a training set and a test set according to the storage date of the antibody structure, ensuring that the data in the test set is not used in the training process of any model. For example, this application uses the data of antibody sequences stored in the SAbDab database before July 1, 2021 for model training, and uses the data of antibody sequences stored after this date for model testing. Then, the training set corresponding to the SAbDab database contains 1719 labeled antibody sequences, and the test set contains 153 labeled antibody sequences.
[0156] Among them, Egypt It is a unit of length, often used to express the distance between atoms or molecules. For example, the value of the root mean square error (RMSD) is This means that the difference between the predicted and actual values is about 1.37 angstroms on average.
[0157] As an example, the present application randomly selected 20,000 pairs of heavy chain and light chain sequences from the antibody sequence database OAS as input data for unsupervised training of the model. At this time, during the model training process, the structured supervised data in the training data set, such as the labeled antibody sequences in the database SAbDab, and the pure sequence unsupervised data, such as the unlabeled antibody sequences in the database OAS, each accounted for half.
[0158] In some embodiments, the training data set constructed in the present application contains 1872 antibody sequences with native structures and 20,000 unstructured antibody sequences. According to the storage date of the antibody sequence, 1719 antibody sequences stored before July 1, 2021 are used as training sets, and 153 antibody sequences are stored as test sets thereafter. All 20,000 unstructured antibody sequences are used for self-supervised learning training.
[0159] It is understandable that the number of antibody structures currently measured by experimental methods such as nuclear magnetic resonance, X-ray crystallography and cryo-electron microscopy is limited. According to statistics from the SabDab database, only several thousand antibody structures have been experimentally measured. This limited labeled data has a significant negative impact on the training effect of the deep learning model. For this reason, the application adopts a teacher-student self-supervised learning method, by introducing a large number of unlabeled antibody sequences, training in a self-supervised manner, and using labeled antibody sequences together for model training, thereby improving the predictive ability of the model.
[0160] S802: Based on the training data set, the structure detection model to be trained is trained by using a teacher-student self-supervised learning method to obtain a trained structure detection model.
[0161] In some embodiments, in the model training stage based on the teacher-student self-supervised learning method, the structure detection model to be trained can be divided into two branches: a student model and a teacher model. As an example, the structure of the student model and the teacher model are consistent, so the student model and the teacher model are used to perform model training based on the teacher-student self-supervised learning method.
[0162] In some embodiments, during the training process, each mini-batch consists of half labeled antibody sequences and half unlabeled antibody sequences. The labeled antibody data uses the natural structure (i.e., the true label) as a supervisory signal, and the unlabeled training antibody sequences use the three-dimensional structure predicted by the teacher model as a supervisory signal.
[0163] In some implementations, the model training in this application is performed for 300 epochs, the minimum batch size is 8, and the optimizer uses AdamW to accelerate convergence. The basic learning rate is set to e -4 , it increases linearly to this value in the first 1000 iterations, then remains unchanged from 1000 to 100000 iterations. After 100000 iterations, the learning rate is decayed by 5% every 1000 iterations to improve training stability and performance. To avoid gradient explosion, the gradients of all parameters of the model are clipped by L2 norm with a clipping threshold of 0.1. In addition, in order to ensure that the sequence lengths within the mini-batch are equal during training and increase the diversity of the data, the length of a single chain (such as a heavy chain or a light chain) is fixed to a set value such as 96. For antibody sequences whose single chain length exceeds the set value of 96, a continuous subsequence with a median length of the single chain of the preset value of 96 is randomly selected; for antibody sequences whose median length of a single chain is less than the preset value of 96, the end of the single chain in the antibody sequence is padded with 0.
[0164] like Fig. 9 As shown, it is a schematic diagram of a network architecture for model training based on teacher-student self-supervised learning provided in an embodiment of the present application. Fig. 9 The architecture diagram shows that the student model 91 and the teacher model 92 are included. The student model 91 and the teacher model 92 have the same structure, with the input being the antibody sequence and the output being the three-dimensional structure of the antibody. For example, for the same input antibody sequence, the three-dimensional structure output by the student model 91 is denoted as T1, and the three-dimensional structure output by the teacher model 92 is denoted as T2. Fig. 9As shown, during the training process, the student model 91 and the teacher model 92 respectively predict the structure of the unlabeled antibody sequence, and use the prediction results of the teacher model 92 as a supervisory signal to guide the learning of the student model 91. Specifically, firstly, a pseudo label is generated according to the prediction results of the teacher model, and then the corresponding loss such as Loss (T1, T2) is calculated, that is, the loss between the three-dimensional structure T1 and the three-dimensional structure T2, and the parameters of the student model 91 are updated by gradient back propagation. The parameters of the teacher model 92 are not updated by directly back propagating the gradient, but are dynamically updated by the exponential moving average (EMA) of the parameters of the student model 91. At this time, the teacher model 92 can stop the stochastic gradient descent (sg) method to update the parameters. Among them, the full name of the exponential moving average is the exponential moving average, which is a trend-type indicator, specifically an exponentially decreasing weighted moving average. The above-mentioned synchronous update strategy dynamically constructs the teacher model 92 during the training process, avoiding dependence on a fixed teacher model, so that the teacher model 92 can continuously optimize itself. In this way, a small amount of annotated structural information can be effectively propagated to a large number of unlabeled antibody sequences, thereby improving the model's ability to predict the structure of antibody sequences.
[0165] In this way, the teacher-student self-supervised learning strategy in this application significantly improves the quality of features learned by the model and can better deal with the problem of insufficient labeled data in the antibody structure prediction task.
[0166] Reference Fig.10 As shown, it is a schematic diagram of the training process of the structure detection model based on the training data set provided in the embodiment of the present application, and the execution subject of the process can be an electronic device or a device in the electronic device. For example, the training data set includes a first training antibody sequence without a label and a second training antibody sequence with a label.
[0167] Specifically, Fig.10 The training process shown includes the following steps:
[0168] S1001: Obtain a teacher model and a student model corresponding to the structure detection model to be trained.
[0169] S1002: Inputting the unlabeled first training antibody sequence in the training data set into the student model to obtain a first detection result, and inputting the first training antibody sequence into the teacher model to obtain a second detection result.
[0170] S1003: Input the labeled second training antibody sequence in the training data set into the student model to obtain a third detection result.
[0171] In some embodiments, the second training antibody sequence can be input into the teacher model to obtain a fourth detection result.
[0172] As an example, the input and output of the teacher model and the student model can be described by the following formulas (2) and (3).
[0173]
[0174] Among them, x1 and x2 represent the unlabeled antibody sequences and labeled antibody sequences in the minimum batch data of the training process, respectively, student and teacher represent the student model and teacher model in self-supervised learning, respectively, and T represents the antibody structure predicted by the model (i.e., the three-dimensional structure of the antibody sequence).
[0175] As an example, Represents the three-dimensional structure obtained by the student model student performing structural detection on the unlabeled antibody sequence x1. For example, when x1 is the first training antibody sequence This is the first test result.
[0176] Represents the three-dimensional structure obtained by the teacher model teacher performing structural detection on the unlabeled antibody sequence x1. For example, when x1 is the first training antibody sequence This is the second test result.
[0177] Represents the three-dimensional structure obtained by the student model student performing structural detection on the labeled antibody sequence x2. For example, when x2 is the second training antibody sequence This is the third test result.
[0178] Represents the three-dimensional structure obtained by the teacher model teacher performing structural detection on the labeled antibody sequence x2. For example, when x2 is the second training antibody sequence This is the fourth test result.
[0179] S1004: Calculate a first loss between the student model and the teacher model based on the first detection result and the second detection result.
[0180] S1005: Calculate a second loss between the student model and the teacher model based on the third detection result and the true label of the second training antibody sequence.
[0181] In some embodiments, the loss function between the student model and the teacher model can be based on the loss function in AlphaFold2. and Functions and IgFold Function confirmed.
[0182] As an example, the loss function between the student model and the teacher model can be determined by formula (4).
[0183]
[0184] Among them, the loss function Including AlphaFold2 and The former measures the similarity between the predicted structure and the actual structure of the antibody sequence, and the latter is used to learn the geometric constraints between residues in the antibody sequence to avoid conflicts between atoms. The loss function comes from IgFold and is used to calculate the L1 norm of the distance between the i-th amino acid and the i+1 and i+2 amino acids in the antibody sequence, where i is a positive integer.
[0185] It can be understood that both the first loss and the second loss can be calculated based on the above formula (4).
[0186] For example, the first loss between the student model and the teacher model It can be calculated using formula (4). That is, the three-dimensional structure predicted by the model student for the unlabeled antibody sequence (such as the first test result mentioned above) is used as T in formula (4). pred , That is, the pseudo label predicted by the teacher model for the antibody sequence (such as the second detection structure mentioned above) is used as T in formula (4) label .
[0187] For example, the second loss between the student model and the teacher model It can be calculated using formula (4). That is, the three-dimensional structure (such as the third test result) predicted by the student model for the antibody sequence corresponding to the labeled data is used as T in formula (4) pred , T gt That is, the true label of the antibody sequence (i.e., the true label corresponding to the second training antibody sequence) is used as T in formula (4). label .
[0188] S1006: Calculate the sum of the first loss and the second loss to obtain a third loss of the structure detection model.
[0189] As an example, the loss function of the structure detection model can be determined by formula (5).
[0190]
[0191] Furthermore, based on the loss function in formula (5), the overall loss of the structure detection model during the training process can be calculated, and the parameters of the student model and the teacher model can be updated according to the loss.
[0192] S1007: Based on the third loss, update the parameters of the student model using the back-propagation gradient method.
[0193] S1008: Update the parameters of the teacher model by performing exponential sliding average processing on the updated parameters of the student model.
[0194] In some embodiments, updating the parameters of the teacher model by the exponential sliding average of the student model parameters can be performed by the following formula (6).
[0195] Θ teacher =(1-α)θ teacher +αΘ student (6)
[0196] Among them, Θ teacher and θ student They represent the parameters of the teacher model and the student model respectively, and α=0.01 is the weighting coefficient.
[0197] Similarly, the present application can be used for each training data group in the training data set according to Fig.10 The process shown is used to update the parameters of the structure detection model.
[0198] In some embodiments, the present application may determine the trained student model as the trained structure detection model.
[0199] In this way, the present application can calculate the first loss corresponding to the unlabeled antibody sequence and the second loss corresponding to the labeled antibody sequence through the teacher-student self-supervised model training method, and then obtain the third loss of the entire model based on the first loss and the second loss. Then, the third loss simultaneously considers the impact of the unlabeled antibody sequence and the labeled antibody sequence on the model loss, which is conducive to improving the structure detection accuracy of the structure detection model after the model training based on the third loss.
[0200] Next, an experiment is conducted on the structure detection process based on the pre-trained structure detection model 01, namely AbFold, in the implementation of this application, and the detection performance of the structure detection model 01 is analyzed based on the experimental results.
[0201] In some embodiments, in order to verify the effectiveness of the model, this application collects 153 double-chain antibody sequences with high-resolution structures as a test set and compares them with existing advanced antibody structure prediction models. These 153 antibody sequences are stored in public databases after July 1, 2021 to ensure that these antibody sequences are not included in the training set of any of the above-mentioned models to be compared, thereby ensuring the fairness of the test. It can be understood that the most challenging part of antibody structure prediction is modeling CDR loops, and this application can use the Chothia numbering scheme to determine the various CDR loops in the test set. Among them, the Chothia numbering scheme is used to identify amino acids at specific positions of antibodies. The scheme is based on the crystal structure of the antibody variable region, defines the loop structure that forms the CDR (complementarity determining region), and corrects the position numbering of the CDR-L1 loop and CDR-H1 loop insertion points to better fit their topological positions.
[0202] In some embodiments, the present application compares the AbFold-based structure detection method in the present application with other antibody structure prediction methods such as AlphaFold2-Multimer, IgFold, ImmuneBuilder, DeepAb and EquiFold on 153 test antibodies. For example, Table 2 shows the prediction accuracy of these models on the test set.
[0203] Table 1:
[0204]
[0205]
[0206] It is understood that this application uses the average root mean square error (RMSD) to evaluate the quality of antibody structure prediction. Figure 2 The experimental results shown in the figure show that AbFold has the lowest error in the structure prediction of CDR loops, and its predicted structure is closer to the natural structure. Among the six CDR loops, AbFold has achieved the lowest average RMSD difference (i.e. ), and the average RMSD of the prediction results of the H1 loop is only slightly better than the best result (i.e. )Difference In the most difficult CDR H3 loop to model, the average RMSD of AbFold is With AlphaFold2-Multimer IgFold ImmuneBuilder DeepAb and EquiFold Compared with AbFold, the average RMSD of AbFold was reduced by 13%, 24%, 5%, 30% and 12% respectively. In the relatively more stable skeleton regions (FRH, FRL), all methods can predict the corresponding structures with high accuracy (FRH: FRL: ); however, in the CDR region with strong structural variability, the RMSD of AbFold is significantly lower than other methods including AlphaFold2-Multimer, indicating that AbFold has a stronger modeling ability for dynamic structures. It is understandable that since deterministic structure prediction methods such as AlphaFold2-Multimer are difficult to model different conformations of the same sequence, the model training tends to learn the mean structure of multiple conformations in the training set, while AbFold based on variational autoencoder introduces noise and variance to make the predicted structure have a certain randomness, thereby avoiding the model forcibly fitting the mean of multiple conformations.
[0207] Reference FIG. 11A to FIG. 11C As shown, it is a schematic diagram of the test results of various models on the test set in the embodiments of the present application. These test results comprehensively compare the performance of six antibody structure prediction methods, namely AbFold, AlphaFold2-Multimer, IgFold, ImmuneBuilder, DeepAb and EquiFold, on the test set.
[0208] like Fig.11A As shown in FIG. 1 , it is a box plot of the prediction accuracy of various antibody structure prediction methods provided in the examples of this application. Fig.11A As shown, compared with other methods, the prediction results of AbFold showed lower values in the lower quartile, median, upper quartile and interquartile range, indicating that the error distribution of AbFold's prediction results was more concentrated, the differences between antibodies were smaller, and the model prediction results were more stable.
[0209] It is understood that in order to gain a deeper understanding of how AbFold optimizes predicted structures and achieves higher accuracy in CDR H3, this application analyzes two examples from the test set, whose protein structure database (protein data bank, PDB) identifiers (IDs) are 7N0A and 7PHW, respectively. Fig. 11B The 7N0A native structure, AbFold predicted structure and IgFold predicted structure are shown in the figure. Fig. 11B The native structure of 7N0A is the expected structure represented by the black part, the AbFold predicted structure is the predicted structure represented by the gray part, and the IgFold predicted structure is the predicted structure represented by the white part. Fig. 11BThe AbFold predicted structure shown is closer to the 7N0A native structure.
[0210] Fig. 11C The figure shows the visualization diagram of the native structure of antibody PHW, the predicted structure of AbFold and the predicted structure of IgFold. Fig. 11C The 7PHW native structure is the expected structure represented by the black part, the AbFold predicted structure is the predicted structure represented by the gray part, and the IgFold predicted structure is the predicted structure represented by the white part. Fig. 11B The AbFold predicted structure shown is closer to the 7PHW native structure.
[0211] Specifically, the application compares the structures predicted by IgFold and AbFold with the native structures. Fig. 11B The CDR H3 loop of this antibody consists of seven residues.
[0212] Average of IgFold This indicates that IgFold failed to accurately predict the conformation of the CDR-H3 loop, showing a low confidence in this region. This indicates that AbFold predicts a relatively more accurate CDR-H3 conformation. AbFold optimizes the orientation of the CDR H3 structure to make it closer to the native structure, thereby significantly reducing the RMSD. Fig. 11C The length of the CDR-H3 loop is 18, more than twice the length of the previous example 7N0A. The average RMSD of the conformations predicted by AbFold The average RMSD of the conformations predicted by IgFold is significantly better than that of This indicates that AbFold can not only accurately predict antibodies with short CDR-H3 loops, but also predict the structures of antibodies with longer CDR-H3 loops.
[0213] In some embodiments, the present application can analyze the influence of the fusion feature (also referred to as fusion information) of the antibody sequence on the structural prediction accuracy of the CDR-H3 structure. It can be understood that there are significant differences between the CDR-H3 structures predicted by AlphaFold2-Multimer based on homologous sequences and IgFold based on protein language models. In order to explore the influence of information fusion on the prediction of antibody CDR-H3 structure, the present application compares the antibody CDR-H3 prediction results of AbFold with AlphaFold2-Multimer and IgFold on the test set separately, and uses the root mean square error (RMSD) as the evaluation indicator.
[0214] Reference Fig. 12AA scatter plot comparing AbFold and AlphaFold2-Multimer is shown, see Fig. 12B A scatter plot comparing AbFold to IgFold is shown.
[0215] like Fig. 12A and Fig. 12B As shown in the figure, AbFold, which combines homologous sequence features and protein language model features, is significantly better than AlphaFold2-Multimer, which only uses homologous sequence features, and IgFold, which only uses protein language model features. Among the 153 tested antibody sequences, the number of antibody sequences whose structures predicted by AbFold are better than those of AlphaFold2-Multimer and IgFold are 88 and 114, respectively, accounting for 57.5% and 74.5%, respectively.
[0216] Reference Fig. 12C , showing a schematic diagram of the comparison between the CDR-H3 structures predicted by the AbFold, AlphaFold2 and IgFold models and the natural structure. Fig. 12C As shown in the figure, from left to right, they are visualizations of the CDR-H3 structures predicted by AlphaFold2 (such as AlphaFold2-Multimer), IgFold, and AbFold, and from top to bottom in rows 1 to 3, they are antibodies 7bh8, 7jwg, and 7aj6. The results show that the prediction results of AbFold are closer to the natural structure. This shows that AbFold can capture more precise structural information from homologous sequence features and language model features through the feature fusion module, thereby more accurately predicting the structure of the antibody CDR-H3 loop.
[0217] In some embodiments, the experimental results of the ablation experiment of the network structure of the AbFold model in the present application are analyzed. Table 2 shows the results of the ablation experiment of the network structure of the AbFold model.
[0218] Table 2:
[0219]
[0220] Specifically, Table 2 shows the values of RMSD (i.e., H3 RMSD) and global distance test total score (GDT-TS) of AbFold and the first ablation model (AbFold-w / o-pRMSD), the second ablation model (AbFold-w / o-VAE), and the third ablation model (AbFold-Attn) in CDR-H3.
[0221] AbFold-w / o-pRMSD means that the AbFold model is based on the AbFold model, but the pRMSD module or technology is removed.
[0222] AbFold-w / o-VAE represents a variant of the AbFold model that does not use variational autoencoder (VAE) as a model component or data preprocessing step when performing protein structure prediction.
[0223] AbFold-Attn means that the attention mechanism is introduced based on the AbFold model, which enables the model to capture the key information in the amino acid sequence more accurately.
[0224] In addition, GDT-TS represents the proportion of atomic positions matching between the predicted protein structure and the true structure at different distance thresholds. GDT-TS integrates the degree of matching at different accuracy levels and is consistent for proteins of different sizes.
[0225] As shown in Table 2, the overall structure and CDR-H3 region prediction accuracy of AbFold are significantly higher than those of the ablation models. Compared with AbFold-w / o-pRMSD, AbFold-Attn and AbFold-w / o-VAE, AbFold's GDT_TS increased by 6.8%, 3.3% and 5.6%, respectively, and the RMSD of CDR-H3 prediction decreased by 23.8%, 15.8% and 19.9%, respectively. Compared with AbFold-w / o-pRMSD, AbFold increased GDT_TS by 6.8% and decreased H3 RMSD by 23.8%, which shows that pRMSD plays a guiding role in the fusion of homologous sequence features and language model features, thereby obtaining more accurate structural information. Compared with AbFold-w / o-VAE, AbFold improves by 5.6% on GDT_TS and reduces by 19.9% on H3 RMSD, which shows that the variational autoencoder not only plays an important role in modeling the highly variable CDR H3 loop, but also has a positive impact on the relatively stable skeleton region. Compared with AbFold-Attn, AbFold improves by 3.3% on GDT_TS and reduces by 15.8% on H3 RMSD. The prediction accuracy of AbFold-Attn based on the attention mechanism is slightly lower than that of AbFold based on the fully connected layer. This may be because there is less training data and it is difficult to fully train the attention network, which is powerful but requires a lot of data.
[0226] In summary, the CDR H3 loop of an antibody plays a vital role in the binding process between an antibody and an antigen. With the continuous advancement of protein structure prediction methods, the prediction accuracy of antibody backbone structures has approached the experimental level. However, the structural prediction of the CDR-H3 loop is still a difficulty and hotspot in antibody research. This application proposes an antibody structure detection model AbFold based on transfer learning. First, MMseqs2 can be used to search the Uniref90 sequence database to obtain homologous sequences and multiple sequence alignments of antibody sequences, and the language model features are extracted through the antibody language model AntiBERTy. Then, transfer learning is used to extract homologous sequence features and language model features from AlphaFold-Multimer and IgFold respectively. Then, these two types of features are fused through the feature fusion module of the fully connected layer to obtain the final fusion features, which are input into the neural network based on the variational autoencoder (i.e., the structure detection module) for dynamic modeling of the antibody structure. The experimental results on 153 test antibodies show that AbFold is more accurate than other antibody structure prediction methods in predicting the CDR-H3 loop structure. By combining homologous sequence features and language model features, AbFold can capture different structural information more comprehensively, making antibody structure prediction more accurate. In addition, based on the antigen sequence and structural information combined with the antibody sequence, an antibody structure prediction method under given antigen conditions can be generated, further improving the accuracy and practicality of antibody structure prediction.
[0227] Next, combine Fig.13 The structure of the electronic device for executing the antibody sequence detection method of the present application is introduced in detail.
[0228] like Fig.13 As shown, the electronic device 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, a button 101 and a display screen 102, etc.
[0229] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0230] The processor 110 may include one or more processing units, for example, a processing module or processing circuit including a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor DSP, a microprocessor (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA). Different processing units may be independent devices or integrated into one or more processors. A storage unit may be provided in the processor 110 for storing instructions and data. In some embodiments, the storage unit in the processor 110 is a cache memory 180. For example, the processor 110 executes Figure 5 , Figure 6 , Figure 8 and Fig.10 The method in .
[0231] The power module 140 may include a power source, a power management component, etc. The power source may be a battery. The power management component is used to manage the charging of the power source and the power supply of the power source to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module is used to receive charging input from the charger; the power management module is used to connect the power source, the charging management module and the processor 110. The power management module receives input from the power source and / or the charging management module, and supplies power to the processor 110, the display screen 102, the camera 170, and the wireless communication module 120.
[0232] The mobile communication module 130 may include, but is not limited to, an antenna, a power amplifier, a filter, a low noise amplifier (LNA), etc. The mobile communication module 130 may provide a solution for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 10. The mobile communication module 130 may receive electromagnetic waves by an antenna, and perform filtering, amplification, and other processing on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 130 may also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna. In some embodiments, at least some of the functional modules of the mobile communication module 130 may be arranged in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 130 may be arranged in the same device as at least some of the modules of the processor 110.
[0233] The wireless communication module 120 may include an antenna, and transmit and receive electromagnetic waves via the antenna. The wireless communication module 120 may provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc., which may be applied to the electronic device 10. The electronic device 10 may communicate with the network and other devices through wireless communication technology.
[0234] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the electronic device 10 may also be located in the same module.
[0235] The display screen 102 is used to display human-computer interaction interfaces, images, videos, etc.
[0236] The sensor module 190 may include a proximity light sensor, a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.
[0237] The audio module 150 is used to convert digital audio information into analog audio signal output, or convert analog audio input into digital audio signal. The audio module 150 can also be used to encode and decode audio signals. In some embodiments, the audio module 150 can be arranged in the processor 110, or some functional modules of the audio module 150 can be arranged in the processor 110. In some embodiments, the audio module 150 can include a speaker, an earpiece, a microphone, and an earphone interface.
[0238] The camera 170 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the image signal processing (ISP) to convert it into a digital image signal. The electronic device 10 can implement the shooting function through the ISP, camera 170, video codec, graphic processing unit (GPU), display screen 102 and application processor. For example, in the present application, the mobile phone can collect images through the ISP, camera 170, etc., and perform the present application on these images in the post-processing stage. Figure 5-Figure 8 Related image processing flow.
[0239] The interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device 10. The external memory card communicates with the processor 110 via the external memory interface to implement a data storage function. The universal serial bus interface is used for the electronic device 10 to communicate with other electronic devices. The subscriber identification module card interface is used to communicate with a SIM card installed in the electronic device 1010, for example, to read a phone number stored in the SIM card, or to write a phone number into the SIM card.
[0240] In some embodiments, the electronic device 10 further includes a button 101, a motor, an indicator, etc. Among them, the button 101 may include a volume button, an on / off button, etc. The motor is used to make the electronic device 10 vibrate, for example, vibrate when the user's electronic device 10 is called, so as to prompt the user to answer the call of the electronic device 10. The indicator may include a laser indicator, a radio frequency indicator, an LED indicator, etc.
[0241] In some embodiments, the present application provides a readable medium having instructions stored thereon, which, when executed on an electronic device, causes the electronic device to execute the above-mentioned image processing method.
[0242] In some embodiments, the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the image processing method described above.
[0243] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.
[0244] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0245] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.
[0246] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a machine (e.g., computer) readable form.
[0247] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0248] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0249] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0250] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it will be apparent to those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application.
Claims
1. A method for detecting an antibody sequence, characterized in that: The method comprises: Obtaining a first antibody sequence, wherein the first antibody sequence includes a paired heavy chain and a light chain; Acquire a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and a homologous sequence, the second feature is used to indicate a relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate a dependency relationship between amino acids in the first antibody sequence; fusing the first feature and the third feature to obtain a fusion feature of the first antibody sequence; Based on the second feature and the fusion feature, the three-dimensional structure of the first antibody sequence is detected.
2. The method according to claim 1, characterized in that The fusing the first feature and the third feature to obtain the fusion feature of the first antibody sequence includes: Converting the root mean square error information of the first antibody sequence into a scaling factor, wherein the root mean square error information includes the root mean square error of each amino acid main chain atom in the first antibody sequence, and the dimension of the scaling factor is one-dimensional; Performing dimension conversion on the third feature according to the scaling factor to obtain a fourth feature, wherein the fourth feature has the same dimension as the first feature; The fourth feature and the first feature are summed to obtain the fused feature.
3. The method according to claim 1, characterized in that The first feature and the second feature are determined by: Searching for homologous sequences of the first antibody sequence; obtaining a multiple sequence alignment representation between the first antibody sequence and the homologous sequence; connecting the heavy chain and the light chain in the first antibody sequence into one sequence to obtain a first protein sequence; performing pairing processing on different amino acids in the first protein sequence to obtain a paired representation of the first protein sequence; The multiple sequence comparison representation and the pairing representation are encoded to obtain the first feature and the second feature, respectively.
4. The method according to claim 1, characterized in that: The third feature is determined by: connecting the heavy chain and the light chain in the first antibody sequence into one sequence through the gaps to obtain a second protein sequence; Inputting the second protein sequence into a pre-trained protein language model, and outputting a first language model encoding and a first attention matrix; The first language model encoding and the first attention matrix are updated based on the attention mechanism to obtain the third feature.
5. The method according to claim 1, characterized in that The detecting and obtaining the three-dimensional structure of the first antibody sequence according to the second feature and the fusion feature comprises: Performing an updating process on the fused feature and the second feature based on an attention mechanism; The three-dimensional structure of the first antibody sequence is detected based on the updated second feature and the updated fusion feature.
6. The method according to claim 5, characterized in that The detecting and obtaining the three-dimensional structure of the first antibody sequence according to the updated second feature and the updated fusion feature comprises: According to the updated second feature and the updated fusion feature, detecting the mean and variance of the translation parameters of the amino acids in the first antibody sequence, and detecting the mean and variance of the rotation parameters of the amino acids in the first antibody sequence; A reparameterization method is adopted to determine the target value of the translation parameter and the target value of the rotation parameter of each amino acid in the first antibody sequence according to the mean and variance of the translation parameter and the mean and variance of the rotation parameter, wherein the translation parameter represents the position of the amino acid, the rotation parameter represents the angle of the amino acid, and the three-dimensional structure includes the position and angle of the amino acid.
7. The method according to claim 6, characterized in that The positions of different amino acids in the first antibody sequence are encoded differently.
8. The method according to claim 2, characterized in that: The three-dimensional structure is determined based on a structure detection model, the structure detection model includes a feature fusion module, the feature fusion module includes a confidence conversion submodule and a feature transformation submodule, the confidence conversion submodule includes three groups of network layers connected in sequence, each group of network layers includes a linear layer and an activation layer connected in sequence, and the feature transformation submodule includes an aggregation layer, a linear layer, an activation layer, a summation layer, a linear layer and an activation layer connected in sequence, and the fusion feature is determined based on the feature fusion module.
9. The method according to claim 8, characterized in that The process of determining the fusion feature by the feature fusion module includes: Inputting the root mean square error information into the first linear layer of the confidence conversion submodule so that the last activation layer of the confidence conversion submodule outputs the scaling factor; The scaling factor and the third feature are input into the aggregation layer in the feature conversion submodule, and the first feature is input into the summation layer in the feature conversion submodule, so that the last activation layer in the feature conversion submodule outputs the fused feature.
10. The method according to claim 8, characterized in that The three-dimensional structure is determined based on a structure detection module in the structure detection model, wherein the structure detection module includes a variable self-decomposition encoder.
11. The method according to any one of claims 1 to 10, characterized in that The first feature is a tensor with a shape of N×384, the third feature is a tensor with a shape of N×64, and the second feature is a tensor with a shape of N×N×128, where N is the total length of amino acids in the heavy chain and light chain in the first antibody sequence.
12. The method according to any one of claims 8 to 10, characterized in that The method further comprises: The structure detection model to be trained is trained by adopting a teacher-student self-supervised learning method to obtain the trained structure detection model.
13. The method according to claim 12, characterized in that The method of training the structure detection model to be trained by using a teacher-student self-supervised learning method to obtain the trained structure detection model includes: Based on the training data set, the structure detection model to be trained is trained by a teacher-student self-supervised learning method to obtain the trained structure detection model, wherein the training data set includes at least one training data group, and the training data group includes an unlabeled antibody sequence and a labeled antibody sequence.
14. The method according to claim 13, characterized in that The training process of the structure detection model includes: Obtain the teacher model and student model corresponding to the structure detection model to be trained; Inputting a first unlabeled training antibody sequence in the training data set into the student model to obtain a first detection result, and inputting the first training antibody sequence into the teacher model to obtain a second detection result; Calculate a first loss according to the first detection result and the second detection result; Inputting the labeled second training antibody sequence in the training data set into the student model to obtain a third detection result; Calculate a second loss based on the third detection result and the true label of the second training antibody sequence; Calculating the sum of the first loss and the second loss to obtain a third loss; According to the third loss, updating the parameters of the student model by back-propagation gradient method; The parameters of the teacher model are updated by performing exponential sliding average processing on the updated parameters of the student model.
15. A readable medium, characterized in that The readable medium stores instructions, which, when executed on an electronic device, enable the electronic device to perform the antibody sequence detection method according to any one of claims 1 to 14.
16. An electronic device, characterized in that: include: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, which is one of the processors of the electronic device, for executing the antibody sequence detection method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Method and system for predicting amino acid sequence in antibody protein CDR region
CN113838523A
Binding affinity prediction method and device based on antigen and antibody sequences
CN114464247A
Sequence-based antigen-antibody affinity prediction method
CN116434839A
Method for preparing novel antibody library and library prepared thereby
US20170362306A1
Protein structure prediction
WO2024072980A1