Antibody sequence detection methods, media, and devices

By integrating deep learning methods with homologous sequence features and language model features, the challenge of detecting antibody CDR structures was solved, and the accuracy of antibody three-dimensional structure detection was improved.

CN119943129BActive Publication Date: 2026-02-10FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510026906.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2026-02-10
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately detect the complementarity-determining region (CDR) structure of antibodies, which affects the detection accuracy of antibody three-dimensional structures.

Method used

By fusing homologous sequence features and language model features, and utilizing deep learning models such as AlphaFold2 and AntiBERTy, combined with the evolutionary constraints and dependencies of amino acids, the three-dimensional structure of antibody sequences is detected.

Benefits of technology

This improved the accuracy of antibody CDR loop structure detection and enhanced the overall structure detection accuracy of antibody sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943129B_ABST
    Figure CN119943129B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of antibody detection, and discloses an antibody sequence detection method, a medium and equipment, which can improve the structural detection precision of CDR-H3 loops in antibodies. The method comprises the following steps: obtaining a first antibody sequence, the first antibody sequence comprising paired heavy chains and light chains; obtaining a first feature, a second feature and a third feature of the first antibody sequence, wherein the first feature is used for indicating evolutionary constraint information between the first antibody sequence and a homologous sequence, the second feature is used for indicating the mutual relationship between amino acid pairs in the first antibody sequence, and the third feature is used for indicating the dependency relationship between amino acids in the first antibody sequence; performing fusion processing on the first feature and the third feature to obtain a fusion feature of the first antibody sequence; and detecting a three-dimensional structure of the first antibody sequence according to the second feature and the fusion feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of antibody detection technology, and in particular to an antibody sequence detection method, medium, and device. Background Technology

[0002] Antibodies are immunoglobulins produced by B lymphocytes and are a core component of the adaptive immune system. Antibodies recognize and neutralize antigens through specific binding mechanisms; therefore, in antibody drug development and application, accurate detection of the antibody's three-dimensional structure is crucial for improving the affinity and specificity between antibodies and antigens. The sequence and structural diversity of the antibody's complementarity determining region (CDR) are key factors determining the specificity and affinity between antibodies and antigens. However, due to the extreme diversity of antibody CDRs and the relatively small number of antibody samples with known structures, the detection of antibody CDR structures is challenging, affecting the accuracy of antibody three-dimensional structure detection. Summary of the Invention

[0003] This application provides an antibody sequence detection method, medium, and device that can improve the accuracy of antibody structure detection.

[0004] In a first aspect, embodiments of this application provide an antibody sequence detection method, the method comprising: obtaining a first antibody sequence, the first antibody sequence comprising a pair of heavy chains and light chains; obtaining a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and homologous sequences, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate the dependency relationship between amino acids in the first antibody sequence; fusing the first feature and the third feature to obtain a fusion feature of the first antibody sequence; and detecting the three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature.

[0005] As an example, the first feature of the first antibody sequence can be called a homology sequence feature, which is the evolutionary constraint information between the first antibody sequence and homologous sequences, and can be used to detect the three-dimensional structure of amino acids in the first antibody sequence. As an example, the third feature of the first antibody sequence can also be called a language model feature, which can be used to detect the three-dimensional structure of amino acids in the first antibody sequence.

[0006] Therefore, the antibody sequence structure detection method provided in this application can integrate the homologous sequence features and language model features of the first antibody sequence to be detected. That is, it considers not only the influence of homologous sequence features on the structure, but also the influence of language model features on the structure. Thus, by comprehensively considering the influence of homologous sequence features and language models on the structure of the first antibody sequence, the structural detection accuracy of the CDR-H3 loop in the complementarity-determining region (CDR) can be improved, thereby enhancing the overall structural detection accuracy of the antibody sequence.

[0007] In one possible implementation of the first aspect described above, the fusion processing of the first feature and the third feature to obtain the fusion feature of the first antibody sequence includes: converting the root mean square error information of the first antibody sequence into a scaling factor, wherein the root mean square error information includes the root mean square error of each amino acid backbone atom in the first antibody sequence, and the scaling factor has a one-dimensional dimension; performing a dimensional transformation on the third feature according to the scaling factor to obtain a fourth feature, wherein the fourth feature has the same dimension as the first feature; and summing the fourth feature and the first feature to obtain the fusion feature. Thus, the first feature and the third feature of the first antibody sequence can be fused using the first root mean square error information to obtain the fusion feature.

[0008] In one possible implementation of the first aspect described above, the first feature and the second feature are determined by: searching for homologous sequences of the first antibody sequence; obtaining a multiple sequence alignment representation between the first antibody sequence and the homologous sequence; linking the heavy and light chains in the first antibody sequence into a single sequence to obtain the first protein sequence; pairing different amino acids in the first protein sequence to obtain a paired representation of the first protein sequence; and encoding the multiple sequence alignment representation and the paired representation to obtain the first feature and the second feature, respectively. As an example, the first feature and the third feature described above can be determined based on homologous sequences of the first antibody sequence using deep learning models such as AlphaFold2 or AlphaFold-Multimer through transfer learning.

[0009] In one possible implementation of the first aspect described above, the third feature is determined as follows: the heavy and light chains in the first antibody sequence are joined together by vacancies to form a second protein sequence; the second protein sequence is input into a pre-trained protein language model, which outputs a first language model encoding and a first attention matrix; the first language model encoding and the first attention matrix are updated based on an attention mechanism to obtain the third feature. As an example, the aforementioned third feature can be obtained by using transfer learning with a protein language model such as AntiBERTy to capture the complex dependencies between amino acids in the first antibody sequence.

[0010] In one possible implementation of the first aspect described above, detecting the three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature includes: updating the fusion feature and the second feature based on an attention mechanism; and detecting the three-dimensional structure of the first antibody sequence based on the updated second feature and the updated fusion feature. Thus, updating the second feature and the fusion feature based on an attention mechanism helps improve the accuracy of detecting the three-dimensional structure of the first antibody sequence.

[0011] In one possible implementation of the first aspect described above, the three-dimensional structure of the first antibody sequence is detected based on the updated second feature and the updated fusion feature, including: detecting the mean and variance of the translation parameters of amino acids in the first antibody sequence, and detecting the mean and variance of the rotation parameters of amino acids in the first antibody sequence, based on the updated second feature and the updated fusion feature; and using a reparameterization method, determining the target values ​​of the translation parameters and rotation parameters of each amino acid in the first antibody sequence based on the mean and variance of the translation parameters and the mean and variance of the rotation parameters, wherein the translation parameter represents the position of the amino acid, and the rotation parameter represents the angle of the amino acid, and the three-dimensional structure includes the position and angle of the amino acids. It is understood that the structure of the CDR ring of an antibody is dynamic and may change due to the influence of the antibody and environmental factors. Therefore, when predicting the antibody structure, this application does not directly predict the specific values ​​of the translation and rotation of each amino acid, but predicts their respective mean and variance, and then obtains the translation and rotation of each amino acid through reparameterization, which is beneficial to improving the accuracy of structure detection.

[0012] In one possible implementation of the first aspect described above, different amino acids in the first antibody sequence have different positional codes. It is understood that during the structural detection of the first antibody sequence, the positions of different amino acids in the first antibody sequence can be identified based on their positional codes.

[0013] In one possible implementation of the first aspect described above, the three-dimensional structure is determined based on a structure detection model. This model includes a feature fusion module, which comprises a confidence transformation submodule and a feature transformation submodule. The confidence transformation submodule includes three sequentially connected network layers, each consisting of a sequentially connected linear layer and an activation layer. The feature transformation submodule includes a sequentially connected aggregation layer, a linear layer, an activation layer, a summation layer, another linear layer, and an activation layer. The fused features are determined based on the feature fusion module. For example, in the feature transformation submodule 142, the aggregation layer is a Multiplate layer, the activation layer is a ReLU layer, and the linear layer can be a fully connected (FC) layer.

[0014] In one possible implementation of the first aspect above, the process of the feature fusion module determining the fused features includes: inputting the root mean square error information into the first linear layer of the confidence transformation submodule, so that the last activation layer of the confidence transformation submodule outputs the scaling factor; inputting the scaling factor and the third feature into the aggregation layer in the feature transformation submodule, and inputting the first feature into the summation layer in the feature transformation submodule, so that the last activation layer in the feature transformation submodule outputs the fused features.

[0015] In one possible implementation of the first aspect described above, the three-dimensional structure is determined based on a structure detection module in a structure detection model, which includes a variational autoencoder. Thus, modeling the dynamics of the antibody structure using a variational autoencoder helps improve the accuracy of structure detection.

[0016] In one possible implementation of the first aspect above, the first feature is a tensor of shape N×384, the third feature is a tensor of shape N×64, and the second feature is a tensor of shape N×N×128, where N is the total length of amino acids in the heavy and light chains of the first antibody sequence.

[0017] In one possible implementation of the first aspect described above, the method further includes: training the structure detection model to be trained using a teacher-student self-supervised learning approach to obtain a trained structure detection model. This application employs a teacher-student self-supervised learning method, which can introduce a large number of unlabeled antibody sequences for self-supervised training, using them in conjunction with labeled antibody sequences for model training, thereby improving the model's predictive ability.

[0018] In one possible implementation of the first aspect described above, a teacher-student self-supervised learning method is used to train the structure detection model to be trained, resulting in a trained structure detection model. This includes: training the structure detection model to be trained using a teacher-student self-supervised learning method based on a training dataset, wherein the training dataset includes at least one training data set, which includes one unlabeled antibody sequence and one labeled antibody sequence. This effectively propagates a small amount of labeled structural information to a large number of unlabeled antibody sequences, thereby improving the model's ability to predict the structure of antibody sequences.

[0019] In one possible implementation of the first aspect above, the training process of the structure detection model includes: obtaining the teacher model and student model corresponding to the structure detection model to be trained; inputting the first unlabeled training antibody sequence from the training data set into the student model to obtain a first detection result, and inputting the first training antibody sequence into the teacher model to obtain a second detection result; calculating a first loss based on the first and second detection results; inputting the second labeled training antibody sequence from the training data set into the student model to obtain a third detection result; calculating a second loss based on the third detection result and the true label of the second training antibody sequence; calculating the sum of the first and second losses to obtain a third loss; updating the parameters of the student model using backpropagation gradient based on the third loss; and updating the parameters of the teacher model by performing exponential moving average processing on the updated parameters of the student model. As an example, during the training process, each mini-batch contains half labeled antibody sequences and half unlabeled antibody sequences. The labeled antibody data uses the native structure (i.e., the true label) as the supervision signal, while the unlabeled training antibody sequences use the three-dimensional structure predicted by the teacher model as the supervision signal.

[0020] Secondly, embodiments of this application provide a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the antibody sequence detection method as described in the first aspect and any possible implementation thereof.

[0021] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for executing the antibody sequence detection method as described in the first aspect and any possible implementation thereof.

[0022] The beneficial effects of the second and third aspects of this application can be referred to the description in the first aspect and its various possible implementations, and will not be repeated here. Attached Figure Description

[0023] Figure 1 According to some embodiments of this application, a structural block diagram of an antibody structure detection model is shown;

[0024] Figure 2 According to some embodiments of this application, a schematic diagram of a homologous sequence feature extraction process is shown;

[0025] Figure 3 According to some embodiments of this application, a schematic diagram of a language model feature extraction process is shown;

[0026] Figure 4 According to some embodiments of this application, a schematic diagram of a feature fusion process is shown;

[0027] Figure 5 According to some embodiments of this application, a flowchart of an antibody sequence detection method is shown;

[0028] Figure 6 According to some embodiments of this application, a flowchart of an antibody sequence detection method based on a structure detection model is shown;

[0029] Figure 7 According to some embodiments of this application, a flowchart of an antibody sequence detection method is shown;

[0030] Figure 8 According to some embodiments of this application, a schematic diagram of the training process of an antibody structure detection model is shown;

[0031] Figure 9 According to some embodiments of this application, a schematic diagram of a network architecture for model training based on a teacher-student self-supervised learning method is shown.

[0032] Figure 10 According to some embodiments of this application, a schematic diagram of a training process for a structure detection model based on a training data set is shown;

[0033] Figure 11A Box plots of prediction accuracy for various antibody structure prediction methods are shown according to some embodiments of this application;

[0034] Figure 11B According to some embodiments of this application, schematic diagrams of test results for various models on a test set are shown;

[0035] Figure 11C According to some embodiments of this application, schematic diagrams of test results for various models on a test set are shown;

[0036] Figure 12A According to some embodiments of this application, a scatter plot comparing AbFold and AlphaFold2-Multimer is shown;

[0037] Figure 12B According to some embodiments of this application, a scatter plot comparing AbFold and IgFold is shown;

[0038] Figure 12C According to some embodiments of this application, a schematic diagram comparing the CDR-H3 structure predicted by the three models AbFold, AlphaFold2 and IgFold with the natural structure is shown.

[0039] Figure 13According to some embodiments of this application, a schematic diagram of the structure of an electronic device is shown. Detailed Implementation

[0040] The illustrative embodiments of this application include, but are not limited to, antibody sequence structure detection methods, media, and electronic devices.

[0041] To more clearly describe the embodiments of this application, some terms provided in the embodiments of this application will be introduced below.

[0042] Antibodies, also known as immunoglobulins, are immune-functional proteins produced in the serum of humans and animals in response to the invasion of pathogens or viruses. They are widely used in medicine, biology, and other fields.

[0043] An antigen is a substance that can stimulate an organism to produce an immune response and can bind to antibodies and other lymphocytes both inside and outside the body, resulting in an immune effect (specific reaction). The binding between antibodies and antigens exhibits specificity and affinity.

[0044] Specificity refers to the high selectivity of the binding reaction between antigens and antibodies; that is, an antibody can only bind to one or a class of specific antigens, and an antigen can only bind to one or a class of specific antibodies. This specific binding of antigens and antibodies is based on the complementary structures of their molecular surfaces.

[0045] Affinity refers to the binding strength between an antigen binding site on an antibody and a corresponding antigen epitope, which depends on the degree of complementarity of their spatial structures.

[0046] An antibody sequence refers to the arrangement of amino acids in an antibody, which determines the antibody's structure and function. In the following examples, the structure of the antibody can also be described as the structure of its antibody sequence.

[0047] The conformation of an antibody typically refers to the relative arrangement and orientation of atoms or groups within the molecule in space. This arrangement and orientation determine the shape and three-dimensional structure of the molecule. For antibodies, the conformation refers to the folding pattern of the peptide chains (including heavy and light chains), the linkage of disulfide bonds between chains, and the relative positions of variable and constant regions within the antibody molecule.

[0048] The three-dimensional structure of an antibody refers to the specific shape and arrangement of the molecule in three-dimensional space. For antibodies, their three-dimensional structure is obtained through direct observation and analysis using modern biophysical techniques such as X-ray crystallography and nuclear magnetic resonance (NMR). The three-dimensional structure of an antibody reveals its detailed molecular conformation, including the folding pattern of the peptide chains, inter- and intra-chain interactions, and the specific positions and configurations of variable and constant regions. This structural information is crucial for understanding the antigen-binding mechanism, biological function, and the design and application of antibody engineering.

[0049] Specifically, antibodies are large protein molecules produced by B lymphocytes. Their monomers can have a Y-shaped structure, consisting of two identical heavy chains and two identical light chains linked by disulfide bonds. As an example, antibodies are typically formed by the folding of the heavy and light chains into a ring structure; for instance, a Y-type antibody contains two symmetrical pairs of heavy and light chains. Correspondingly, the antibody sequence can be a sequence composed of amino acids from paired heavy and light chains.

[0050] The key functional region of an antibody is located at the apex of its ring structure, in the variable region (Fv), which is responsible for specific antibody recognition. The variable region contains a complementarity-determining region (CDR) and a backbone region (Fr), the latter being the portion of the variable region excluding the CDR. The CDR consists of six highly variable rings: three rings (CDR-H1, CDR-H2, CDR-H3) are located on the heavy chain, and three rings (CDR-L1, CDR-L2, CDR-L3) are located on the light chain. These rings are responsible for antigen recognition and binding. The sequence and structural diversity of the CDR is a key factor determining antibody specificity and affinity. The CDR-H3 ring, in particular, exhibits significant structural diversity. Located at the boundary between the heavy and light chains, its structure is influenced by the interactions between the two chains, making its structural detection challenging.

[0051] Amino acids are the basic building blocks of proteins. Each amino acid contains one or more amino groups (-NH2) and carboxyl groups (-COOH), as well as a specific side chain group (R group). During protein synthesis, amino acids are linked by peptide bonds to form polypeptide chains, which then fold into proteins with specific functions, such as antibodies.

[0052] An amino acid pair is a combination of two adjacent or non-adjacent amino acids in a protein sequence. The interactions between amino acid pairs are one of the key factors determining the three-dimensional structure and function of proteins. These interactions include hydrogen bonds, hydrophobic interactions, ionic bonds, and van der Waals forces, which together maintain the stability and function of proteins.

[0053] An amino acid residue is the structural part remaining after an amino acid that makes up a polypeptide loses a water molecule when some of its groups (such as amino and carboxyl groups) participate in the formation of a peptide bond. In other words, an amino acid residue is the part remaining after an amino acid forms a peptide bond. In a protein sequence, each amino acid residue occupies a specific position and is linked to other amino acid residues through a peptide bond.

[0054] As mentioned earlier, the structure of the complementarity-determining region, especially the CDR-H3 loop, is difficult to detect, which affects the accuracy of antibody three-dimensional structure detection.

[0055] In some embodiments, deep learning models such as AlphaFold2 or AlphaFold-Multimer can determine homology sequence features, such as evolutionary constraints, of the antibody sequence to be detected based on homology sequences, and detect the three-dimensional structure (i.e., protein structure) of the antibody sequence based on these homology sequence features. Homology sequences refer to amino acid sequences that share a common evolutionary ancestor. The evolutionary trajectory of the amino acid sequence of a protein (such as an antibody) is constrained by its function; its evolutionary constraints can be inferred through a set of homology sequences, thereby enabling the detection of protein structure.

[0056] In other embodiments, deep learning models such as IgFold can determine the language model features of the antibody sequence to be detected based on pre-trained protein language models such as BERT and AntiBERTy, and then detect the three-dimensional structure of the antibody sequence based on these language model features. These language model features can characterize the complex dependencies between amino acids in the antibody sequence. It can be understood that protein language models work by simulating the interdependencies between amino acids in proteins. Specifically, the model uses a neural network architecture called Transformer to learn sequence dependencies from data in contexts of arbitrary length. When processing antibody sequences, protein language models can analyze the position, type, and interactions of each amino acid with other amino acids, thereby capturing the dependencies between amino acids.

[0057] However, for the same antibody, the aforementioned detection methods based on homology sequences and those based on language model features often show significant differences in detecting CDR loops, such as the CDR-H3 loop structure. Furthermore, for many antibodies, one method may be accurate while the other has a larger error. It's understandable that homology sequence-based detection methods only consider the influence of homology sequence features on the structure, while language model-based detection methods only consider the influence of language model features. However, the influence of homology sequence features and language model features on the structure of different antibody sequences is usually different. Therefore, when the structure of an antibody sequence is highly influenced by homology sequence features, detecting the structure using only language model features will result in a larger error. Conversely, when the structure of an antibody sequence is highly influenced by language model features, detecting the structure using only homology sequence features will also result in a larger error.

[0058] Therefore, to improve the accuracy of detecting the CDR H3 ring structure in antibodies, this application provides an antibody sequence detection method. Specifically, the method includes: obtaining a first antibody sequence, the first antibody sequence comprising paired heavy and light chains, i.e., the antibody sequence comprising amino acids in three loops of the heavy chain and three loops of the light chain; obtaining a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and homologous sequences, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence (i.e., amino acid pairing relationship), and the third feature is used to indicate the dependency relationship between amino acids in the first antibody sequence; fusing the first feature and the third feature to obtain a fusion feature of the first antibody sequence; and detecting the three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature.

[0059] As an example, the first and third features mentioned above can be determined based on homologous sequences of the first antibody sequence using deep learning models such as AlphaFold2 or AlphaFold-Multimer. In this case, the first feature of the first antibody sequence (also known as the homologous sequence feature), that is, the evolutionary constraint information between the first antibody sequence and homologous sequences, can be used to detect the three-dimensional structure of amino acids in the first antibody sequence.

[0060] As an example, the aforementioned third feature can be obtained by using protein language models such as AntiBERTy to capture the complex dependencies between amino acids in the first antibody sequence. In this case, the third feature of the first antibody sequence (also known as a language model feature) can be used to detect the three-dimensional structure of the amino acids in the first antibody sequence.

[0061] Therefore, the antibody sequence structure detection method provided in this application can improve the accuracy of CDR H3 loop structure detection by fusing homologous sequence features and language model features of the first antibody sequence to be detected, combining the advantages of both methods, thereby improving the overall structure detection accuracy of the antibody sequence. Specifically, this application can perform antibody structure detection based on homologous sequence features and language models of the first antibody sequence, comprehensively considering the influence of homologous sequence features and language models on the structure of the first antibody sequence, which helps to reduce detection errors.

[0062] In some embodiments of this application, the subject performing the antibody sequence detection method may be an electronic device or a device in an electronic device for performing antibody structure detection.

[0063] As examples, electronic devices applicable to this application include, but are not limited to, mobile phones, tablets, wearable electronic devices, in-vehicle electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). Of course, the aforementioned electronic devices are not limited to the examples above, but also include hardware servers, cloud servers, and other devices; this application does not limit these to specific devices.

[0064] In some embodiments, the electronic device in this application can perform antibody sequence detection methods using a deep learning-based antibody structure detection model. For example, the antibody structure detection model provided in this application may be called AbFold, or a detection network, etc. The electronic device used in the antibody structure detection model of this application may be the same or different during the model training and model usage phases; this application does not specifically limit this.

[0065] Next, combined Figures 1 to 4 The structure and function of the pre-trained structure detection model in this application are introduced.

[0066] Reference Figure 1 This is a structural block diagram of an antibody structure detection model provided for the implementation of this application. Figure 1 As shown, the structure detection model 01 includes the following modules: homology feature extraction module 11, language model feature extraction module 12, root mean square error determination module 13, feature fusion module 14, feature update module 15, amino acid position encoding module 16, and structure detection module 17.

[0067] The homology feature extraction module 11 is used to search for homologous sequences of the input antibody sequence, such as the first antibody sequence, extract evolutionary constraint information (i.e., first feature) between the input antibody sequence and the homologous sequence, and extract the correlation relationship of amino acid pairs in the input antibody sequence (i.e., second feature).

[0068] The language model feature extraction module 12 is used to extract the dependencies between amino acids in the antibody sequence (i.e., the third feature).

[0069] Root mean square error determination module 13 is used to calculate the root mean square error of each amino acid backbone atom (C) in the input antibody sequence. α C β The root mean square error (RMSE) of (N, O) is denoted as pRMSD. RMSE, also known as root mean square deviation (RMSD), is a commonly used measure of the difference between measured values.

[0070] Feature fusion module 14 is used to fuse the first and third features of the input antibody sequence to obtain fused features. For example, feature fusion module 14 can fuse the first and third features of the input antibody sequence based on the root mean square error of each amino acid backbone atom in the input antibody sequence. The specific process of fusion processing will be described below and will not be detailed here.

[0071] The feature update module 15 is used to update the fusion features and second features of the input antibody sequence based on attention mechanisms such as triangular attention mechanisms.

[0072] The amino acid position encoding module 16 is used to encode the position of amino acids in the input antibody sequence to obtain the amino acid position encoding features of the input antibody sequence.

[0073] Structure detection model 17 is used to perform antibody structure detection on the input antibody sequence based on the third feature and fusion feature of the input antibody sequence, so as to obtain the three-dimensional structure of the antibody sequence. For example, structure detection model 17 can perform antibody structure detection on the input antibody sequence based on the third feature, fusion feature and amino acid position coding feature of the input antibody sequence.

[0074] Understandable. Figure 1 The structure of the structure detection model 01 shown is only an example. In other embodiments, the structure detection model 01 may include more or fewer modules. For example, in some embodiments, the structure detection model 01 may not include modules such as the amino acid position encoding module 16.

[0075] In some embodiments, the structure detection model 01 in this application, such as AbFold, can extract features such as evolutionary constraints between the input antibody sequence and homologous sequences, correlations between amino acid pairs, and dependencies between amino acids through transfer learning. That is, it can extract features such as the first feature, second feature, and third feature of the input antibody sequence through transfer learning.

[0076] As an example, the homology feature extraction module 11 can use AlphaFold2 to extract the first and second features of the input antibody sequence through transfer learning, for example, by using the Evoformer module (an encoder) in AlphaFold2 to extract the first and second features.

[0077] It is understandable that transfer learning can transfer knowledge learned in a source domain to a target domain, thereby solving the problem of insufficient labeled data in the target domain. In antibody structure prediction tasks, the limited number of high-quality antibody structures makes it difficult for neural networks to fully learn the structural features of antibodies from a small amount of labeled data. Therefore, this application uses general protein structure prediction as the source domain and antibody structure prediction as the target domain. By extracting key features of antibodies through transfer learning, the accuracy of antibody structure prediction can be improved. Among them, AlphaFold2 is an advanced model in the field of general protein structure prediction, and its training process uses a large number of known structures as well as protein structures generated through self-distillation.

[0078] In some embodiments, the homology sequence feature extraction module 11 can extract features such as the first feature and the second feature from the homology sequences of the input antibody sequence using AlphaFold2-Multimer through transfer learning. AlphaFold2-Multimer provides five sets of model parameters; the structure detection model 01 in this application uses the first set of parameters for structure prediction and feature extraction. Since homology sequence searching with AlphaFold2-Multimer is time-consuming (approximately one hour per antibody), this application employs an acceleration scheme from ColabFold (a deep learning model), specifically using search tools such as MMseqs2 to search protein databases like Uniref90, reducing the homology sequence search time for each antibody to the minute level.

[0079] Reference Figure 2 This is a flowchart of homologous sequence feature extraction provided in an embodiment of this application. This extraction process can be executed by the aforementioned homologous feature extraction module 11. Specifically, Figure 2The illustrated process includes: the homology feature extraction module 11 acquires the input antibody sequence (input sequence) composed of paired heavy and light chain amino acids; the MMseqs2 search tool searches the Uniref90 database to identify homologous sequences of the input antibody sequence; and a multiple sequence alignment (MSA) representation is generated between the input antibody sequence and the homologous sequence. Figure 2 In the MSA representation shown, the dark squares represent sequences with the same evolutionary constraint information. Furthermore, the homology feature extraction module 11 can use AlphaFold2-Multimer to connect the amino acids of the heavy and light chains in the input antibody sequence into a single sequence through transfer learning, and perform pairing processing on the amino acids in this sequence to obtain the paired representation (pair repr, i.e., pair representation) of the sequence. Figure 2 The pair representation includes N×N dimensional data blocks, each corresponding to two amino acids in the input antibody sequence.

[0080] Furthermore, the homology feature extraction module 11 can use transfer learning to encode the MSA representation and pair representation of the input antibody sequence using the Evoformer module (encoder) in AlphaFold2-Multimer, such as performing multiple encodings (e.g., 3 times) to output tensors S1 and Z. Tensor S1 can have a shape of N×384, representing the multidimensional features of the input antibody sequence, specifically the evolutionary constraint information between the input antibody sequence and homologous sequences. Tensor Z can have a shape of N×N×128, representing the pairwise information between any two amino acids in the input antibody sequence, i.e., the correlation between amino acid pairs in the input antibody sequence. N is the sum of the heavy chain and light chain lengths in the input antibody sequence, i.e., the total number of amino acids. In this case, tensor S1 represents the first feature of the input antibody sequence, and tensor Z represents the second feature of the input antibody sequence.

[0081] Furthermore, as an example, the language model feature extraction module 12 can use transfer learning to extract the third feature of the input antibody sequence, namely the dependencies between amino acids, using the protein language model AntiBERTy and the feature encoding module in IgFold. It can be understood that AntiBERTy is a protein language model pre-trained on a large number of natural antibody sequences.

[0082] Reference Figure 3 This is a flowchart of language model feature extraction provided in an embodiment of this application. This extraction process can be executed by the language model feature extraction module 12 described above. Specifically, Figure 3The illustrated process includes: the language model feature extraction module 12 connects the heavy and light chains of the input antibody sequence into a single sequence using spaces (or separators), inputs this sequence into the protein language model AntiBERTy, and generates a language model encoding and attention matrix for this sequence. Furthermore, the language model feature extraction module 12 can use transfer learning with the IgFold feature encoding module to interact with the generated language model encoding and attention matrix, further extracting the tensor S2 of the input antibody sequence. For example, the dimension of tensor S2 is N×64. Here, tensor S2 represents the third feature of the input antibody sequence, that is, tensor S2 is used to indicate the dependencies between amino acids in the input antibody sequence.

[0083] In some embodiments, the root mean square error calculation module 13 can use IgFold to predict the root mean square error (pRMSD) of each amino acid backbone atom in the input antibody sequence through transfer learning. It is understood that root mean square error (RMSD) is typically used to measure the deviation between the model-predicted structure and the actual structure, thereby reflecting the accuracy of various parts of the structure detection model and assessing the reliability of the predicted structure. Studies have shown that the predicted pRMSD is highly correlated with the actual error in structure detection, and therefore can be used as a basis for assessing the reliability of the predicted structure. This application can use IgFold to obtain the pRMSD of the predicted three-dimensional structure of the input antibody sequence to assess the reliability of antibody sequence features extracted from the protein language model, such as the second feature, and use it in the subsequent feature fusion process. Furthermore, the dimension of the root mean square error (pRMSD) of each amino acid backbone atom in the input antibody sequence is N×4.

[0084] In some embodiments, the feature fusion module 14 includes a confidence transformation submodule and a feature transformation submodule. The confidence transformation submodule includes three sets of network layers connected in sequence, each set including a linear layer and an activation layer connected in sequence. For example, the linear layer in the confidence transformation submodule 141 is a fully connected layer (FC layer), and the activation layer is a sigmoid layer. Furthermore, the feature transformation submodule 142 includes an aggregation layer, a linear layer, an activation layer, a summation layer, another linear layer, and an activation layer connected in sequence. For example, the aggregation layer in the feature transformation submodule 142 is a Multiplate layer, the activation layer is a ReLU layer, and the linear layer can be a fully connected layer (FC).

[0085] Reference Figure 4 This is a flowchart of the feature fusion process provided in an embodiment of this application. This process can be executed by the feature fusion module 14 described above. Figure 4 As shown, the feature fusion module 14 includes a confidence conversion submodule 141 and a feature transformation submodule 142.

[0086] As an example, Figure 4 The confidence conversion submodule 141 in the middle includes FC(4,32) layer, Sigmoid layer, FC(32,32) layer, Sigmoid layer, FC(32,1) layer and Sigmoid layer in sequence. Figure 4 The feature transformation submodule 142 in the middle includes, in sequence, a Multiplate layer, an FC(64,384) layer, a ReLU layer, an Add layer, an FC(384,384) layer, and a Sigmoid layer. Specifically, as shown in... Figure 4 As shown, the N×4 root mean square error pRMSD is input to the confidence transformation submodule 141, and is converted into a 1-dimensional scaling factor through multiple layers of the confidence transformation submodule 141. Then, this scaling factor and the tensor S2 of the input antibody sequence are input to the Multiplate layer of the feature transformation submodule 142, and the tensor S1 of the input antibody sequence is input to the Add layer of the feature transformation submodule 142. Tensor S2 is then summed with tensor S1 in the Add layer after passing through an FC(64,384) layer and a ReLU layer. The resulting tensor is then transformed through an FC(64,384) layer and a Sigmoid layer, and output as tensor S. At this point, tensor S represents the fusion feature of the input antibody sequence, and tensor S has an N×384 dimension.

[0087] In some embodiments, the feature update module 15 may update the fusion features such as tensor S and the second feature such as tensor z of the input antibody sequence based on an attention update mechanism.

[0088] As an example, the feature update module 15 can be based on a graph attention network (GAT) and employ graph attention mechanism and triangular update mechanism to update the integration features and second features of the input antibody sequence.

[0089] Graph attention networks (GNNs) are models based on graph neural networks used to process graph-structured data. GNNs optimize the update process of node features by introducing an attention mechanism. In a GNN, each node calculates attention weights based on the features of its neighbors. These weights are used to weighted aggregate the features of its neighbors, thereby updating the feature representation of the current node.

[0090] In some embodiments, the amino acid position encoding module 16, when encoding the position of the input antibody sequence, starts the encoding of the heavy chain and light chain from 0, and increments the position code by 1 for each amino acid. To distinguish between the heavy chain and the light chain, each position code of the light chain is uniformly incremented by 256 after the initial encoding is completed. The final dimension of the amino acid position encoding feature of the input antibody sequence is N×1. It can be understood that the type and position of amino acids can uniquely identify an antibody sequence, and accurate amino acid position encoding can significantly improve the accuracy of structure prediction. Furthermore, amino acid position encoding can also help the Transformer-based structure detection module 17 identify the positions of different amino acids in the input antibody sequence.

[0091] In some embodiments, the structure detection module 17 can model the antibody structure dynamics of the input antibody sequence using a variational autoencoder based on the tensor S of the updated fusion features and the tensor z of the second features. The structure of the input antibody sequence, i.e., the three-dimensional structure, includes translation and rotation parameters for each amino acid; the translation parameters indicate the position of the amino acid, and the rotation parameters indicate the position of the amino acid.

[0092] Furthermore, in some embodiments, the structure detection module 17 can use an invariant point attention mechanism for antibody structure detection. It is understood that the core idea of ​​the invariant point attention mechanism is to maintain attention on certain key points when processing data with spatial or structural features, and to aggregate and infer information based on these key points. In protein structure prediction, this mechanism is used to calculate the individual parts of a protein separately, thereby constructing an accurate three-dimensional structure.

[0093] As an example, the implementation of the invariant point attention mechanism includes keypoint selection, attention weight calculation, information aggregation, and 3D structure construction. Specifically, in the keypoint selection process, the model determines which atomic or molecular structural points are keypoints. These keypoints are typically critical residues or atoms with specific functions in the protein structure. In the attention weight calculation process, for each keypoint, the model calculates its attention weight relative to surrounding atoms or structural points. These weights reflect the degree of association and importance between the keypoint and its surrounding points. In the information aggregation process, based on the attention weights, the model aggregates information from surrounding points onto the keypoint, thus forming a more accurate and robust structural representation. In the 3D structure construction process, the model constructs the 3D structure of the protein based on these keypoints and their aggregated information.

[0094] It is understandable that research has found the structure of the CDR ring of antibodies to be dynamic and may change due to the influence of antibody and environmental factors. Therefore, this application does not use traditional deterministic structure prediction networks, but instead uses a variational autoencoder to model this dynamic nature of antibody structure. Specifically, this application represents the antibody structure by predicting the translation and rotation of amino acids in the antibody. When predicting the antibody structure, this application does not directly predict the specific values ​​of the translation and rotation of each amino acid, but predicts their respective means and variances, and then obtains the translation and rotation of each amino acid through reparameterization.

[0095] As an example, the structure detection module 17 in this application can use formula (1) to calculate the translation and rotation parameters of each amino acid in the input antibody sequence.

[0096]

[0097] Where μ is the mean, σ is the variance, and ∈ is a random number sampled from the standard normal distribution.

[0098] For example, the structure detection module 17 can first predict the mean μ and variance σ of the translation parameters of the amino acids in the input antibody sequence. Then, for the translation parameter of an amino acid in the input antibody sequence, the mean μ, the variance σ, and a random number ∈ can be used to calculate the translation parameter of the amino acid using formula (1). Similarly, the structure detection module 17 can use a reparameterization method to determine the rotation parameters of each amino acid in the input antibody sequence.

[0099] Next, combined Figures 5 to 7 This paper provides a detailed explanation of the method and process for the model usage stage of the deep learning-based antibody structure detection model in this application.

[0100] Reference Figure 5 The diagram shown is a flowchart of an antibody sequence detection method provided in this application. The subject of this method can be an electronic device or a device in the electronic device.

[0101] Specifically, such as Figure 5 The process shown includes the following steps:

[0102] S501: Obtain the first antibody sequence, which includes a pair of heavy and light chains.

[0103] As an example, the first antibody sequence may include a pair of amino acids from the heavy chain and light chain. As another example, the first antibody sequence may include multiple pairs of amino acids from the heavy chain and light chain; for example, the antibody sequence for a Y-shaped antibody includes two pairs of heavy chains and light chains.

[0104] In some embodiments, the first antibody sequence can be used as the input antibody sequence to input into a pre-trained structural detection model 01, such as AbFold.

[0105] S502: Obtain the first feature, the second feature, and the third feature of the first antibody sequence, wherein the first feature is used to indicate the evolutionary constraint information between the first antibody sequence and homologous sequences, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate the dependency relationship between amino acids in the first antibody sequence.

[0106] In some embodiments, the first antibody sequence is input into the homology feature extraction module 11 in the pre-trained structure detection model 01, so as to extract the first feature and the second feature of the first antibody sequence through the homology feature extraction module 11. For example, the first feature is a tensor with a shape of N×384 (such as S1), and the second feature is a tensor with a shape of N×N×128 (such as tensor z).

[0107] In some embodiments, the first antibody sequence is input into the language model feature extraction module 12 in the pre-trained structure detection model 01 to extract the third feature of the first antibody sequence through the language model feature extraction module 12.

[0108] S503: The first and third features are fused to obtain the fusion feature of the first antibody sequence.

[0109] In some embodiments, this application may convert the first and third features of the first antibody sequence into features of the same dimension, and then fuse these two features to obtain fused features.

[0110] In some embodiments, the first feature and the third feature can be input into the feature fusion module 13 in the pre-trained structure detection model 01 to fuse the first feature and the third feature to obtain a fused feature. For example, the fused feature is a tensor (such as tensor S) with a dimension of N×384 and a shape of N×384.

[0111] S504: Detect the three-dimensional structure of the first antibody sequence based on the second feature and fusion feature.

[0112] In some embodiments, the fusion feature and the second feature of the first antibody sequence can be input into the structure detection module 14 in the pre-trained structure detection model 01, so that the first feature and the third feature can be fused by the structure detection module 13 to obtain the fusion feature.

[0113] It can be understood that the first and second features are homologous sequence features related to antibody structure, while the second feature is a language model feature related to antibody structure. Therefore, the fusion feature of the first antibody sequence can combine the homologous sequence features and language model features of the first antibody sequence. Thus, the antibody sequence detection method provided in this application can improve the accuracy of CDR H3 loop structure detection by fusing the homologous sequence features and language model features of the first antibody sequence to be detected, combining the influence of these two features on antibody structure detection, thereby improving the overall structure detection accuracy of the antibody sequence.

[0114] In some embodiments, such as Figure 6 The diagram shown is a flowchart illustrating an antibody sequence detection method provided in this application. The subject executing this method can be an electronic device or a device within that electronic device. As an example, Figure 6 The illustrated process can be executed by various modules of the pre-trained structure detection model 01 in this application, such as those in AbFold. It is understood that the following description... Figure 6 and Figure 5 The same steps will not be repeated here; the main difference is that... Figure 6 The processing flow of each module in the structural detection model 01 is shown.

[0115] Specifically, such as Figure 6 The process shown includes the following steps:

[0116] S601: Input the first antibody sequence into the structure detection model 01. The first antibody sequence includes a pair of heavy chains and light chains.

[0117] S602: The homology sequence feature extraction module 11 in the structure detection model 01 searches for homology sequences of the first antibody sequence and obtains the multi-sequence comparison representation between the first antibody sequence and the homology sequence.

[0118] In some embodiments, the homology sequence feature extraction module 11 uses MMseqs2 to search protein databases such as Uniref90 to obtain homology sequences of the first antibody sequence, with a relatively short search time.

[0119] In some embodiments, the homology sequence feature extraction module 11 uses the MSA representation between the first antibody sequence and the homology sequence to generate evolutionary constraint information indicating the relationship between the first antibody sequence and the homology sequence.

[0120] S603: The homology sequence feature extraction module 11 in the structure detection model 01 connects the heavy chain and light chain in the first antibody sequence into a single sequence to obtain the first protein sequence, and obtains the paired representation of the first protein sequence.

[0121] In some embodiments, the homology sequence feature extraction module 11 performs pairing processing on different amino acids in the first protein sequence to obtain a paired representation of the first protein sequence.

[0122] For example, the homology sequence feature extraction module 11 can pair up two amino acids in the first protein sequence to obtain the pair representation of the first protein sequence.

[0123] S604: The homologous sequence feature extraction module 11 in the structure detection model 01 encodes the multiple sequence contrast representation and pairing representation to obtain the first feature and the second feature, respectively.

[0124] In some embodiments, the homology sequence feature extraction module 11 can use transfer learning to encode the multiple sequence contrast representation and pairing representation of the first protein sequence using the Evoformer module in AlphaFold2-Multimer, to obtain the first feature of the first antibody sequence, such as tensor S1, and the second feature, such as tensor S2.

[0125] S605: The amino acid position encoding module 16 in the structure detection model 01 encodes the position of each amino acid in the first protein sequence to obtain the position encoding of each amino acid in the first antibody sequence.

[0126] The amino acid positions in the first antibody sequence are encoded differently at different locations.

[0127] For example, the amino acid position encoding module 16 can sequentially encode the first N / 2 amino acids belonging to the heavy chain in the first protein sequence as 0 to N / 2, and sequentially encode the last N / 2 amino acids belonging to the light chain in the first protein sequence as 256 to N / 2+256. This allows the position encoding of each amino acid to distinguish its position and whether it belongs to the heavy or light chain.

[0128] S606: The language model feature extraction module 12 in the structure detection model 01 connects the heavy chain and light chain in the first antibody sequence into a sequence through vacancies to obtain the second protein sequence, and obtains the first language model encoding and the first attention matrix of the second protein sequence.

[0129] In some embodiments, the language model feature extraction module 12 inputs the second protein sequence into a pre-trained protein language model and outputs a first language model encoding and a first attention matrix.

[0130] In some embodiments, the language model feature extraction module 12 can use transfer learning to identify the second protein sequence using the protein language model AntiBERTy, and determine the first language model encoding and the first attention matrix of the second protein sequence.

[0131] S607: The language model feature extraction module 12 in the structure detection model 01 updates the first language model encoding and the first attention matrix based on the attention mechanism to obtain the third feature.

[0132] In some embodiments, the language model feature extraction module 12 can use each amino acid of the second protein sequence as a node and construct edges using the first attention matrix to obtain graph structure data. Then, based on a triangular update mechanism, the graph structure data is iteratively updated, for example, three times, to obtain the updated first language model encoding and the first attention matrix. The updated first language model encoding is then determined as the third feature.

[0133] S608: The root mean square error calculation module 14 in the structure detection model 01 obtains the root mean square error information of each amino acid backbone atom in the first antibody sequence.

[0134] In some embodiments, the root mean square error calculation module 14 can calculate the root mean square error for each amino acid backbone atom (C) in the first antibody sequence. α C β The root mean square error (pRMSD) of the main chain atoms of amino acids (N, O) is calculated. α Root mean square error, main chain atom C β The root mean square error (RMSE) of the amino acid in the first antibody sequence is calculated as follows: RMSE of the main chain atom N and RMSE of the main chain atom O. Therefore, the RMSE information for each amino acid in the main chain of the first antibody sequence is obtained as N×4 units of data.

[0135] S609: The feature fusion module 14 in the structure detection model 01 converts the root mean square error information into scaling factors, performs dimensional transformation on the third feature according to the scaling factors to obtain the fourth feature, and obtains the fused feature based on the fourth feature and the first feature.

[0136] The scaling factor has a one-dimensional dimension, and the fourth feature has the same dimension as the first feature.

[0137] In some embodiments, the feature fusion module 14 can input the N×4-dimensional root mean square error information pRMSD of the first antibody sequence into the first linear layer, FC(4,32), of the confidence conversion submodule 141 within the feature fusion module 14, so that the last activation layer (Sigmoid layer) of the confidence conversion submodule 141 outputs a scaling factor. At this time, the N×4-dimensional root mean square error information pRMSD passes through the FC(4,32) layer, Sigmoid layer, FC(32,32) layer, Sigmoid layer, FC(32,1) layer, and Sigmoid layer in the confidence conversion submodule 141, causing the confidence conversion submodule 141 to output a scaling factor.

[0138] In some embodiments, the feature fusion module 14 inputs the scaling factor and the third feature into the aggregation layer of the feature conversion submodule 142, and inputs the first feature into the summation layer (such as the Add layer) of the feature conversion submodule 142, so that the last activation layer (Sigmoid) in the feature conversion submodule 142 outputs the fused feature. That is, the aforementioned scaling factor and the third feature of the first antibody sequence, i.e., tensor S2, are input into the Multiplate layer of the feature conversion submodule 142, and converted into a fourth feature through the FC(64,384) layer and the ReLU layer. The dimension of the fourth feature is N×384, just like the dimension of the first feature. Then, the tensor S1 corresponding to the first feature of the first antibody sequence is input into the Add layer of the feature conversion submodule 142, so that the tensor corresponding to the fourth feature is summed with the tensor S1 corresponding to the first feature in the Add layer. Then, the tensor obtained by summation is converted again through the FC(64,384) layer and the Sigmoid layer and the tensor S corresponding to the fused feature is output.

[0139] In some embodiments, the feature fusion module 14 sums the fourth feature and the first feature to obtain a fused feature.

[0140] Specifically, the feature transformation submodule 142 in the feature fusion module 14 inputs the tensor S1 corresponding to the first feature of the first antibody sequence into the Add layer of the feature transformation submodule 142, so that the tensor corresponding to the fourth feature is summed with the tensor S1 corresponding to the first feature in the Add layer. Then, the tensor obtained by summation is transformed through the FC(64,384) layer and the Sigmoid layer and the tensor S corresponding to the fused feature is output.

[0141] S610: The feature update module 15 in the structure detection model 01 updates the fused features and the second feature based on the attention mechanism.

[0142] In some embodiments, the feature update module 15 may update the fused features and the second feature based on an attention mechanism such as a triangular attention mechanism.

[0143] S611: The structure detection module 17 in the structure detection model 01 detects the mean and variance of the translation parameters of amino acids in the first antibody sequence based on the updated second feature and the updated fusion feature, and also detects the mean and variance of the rotation parameters of amino acids in the first antibody sequence.

[0144] In some embodiments, during the process of detecting the mean and variance of the translation parameters of each amino acid, the structure detection module 17 can identify amino acids at different positions based on the positional encoding of each amino acid in the first antibody sequence. In this case, the structure detection module 17 can employ a Transformer architecture.

[0145] S612: The structure detection module 17 in the structure detection model 01 adopts a reparameterization method to determine the target values ​​of the translation parameters and rotation parameters of each amino acid in the first antibody sequence based on the mean and variance of the translation parameters and the mean and variance of the rotation parameters.

[0146] As an example, the structure detection module 17 can use the above formula (1) to determine the target values ​​of the translation parameters and rotation parameters of each amino acid in the first antibody sequence, so as to determine the three-dimensional structure of the first antibody sequence.

[0147] In this context, the translation parameter represents the position of the amino acid, the rotation parameter represents the angle of the amino acid, and the three-dimensional structure includes both the position and angle of the amino acids. Therefore, the position and angle of each amino acid in the first antibody sequence are used to characterize the three-dimensional structure of the first antibody sequence.

[0148] Thus, in the antibody sequence detection method provided in this application, the pre-trained structural antibody detection model can quickly search protein sequence databases using the search tool MMseqs2, which is beneficial for quickly obtaining homologous sequences of the first antibody sequence. Furthermore, this structural antibody detection model can use transfer learning to extract multidimensional features (the first feature) of the homologous sequence from AlphaFold2, and the interaction information between pairs of amino acids in the sequence (the second feature). Simultaneously, this structural detection model can also use the pre-trained protein language model AntiBERTy through transfer learning to extract language model features of the first antibody sequence, and further use feature encoding modules in IgFold, such as the Evoformer module, to quickly generate multidimensional feature information (the third feature). Therefore, the pre-trained structural antibody detection model of this application can integrate homologous sequence features and language model features of the first antibody sequence for antibody structure detection, thereby improving the overall structural detection accuracy of the antibody sequence.

[0149] In some embodiments, refer to Figure 7 The diagram shown is a flowchart illustrating the structural detection of an antibody sequence according to an embodiment of this application. Specifically, Figure 7The antibody structure detection process of the illustrated structure detection model 01, AbFold, includes the following steps: On the input side, AbFold receives an input antibody sequence composed of heavy and light chains (e.g., the first antibody sequence). It uses MMseqs2 to search the Uniref90 protein sequence database to obtain homologous sequences of the input antibody sequence and generates a multiple sequence alignment (MSA) representation. Next, AbFold uses transfer learning to extract multidimensional features (i.e., the first feature, such as tensor S1) and the interaction information between pairs of amino acids in the sequence (i.e., the second feature, such as tensor z) from AlphaFold2. Simultaneously, the input antibody sequence is fed into the pre-trained protein language model AntiBERTy to extract the sequence's language model features, which are then further input into the IgFold feature encoding module to generate multidimensional feature information (i.e., the third feature, such as tensor S2). Then, AbFold's deep neural network fuses the multidimensional feature information represented by tensors S1 and S2 to obtain a fused feature (e.g., tensor S). Through a triangular self-attention mechanism, the tensor S corresponding to the fused feature and the tensor z corresponding to the second feature are mutually updated. Finally, the updated tensors S and z are input into the structure prediction module, which uses an invariant point attention mechanism and a variational autoencoder to dynamically predict the antibody's three-dimensional structure. Thus, AbFold starts with the amino acid sequence of the input antibody sequence, extracting not only information from homologous sequences but also sequence features using a pre-trained antibody language model. By fusing homologous sequence features and language model features through a deep neural network, AbFold dynamically models the antibody's three-dimensional structure, which helps improve the accuracy of antibody three-dimensional structure detection.

[0150] Next, combined Figures 8 to 10 This application provides a detailed description of the model training phase of the deep learning-based antibody structure detection model.

[0151] Reference Figure 8 The diagram shown is a schematic of the training process of an antibody structure detection model provided in this application. The subject executing this process can be an electronic device or a device in the electronic device.

[0152] Specifically, Figure 8 The training process shown includes the following steps:

[0153] S801: Obtain the training dataset, which includes at least one training data set, each training data set including an unlabeled antibody sequence and a labeled antibody sequence.

[0154] In some embodiments, the labeled and unlabeled antibody sequences in a training data group in the training dataset can be randomly selected or pre-divided from the training dataset, and this application does not specifically limit this.

[0155] As an example, to train and evaluate the performance of the structure detection model 01, namely the AbFold model, this application screened 1872 high-quality antibody structures from the structure antibody database SAbDab, with a maximum sequence similarity of 99% and a minimum resolution of 4.0 Å. To facilitate effective comparison with existing antibody structure prediction models (such as IgFold), this application divides the dataset corresponding to the SAbDab database into training and testing sets based on the storage date of the antibody structures, ensuring that the data in the testing set has not been used in the training of any model. For example, this application uses antibody sequence data stored in the SAbDab database before July 1, 2021 for model training, and uses antibody sequence data stored after that date for model testing. Therefore, the training set corresponding to the SAbDab database contains 1719 labeled antibody sequences, and the testing set contains 153 labeled antibody sequences.

[0156] Among them, E It is a unit of length, commonly used to represent the distance between atoms or molecules. For example, the root mean square error (RMSD) is... This indicates that the difference between the predicted and actual values ​​is approximately 1.37 angstroms on average.

[0157] As an example, this application randomly selected 20,000 pairs of heavy and light chain sequences from the antibody sequence database OAS as input data for unsupervised training of the model. In this case, during model training, the training dataset consists of half structured supervised data (such as labeled antibody sequences from the SAbDab database) and half unsupervised data (such as unlabeled antibody sequences from the OAS database).

[0158] In some embodiments, the training dataset constructed in this application includes 1872 antibody sequences with native structures and 20,000 unstructured antibody sequences. Based on the storage date of the antibody sequences, 1719 antibody sequences stored before July 1, 2021, are used as the training set, and 153 antibody sequences stored thereafter are used as the test set. All 20,000 unstructured antibody sequences are used for self-supervised learning training.

[0159] It is understandable that the number of antibody structures currently determined using experimental methods such as NMR, X-ray crystallography, and cryo-electron microscopy is limited. According to statistics from the SabDab database, only a few thousand antibody structures have been experimentally determined. This limited amount of labeled data has a significant negative impact on the training performance of deep learning models. Therefore, this application employs a teacher-student self-supervised learning method. By introducing a large number of unlabeled antibody sequences, a self-supervised training approach is used, combining them with labeled antibody sequences for model training, thereby improving the model's predictive ability.

[0160] S802: Based on the training dataset, a teacher-student self-supervised learning method is used to train the structure detection model to be trained, and the trained structure detection model is obtained.

[0161] In some embodiments, during the model training phase based on teacher-student self-supervised learning, the structure detection model to be trained can be divided into two branches: a student model and a teacher model. As an example, the student model and the teacher model have the same structure, and thus, the student model and the teacher model are used for model training based on teacher-student self-supervised learning.

[0162] In some embodiments, during training, each mini-batch consists of half labeled antibody sequences and half unlabeled antibody sequences. The labeled antibody data uses the native structure (i.e., the real label) as a supervision signal, while the unlabeled training antibody sequences use the three-dimensional structure predicted by the teacher model as a supervision signal.

[0163] In some implementations, the model training in this application lasted for 300 epochs, with a minimum batch size of 8, and the AdamW optimizer was used to accelerate convergence. The base learning rate was set to e^(-ε / ε). -4 The learning rate is linearly increased to this value in the first 1000 iterations, then remains constant from 1000 to 100000 iterations. After 100000 iterations, the learning rate is decayed by 5% every 1000 iterations to improve training stability and performance. To avoid gradient explosion, the gradients of all model parameters are pruned using the L2 norm with a pruning threshold of 0.1. Furthermore, to ensure equal sequence lengths within mini-batches and increase data diversity during training, the length of a single strand (e.g., heavy or light strand) is fixed at a predetermined value, such as 96. For antibody sequences with a single strand length exceeding the predetermined value of 96, a contiguous subsequence with a median length of 96 is randomly selected; for antibody sequences with a median length less than the predetermined value of 96, the end of each single strand is padded with zeros.

[0164] like Figure 9 The diagram shown is a schematic of a network architecture for model training based on a teacher-student self-supervised learning method provided in an embodiment of this application. Figure 9 The architecture diagram shows a student model (91) and a teacher model (92). Student model 91 and teacher model 92 have identical structures; the input is an antibody sequence, and the output is the three-dimensional structure of the antibody. For example, for the same input antibody sequence, the three-dimensional structure output by student model 91 is denoted as T1, and the three-dimensional structure output by teacher model 92 is denoted as T2. Figure 9As shown, during training, student model 91 and teacher model 92 predict the structure of the unlabeled antibody sequence, and the prediction results of teacher model 92 are used as supervision signals to guide the learning of student model 91. Specifically, pseudo-labels are first generated based on the prediction results of the teacher model, and then the corresponding loss, such as Loss(T1,T2), which is the loss between three-dimensional structures T1 and T2, is calculated. The parameters of student model 91 are then updated through gradient backpropagation. The parameters of teacher model 92 are not updated through direct gradient backpropagation, but rather dynamically through the exponential moving average (ema) of the parameters of student model 91. At this point, teacher model 92 can stop updating its parameters using stochastic gradient descent (sg). The exponential moving average is a trend-following indicator, specifically a moving average with exponentially decreasing weights. This synchronous update strategy dynamically constructs teacher model 92 during training, avoiding reliance on a fixed teacher model, thus enabling teacher model 92 to continuously optimize itself. In this way, a small amount of labeled structural information can be effectively propagated to a large number of unlabeled antibody sequences, thereby improving the model's ability to predict the structure of antibody sequences.

[0165] Thus, the teacher-student self-supervised learning strategy in this application significantly improves the quality of features learned by the model and can better address the problem of insufficient labeled data in antibody structure prediction tasks.

[0166] Reference Figure 10 The diagram illustrates a training process for a structure detection model based on a training data set, as provided in an embodiment of this application. The execution entity of this process can be an electronic device or a device within that electronic device. For example, the training data set includes an unlabeled first training antibody sequence and a labeled second training antibody sequence.

[0167] Specifically, Figure 10 The training process shown includes the following steps:

[0168] S1001: Obtain the teacher model and student model corresponding to the structure detection model to be trained.

[0169] S1002: Input the first unlabeled training antibody sequence from the training data set into the student model to obtain the first detection result, and input the first training antibody sequence into the teacher model to obtain the second detection result.

[0170] S1003: Input the second training antibody sequence with the label in the training data set into the student model to obtain the third detection result.

[0171] In some embodiments, a second training antibody sequence can be input into the teacher model to obtain a fourth detection result.

[0172] As an example, the inputs and outputs of the teacher model and the student model can be illustrated by the following formulas (2) and (3).

[0173]

[0174] Where x1 and x2 represent the unlabeled and labeled antibody sequences in the minimum batch of data during the training process, respectively; student and teacher represent the student model and teacher model in self-supervised learning, respectively; and T represents the antibody structure predicted by the model (i.e., the three-dimensional structure of the antibody sequence).

[0175] As an example, This represents the 3D structure obtained by the student model performing structure detection on the unlabeled antibody sequence x1, for example, when x1 is the first training antibody sequence. This is the first test result.

[0176] This represents the three-dimensional structure obtained by the teacher model (teacher) through structure detection of the unlabeled antibody sequence x1, for example, when x1 is the first training antibody sequence. This is the second test result.

[0177] This represents the three-dimensional structure obtained by the student model using structure detection on the labeled antibody sequence x2, for example, when x2 is the second training antibody sequence. This is the third test result.

[0178] This represents the three-dimensional structure obtained by the teacher model (teacher) through structure detection of the labeled antibody sequence x2, for example, when x2 is the second training antibody sequence. This is the fourth test result.

[0179] S1004: Calculate the first loss between the student model and the teacher model based on the first and second detection results.

[0180] S1005: Calculate the second loss between the student model and the teacher model based on the third detection result and the true label of the second training antibody sequence.

[0181] In some embodiments, the loss function between the student model and the teacher model can be based on AlphaFold2. and Functions and IgFold The function is determined.

[0182] As an example, the loss function between the student model and the teacher model can be determined by formula (4).

[0183]

[0184] Among them, the loss function Including AlphaFold2 and The functions are: the former measures the similarity between the predicted and actual structures of the antibody sequence, while the latter is used to learn the geometric constraints between residues in the antibody sequence to avoid inter-atom conflicts. Additional... The loss function is derived from IgFold and is used to calculate the L1 norm of the distance between the i-th amino acid and the (i+1)-th and (i+2)-th amino acids in the antibody sequence, where i is a positive integer.

[0185] It is understandable that the first loss and the second loss mentioned above can both be calculated based on the above formula (4).

[0186] For example, the first loss between the student model and the teacher model Formula (4) can be used for calculation. Wherein, That is, the three-dimensional structure predicted by the model student for the unlabeled antibody sequence (such as the first detection result mentioned above) is used as T in formula (4). pred , That is, the pseudo-label predicted by the teacher model for the antibody sequence (such as the second detection structure mentioned above) is used as T in formula (4). label .

[0187] For example, the second loss between the student model and the teacher model. Formula (4) can be used for calculation. That is, the three-dimensional structure predicted by the student model for the antibody sequence corresponding to the labeled data (such as the third detection result) is used as T in formula (4). pred T gt That is, the real label of the antibody sequence (i.e., the real label corresponding to the second training antibody sequence) is used as T in formula (4). label .

[0188] S1006: Calculate the sum of the first loss and the second loss to obtain the third loss of the structural detection model.

[0189] As an example, the loss function of the structure detection model can be determined by formula (5).

[0190]

[0191] Furthermore, based on the loss function in formula (5), the overall loss of the structure detection model during training can be calculated, and the parameters of the student model and the teacher model can be updated according to the loss.

[0192] S1007: Based on the third loss, update the parameters of the student model using the backpropagation gradient method.

[0193] S1008: Update the parameters of the teacher model by applying an exponential moving average to the updated parameters of the student model.

[0194] In some embodiments, the parameters of the teacher model can be updated by an exponential moving average of the parameters of the student model, which can be performed by the following formula (6).

[0195] Θ teacher =(1-α)Θ teacher +αΘ student (6)

[0196] Where, Θ teacher and Θ student These represent the parameters of the teacher model and the student model, respectively, with α = 0.01 being the weighting coefficient.

[0197] Similarly, this application can apply the following to each training data set in the training dataset: Figure 10 The process shown is used to update the parameters of the structural detection model.

[0198] In some embodiments, this application may determine the trained student model as the trained structure detection model.

[0199] Thus, this application can calculate the first loss corresponding to unlabeled antibody sequences and the second loss corresponding to labeled antibody sequences through a teacher-student self-supervised model training method, and then obtain the overall third loss of the model based on the first and second losses. This third loss simultaneously considers the impact of both unlabeled and labeled antibody sequences on the model loss, which is beneficial for improving the structure detection accuracy of the structure detection model trained based on the third loss.

[0200] Next, we will conduct experiments on the structure detection process based on the pre-trained structure detection model 01, namely AbFold, in the implementation of this application, and analyze the detection performance of the structure detection model 01 based on the experimental results.

[0201] In some embodiments, to verify the effectiveness of the model, this application collected 153 double-stranded antibody sequences with high-resolution structures as a test set and compared them with existing state-of-the-art antibody structure prediction models. All 153 antibody sequences were stored in publicly available databases after July 1, 2021, ensuring that these sequences were not included in the training set of any of the models being compared, thus guaranteeing the fairness of the test. It is understood that the most challenging part of antibody structure prediction is modeling CDR loops. This application can use the Chothia numbering scheme to determine each CDR loop in the test set. The Chothia numbering scheme is used to identify amino acids at specific positions on the antibody. This scheme is based on the crystal structure of the antibody's variable region, defines the loop structure forming the CDR (complementarity-determining region), and corrects the position numbers of the CDR-L1 and CDR-H1 loop insertion points to better fit their topological positions.

[0202] In some embodiments, this application compares the AbFold-based structure detection method of this application with other antibody structure prediction methods such as AlphaFold2-Multimer, IgFold, ImmuneBuilder, DeepAb, and EquiFold on 153 test antibodies. For example, Table 2 shows the prediction accuracy of these models on the test set.

[0203] Table 1:

[0204]

[0205]

[0206] Understandably, this application uses the root mean square error (RMSD) to evaluate the quality of antibody structure prediction. For example... Figure 2 The experimental results show that AbFold has the lowest error in predicting the structure of CDR rings, and its predicted structure is closer to the natural state structure. Among the six CDR rings, AbFold achieved the lowest average RMSD difference (i.e., ...) on five CDR rings except for the H1 ring. ), and the average RMSD of the prediction results for the H1 ring is only slightly lower than the optimal result (i.e. )Difference In the most difficult-to-model CDR H3 ring, AbFold's average RMSD is: With AlphaFold2-Multimer IgFold ImmuneBuilder DeepAb and EquiFold In comparison, AbFold reduced the average RMSD by 13%, 24%, 5%, 30%, and 12%, respectively. In the relatively more structurally stable skeleton regions (FRH, FRL), all methods were able to predict the corresponding structures with high accuracy (FRH: FRL: However, in the CDR region, where structural variability is high, AbFold's RMSD is significantly lower than other methods, including AlphaFold2-Multimer, indicating that AbFold has a stronger modeling ability for dynamic structures. This is understandable because deterministic structural prediction methods, such as AlphaFold2-Multimer, struggle to model different conformations of the same sequence. During training, these methods tend to learn the mean structure of multiple conformations in the training set. AbFold, based on a variational autoencoder, introduces noise and variance to give the predicted structure a degree of randomness, thus avoiding the model forcibly fitting the mean of multiple conformations.

[0207] Reference Figures 11A to 11C The diagram shows the test results of various models on the test set in the embodiments of this application. These test results comprehensively compare the performance of six antibody structure prediction methods, namely AbFold, AlphaFold2-Multimer, IgFold, ImmuneBuilder, DeepAb, and EquiFold, on the test set.

[0208] like Figure 11A The figure shown is a box plot of the prediction accuracy of various antibody structure prediction methods provided in the embodiments of this application. Figure 11A As shown, compared with other methods, AbFold's prediction results show lower values ​​in the lower quartile, median, upper quartile, and interquartile range, indicating that the error distribution of AbFold's prediction results is more concentrated, the differences between antibodies are smaller, and the model prediction results are more stable.

[0209] Understandably, in order to gain a deeper understanding of how AbFold optimizes predicted structures and achieves higher accuracy in CDR H3, this application analyzed two examples from the test set, whose protein data bank (PDB) identifiers (IDs) are 7N0A and 7PHW, respectively. Figure 11B The diagram shows a visualization of the natural structure of 7N0A, the AbFold predicted structure, and the IgFold predicted structure. Figure 11B In the diagram, the natural structure of 7N0A is represented by the black portion (experimental structure), the AbFold predicted structure is represented by the gray portion (predicted structure), and the IgFold predicted structure is represented by the white portion (predicted structure). Clearly, Figure 11BThe AbFold predicted structure shown is closer to the natural structure of 7N0A.

[0210] Figure 11C A visual representation of the native structure of antibody PHW, the predicted structure of AbFold, and the predicted structure of IgFold is shown. Similarly, Figure 11C In the diagram, the 7PHW native structure is represented by the black portion (expected structure), the AbFold predicted structure by the gray portion (predicted structure), and the IgFold predicted structure by the white portion (predicted structure). Clearly, Figure 11B The AbFold predicted structure shown is closer to the natural structure of 7PHW.

[0211] Specifically, this application compares the structures predicted by IgFold and AbFold with the natural structures. For 7N0A (e.g. Figure 11B As shown in the figure, the CDR H3 ring of this antibody consists of seven residues.

[0212] IgFold's average This indicates that IgFold failed to accurately predict the conformation of the CDR-H3 ring, exhibiting low confidence in this region. In contrast, AbFold's average... This indicates that AbFold predicted a relatively more accurate CDR-H3 conformation. AbFold optimized the orientation of the CDR-H3 structure, making it closer to the natural structure, thereby significantly reducing RMSD. For 7PHW (such as... Figure 11C As shown), the length of the CDR-H3 ring is 18, more than twice the length of the previous example 7N0A. The AbFold-predicted conformation has an average RMSD. The average RMSD of the conformation is significantly better than that predicted by IgFold. This indicates that AbFold can accurately predict not only antibodies with short CDR-H3 rings, but also antibody structures with longer CDR-H3 rings.

[0213] In some embodiments, this application can analyze the impact of antibody sequence fusion features (also known as fusion information) on the accuracy of CDR-H3 structure prediction. It is understood that there are significant differences between the CDR-H3 structures predicted by AlphaFold2-Multimer based on homologous sequences and IgFold based on protein language models. To explore the impact of information fusion on antibody CDR-H3 structure prediction, this application separately compared the antibody CDR-H3 prediction results of AbFold with AlphaFold2-Multimer and IgFold on the test set, using the root mean square error (RMSD) as the evaluation metric.

[0214] Reference Figure 12AA scatter plot comparing AbFold and AlphaFold2-Multimer is shown below. Figure 12B A scatter plot comparing AbFold and IgFold is shown.

[0215] like Figure 12A and Figure 12B As shown, AbFold, which integrates homology sequence features and protein language model features, significantly outperforms AlphaFold2-Multimer, which uses only homology sequence features, and IgFold, which uses only protein language model features. Among the 153 tested antibody sequences, AbFold predicted structures superior to AlphaFold2-Multimer and IgFold in 88 and 114 sequences, respectively, representing 57.5% and 74.5%.

[0216] Reference Figure 12C This diagram illustrates a comparison between the CDR-H3 structures predicted by the three models AbFold, AlphaFold2, and IgFold and the natural structure. Figure 12C As shown, from left to right, these are visualizations of the CDR-H3 structures predicted by AlphaFold2 (such as AlphaFold2-Multimer), IgFold, and AbFold. Rows 1 to 3, from top to bottom, represent antibodies 7bh8, 7jwg, and 7aj6. The results indicate that AbFold's prediction is closer to the native structure. This demonstrates that AbFold can capture more precise structural information from homologous sequence features and language model features through its feature fusion module, thereby more accurately predicting the structure of the antibody's CDR-H3 loop.

[0217] In some embodiments, the experimental results of ablation experiments on the network structure of the AbFold model in this application are analyzed. Table 2 shows a schematic diagram of the ablation experimental results of the AbFold model network structure.

[0218] Table 2:

[0219]

[0220] Specifically, Table 2 shows the RMSD (i.e. H3 RMSD) and global distance test total score (GDT-TS) values ​​of AbFold with the first ablation model (AbFold-w / o-pRMSD), the second ablation model (AbFold-w / o-VAE), and the third ablation model (AbFold-Attn) in CDR-H3.

[0221] AbFold-w / o-pRMSD represents a variant of the AbFold model that removes the root mean square error calculation module (pRMSD) or related techniques.

[0222] AbFold-w / o-VAE represents a variant of AbFold model that does not employ a variational autoencoder (VAE) as a model component or data preprocessing step when predicting protein structures.

[0223] AbFold-Attn introduces an attention mechanism into the AbFold model, which enables the model to more accurately capture key information in the amino acid sequence.

[0224] Furthermore, GDT-TS represents the proportion of atomic position matches between predicted and actual protein structures at different distance thresholds. GDT-TS integrates the degree of matching at different accuracy levels and is consistent for proteins of different sizes.

[0225] Table 2 shows that AbFold's overall structure and CDR-H3 region prediction accuracy are significantly higher than those of the other ablation models. Compared to AbFold-w / o-pRMSD, AbFold-Attn, and AbFold-w / o-VAE, AbFold's GDT_TS improved by 6.8%, 3.3%, and 5.6%, respectively, while its RMSD prediction for CDR-H3 decreased by 23.8%, 15.8%, and 19.9%, respectively. Compared to AbFold-w / o-pRMSD, AbFold improved GDT_TS by 6.8% and decreased H3 RMSD by 23.8%, indicating that pRMSD plays a guiding role in the fusion of homologous sequence features and language model features, thus obtaining more accurate structural information. Compared to AbFold-w / o-VAE, AbFold improves GDT_TS by 5.6% and reduces H3 RMSD by 19.9%, indicating that variational autoencoders not only play a crucial role in modeling the highly variable CDR H3 ring, but also have a positive impact on the relatively stable skeleton region. Compared to AbFold-Attn, AbFold improves GDT_TS by 3.3% and reduces H3 RMSD by 15.8%. The prediction accuracy of AbFold-Attn, based on the attention mechanism, is slightly lower than that of AbFold, which is based on fully connected layers. This may be because there is less training data, making it difficult to fully train a powerful but data-intensive attention network.

[0226] In summary, the CDR-H3 loop of an antibody plays a crucial role in the binding process between the antibody and the antigen. With the continuous advancement of protein structure prediction methods, the prediction accuracy of antibody backbone structures has approached experimental levels. However, the prediction of the CDR-H3 loop structure remains a challenge and a hot topic in antibody research. This application proposes an antibody structure detection model, AbFold, based on transfer learning. First, homologous sequences and their multiple sequence alignments for the antibody sequence can be obtained by searching the Uniref90 sequence database using MMseqs2, and language model features are extracted using the AntiBERTy antibody language model. Then, transfer learning is used to extract homologous sequence features and language model features from AlphaFold-Multimer and IgFold, respectively. Next, the two types of features are fused through a fully connected layer feature fusion module to obtain the final fused features, which are then input into a variational autoencoder-based neural network (i.e., the structure detection module) for dynamic modeling of the antibody structure. Experimental results on 153 test antibodies show that AbFold outperforms other antibody structure prediction methods in terms of CDR-H3 loop structure prediction accuracy. By combining homologous sequence features and language model features, AbFold can capture more comprehensive structural information, making antibody structure prediction more accurate. Furthermore, based on antigen sequence and structural information bound to the antibody sequence, an antibody structure prediction method can be generated under given antigen conditions, further improving the accuracy and practicality of antibody structure prediction.

[0227] Next, combined Figure 13 The structure of the electronic device for performing the antibody sequence detection method of this application is described in detail.

[0228] like Figure 13 As shown, the electronic device 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, buttons 101, and a display screen 102, etc.

[0229] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 10. In other embodiments of this application, the electronic device 10 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0230] Processor 110 may include one or more processing units, such as processing modules or circuits of a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), microprocessor (MCU), artificial intelligence (AI) processor, or field-programmable gate array (FPGA). Different processing units may be independent devices or integrated into one or more processors. Processor 110 may include storage units for storing instructions and data. In some embodiments, the storage unit in processor 110 is a cache memory 180. For example, processor 110 executes... Figure 5 , Figure 6 , Figure 8 and Figure 10 The method in the middle.

[0231] The power module 140 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component manages the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects to the power supply and the processor 110. The power management module receives input from the power supply and / or the charging management module to supply power to the processor 110, the display 102, the camera 170, and the wireless communication module 120, etc.

[0232] The mobile communication module 130 may include, but is not limited to, antennas, power amplifiers, filters, and low-noise amplifiers (LNAs). The mobile communication module 130 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the electronic device 10. The mobile communication module 130 can receive electromagnetic waves via the antenna, filter and amplify the received electromagnetic waves, and then transmit them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device.

[0233] The wireless communication module 120 may include an antenna, which enables the transmission and reception of electromagnetic waves. The wireless communication module 120 can provide solutions for wireless communication applications on the electronic device 10, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The electronic device 10 can communicate with networks and other devices through wireless communication technologies.

[0234] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the electronic device 10 may also be located in the same module.

[0235] The display screen 102 is used to display human-computer interaction interfaces, images, videos, etc.

[0236] The sensor module 190 may include proximity sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0237] The audio module 150 is used to convert digital audio information into analog audio signals for output, or to convert analog audio input into digital audio signals. The audio module 150 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 150 may be located in the processor 110, or some functional modules of the audio module 150 may be located in the processor 110. In some embodiments, the audio module 150 may include a speaker, a handset, a microphone, and a headphone jack.

[0238] Camera 170 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to image signal processing (ISP) to be converted into a digital image signal. Electronic device 10 can implement the shooting function through ISP, camera 170, video codec, graphics processing unit (GPU), display screen 102, and application processor. For example, in this application, a mobile phone can acquire images through ISP, camera 170, etc., and perform the post-processing of these images according to this application. Figures 5-8 The relevant image processing workflow.

[0239] Interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a microSD card, to expand the storage capacity of electronic device 10. The external memory card communicates with processor 110 through the external memory interface to perform data storage. The USB interface is used for communication between electronic device 10 and other electronic devices. The SIM card interface is used to communicate with the SIM card installed in electronic device 1010, for example, to read or write phone numbers stored in the SIM card.

[0240] In some embodiments, the electronic device 10 further includes buttons 101, a motor, and indicators. The buttons 101 may include volume buttons, a power button, etc. The motor is used to generate a vibration effect in the electronic device 10, for example, vibrating when the user's electronic device 10 is called to prompt the user to answer the call. The indicators may include laser indicators, radio frequency indicators, LED indicators, etc.

[0241] In some embodiments, this application provides a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the image processing method described above.

[0242] In some embodiments, this application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, one of the processors of the electronic device, for performing the image processing method described above.

[0243] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0244] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0245] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0246] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other propagation signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0247] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0248] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0249] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0250] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. An antibody sequence detection method, characterized in that, The method includes: Obtain a first antibody sequence, the first antibody sequence comprising a pair of heavy chains and light chains; Obtain a first feature, a second feature, and a third feature of the first antibody sequence, wherein the first feature is used to indicate evolutionary constraint information between the first antibody sequence and homologous sequences, the second feature is used to indicate the relationship between amino acid pairs in the first antibody sequence, and the third feature is used to indicate the dependency relationship between amino acids in the first antibody sequence. The first feature and the third feature are fused to obtain the fusion feature of the first antibody sequence; The three-dimensional structure of the first antibody sequence is detected based on the second feature and the fusion feature. The third feature is determined in the following manner: The heavy and light chains in the first antibody sequence are linked together by vacancies to form a single sequence, thus obtaining the second protein sequence. The second protein sequence is input into a pre-trained protein language model, which outputs a first language model encoding and a first attention matrix. The third feature is obtained by updating the first language model encoding and the first attention matrix based on the attention mechanism; The process of fusing the first feature and the third feature to obtain the fusion feature of the first antibody sequence includes: The root mean square error information of the first antibody sequence is converted into a scaling factor. The root mean square error information includes the root mean square error of each amino acid backbone atom in the first antibody sequence. The scaling factor has a one-dimensional dimension. The third feature is transformed according to the scaling factor to obtain the fourth feature, wherein the fourth feature has the same dimension as the first feature; The fourth feature and the first feature are summed to obtain the fused feature.

2. The method according to claim 1, characterized in that, The first feature and the second feature are determined in the following manner: Search for homologous sequences of the first antibody sequence; Obtain a multiple sequence comparison representation between the first antibody sequence and the homologous sequence; The heavy chain and light chain in the first antibody sequence are linked together to form a single sequence, thus obtaining the first protein sequence. The different amino acids in the first protein sequence are paired to obtain the paired representation of the first protein sequence; The multiple sequence contrast representation and the pairing representation are encoded to obtain the first feature and the second feature, respectively.

3. The method according to claim 1, characterized in that, The step of detecting the three-dimensional structure of the first antibody sequence based on the second feature and the fusion feature includes: The fused features and the second feature are updated based on an attention mechanism; The three-dimensional structure of the first antibody sequence is detected based on the updated second feature and the updated fusion feature.

4. The method according to claim 3, characterized in that, The step of detecting the three-dimensional structure of the first antibody sequence based on the updated second feature and the updated fusion feature includes: Based on the updated second feature and the updated fusion feature, the mean and variance of the translation parameters of amino acids in the first antibody sequence are detected, and the mean and variance of the rotation parameters of amino acids in the first antibody sequence are also detected. Using a reparameterization method, the target values ​​of the translation parameters and rotation parameters for each amino acid in the first antibody sequence are determined based on the mean and variance of the translation parameters and the mean and variance of the rotation parameters. The translation parameters represent the positions of the amino acids, and the rotation parameters represent the angles of the amino acids. The three-dimensional structure includes both the positions and angles of the amino acids.

5. The method according to claim 4, characterized in that, The positions of different amino acids in the first antibody sequence encode different values.

6. The method according to claim 1, characterized in that, The three-dimensional structure is determined based on a structure detection model, which includes a feature fusion module. The feature fusion module includes a confidence transformation submodule and a feature transformation submodule. The confidence transformation submodule includes three sets of network layers connected in sequence. Each set of network layers includes a linear layer and an activation layer connected in sequence. The feature transformation submodule includes an aggregation layer, a linear layer, an activation layer, a summation layer, a linear layer, and an activation layer connected in sequence. The fused features are determined based on the feature fusion module.

7. The method according to claim 6, characterized in that, The process by which the feature fusion module determines the fused feature includes: The root mean square error information is input into the first linear layer of the confidence transformation submodule, so that the last active layer of the confidence transformation submodule outputs the scaling factor; The scaling factor and the third feature are input into the aggregation layer of the feature transformation submodule, and the first feature is input into the summation layer of the feature transformation submodule, so that the last activation layer in the feature transformation submodule outputs the fused feature.

8. The method according to claim 6, characterized in that, The three-dimensional structure is determined based on the structure detection module in the structure detection model, and the structure detection module includes a variable autoencoder.

9. The method according to any one of claims 1 to 8, characterized in that, The first feature is a tensor of shape N×384, the third feature is a tensor of shape N×64, and the second feature is a tensor of shape N×N×128, where N is the total length of amino acids in the heavy and light chains of the first antibody sequence.

10. The method according to any one of claims 6 to 8, characterized in that, The method further includes: The structure detection model to be trained is trained using a teacher-student self-supervised learning method to obtain the trained structure detection model.

11. The method according to claim 10, characterized in that, The method of training the structure detection model to be trained using a teacher-student self-supervised learning approach to obtain the trained structure detection model includes: Based on the training dataset, a structure detection model to be trained is trained using a teacher-student self-supervised learning method to obtain the trained structure detection model. The training dataset includes at least one training data set, which includes an unlabeled antibody sequence and a labeled antibody sequence.

12. The method according to claim 11, characterized in that, The training process of the structure detection model includes: Obtain the teacher and student models corresponding to the structure detection model to be trained; The first training antibody sequence without a label in the training data set is input into the student model to obtain a first detection result, and the first training antibody sequence is input into the teacher model to obtain a second detection result; Calculate the first loss based on the first detection result and the second detection result; The second training antibody sequence with a label in the training data set is input into the student model to obtain the third detection result; The second loss is calculated based on the third detection result and the true label of the second training antibody sequence; Calculate the sum of the first loss and the second loss to obtain the third loss; Based on the third loss, the parameters of the student model are updated using the backpropagation gradient method; The parameters of the teacher model are updated by applying an exponential moving average to the updated parameters of the student model.

13. A readable medium, characterized in that, The readable medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the antibody sequence detection method as described in any one of claims 1 to 12.

14. An electronic device, characterized in that, include: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, one of the processors of the electronic device, for performing the antibody sequence detection method as described in any one of claims 1 to 12.