Method, apparatus, device, and storage medium for determining antigen-antibody binding sites

By binding the sequence and structural information of the antibody, multi-sequence alignment and de novo folding technology are used to solve the accuracy of antigen antibody binding site prediction, and the rapid support for antibody drug development is achieved.

CN115116543BActive Publication Date: 2025-08-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210407121.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-08-05
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict antigen antibody binding sites under different antibody sequence lengths, especially for template-free antibodies. The existing methods cannot effectively utilize structural information, resulting in low prediction accuracy of binding site.

Method used

Through the sequence information and structural information of the antibody, the structural similarity and sequence similarity of the antibody to be predicted and the known antibody are determined based on the key regions of the antibody. A unified modeling method, including multi-sequence alignment and de novo folding technology, is adopted to predict the antibody binding site.

Benefits of technology

Accurate prediction of antigen antibody binding sites under different sequence lengths is achieved, rapid response to binding site changes caused by antigen mutations, and promoting the development of vaccines and antibody drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116543B_ABST
    Figure CN115116543B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method, apparatus, device and computer-readable storage medium for determining an antigen-antibody binding site. The method provided by the embodiments of the present disclosure determines the structural similarity and sequence similarity of the antibody to be predicted with other known antibodies based on the key regions of the antibody by combining the sequence information and structural information of the antibody, and predicts the binding site of the known antibody and antigen that is most similar to the antibody to be predicted as the binding site of the antibody to be predicted and the antigen, thereby achieving accurate prediction of the antibody binding site. The method of the embodiments of the present disclosure can assist antibody drugs for specific binding sites, and quickly and accurately predict changes in binding sites caused by antigen mutations, thereby accelerating the research and development of vaccines or antibody drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence and biology, and more specifically, to a method, apparatus, device, and storage medium for determining antigen-antibody binding sites. Background Art

[0002] Antibody binding site prediction is a major unsolved problem in immunology and a prerequisite for artificial intelligence-assisted vaccine and synthetic antibody design. Currently, the most accurate method for determining binding sites is to observe residues that are spatially close to the three-dimensional binding region, which can be obtained through experimental techniques such as X-ray crystallography. However, these experimental methods are time-consuming and expensive, and computational methods that can overcome these issues are needed to facilitate faster development of therapeutics.

[0003] Existing technologies have proven that sequence-based clustering can identify antibodies with similar binding sites, such as "cluster cloning" of sequences through pedigree clustering. In addition, the rational mining of antibody structural information can also help to effectively classify the binding sites between antibodies and antigens. However, for antibodies belonging to different pedigrees but with binding sites in the same region, or for antibodies with large differences in sequence length, the technical solution of pedigree clustering through antibody sequences is not ideal, and in the antibody database, only a very small number of antibodies have obtained their true structure through experimental analysis. This results in the fact that in most cases reliable structural information cannot be used directly, and can only be obtained through predicted structure. Existing methods based on antibody structural information cannot effectively classify antibody sequences without templates.

[0004] Therefore, an efficient and accurate method for predicting antigen-antibody binding sites is needed, which can determine the antigen-antibody binding sites of template-free antibodies when the lengths of antibody sequences are different. Summary of the Invention

[0005] In order to solve the above problems, the present disclosure combines the sequence information and structural information of antibodies, and jointly determines the known antibodies most similar to the antibody to be predicted based on the structural similarity and sequence similarity between the antibody to be predicted and other known antibodies, thereby achieving accurate prediction of antibody binding sites.

[0006] Embodiments of the present disclosure provide a method, apparatus, device, and computer-readable storage medium for determining an antigen-antibody binding site.

[0007] Embodiments of the present disclosure provide a method for determining an antigen-antibody binding site, comprising: obtaining an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted; determining, based on at least a portion of the heavy chain and at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted, respectively, a heavy chain sequence similarity and a light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted; determining, based on the plurality of known antibodies and the heavy chain of the antibody to be predicted, a structural similarity between each of the plurality of known antibodies and the antibody to be predicted; and determining, based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted, a known antibody among the plurality of known antibodies that is most similar to the antibody to be predicted, and using the binding site of the known antibody with the antigen as the binding site of the antibody to be predicted with the antigen.

[0008] An embodiment of the present disclosure provides an antigen-antibody binding site determination device, comprising: an antibody acquisition module configured to acquire an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted; a sequence alignment module configured to determine, based on at least a portion of the heavy chain and at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted, a heavy chain sequence similarity and a light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted; a structure alignment module configured to determine, based on the plurality of known antibodies and the heavy chain of the antibody to be predicted, a structural similarity between each of the plurality of known antibodies and the antibody to be predicted; and a site determination module configured to determine, based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted, a known antibody that is most similar to the antibody to be predicted among the plurality of known antibodies, and to use the binding site of the known antibody with the antigen as the binding site of the antibody to be predicted with the antigen.

[0009] An embodiment of the present disclosure provides an antigen-antibody binding site determination device, comprising: one or more processors; and one or more memories, wherein a computer executable program is stored in the one or more memories, and when the computer executable program is executed by the processor, the antigen-antibody binding site determination method described above is executed.

[0010] An embodiment of the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon. When the instructions are executed by a processor, the instructions are used to implement the above-mentioned method for determining antigen-antibody binding sites.

[0011] Embodiments of the present disclosure provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for determining an antigen-antibody binding site according to an embodiment of the present disclosure.

[0012] Compared with existing methods of performing family tree clustering based on antibody sequences or technologies of determining binding sites by mining antibody structural information, the methods provided by the embodiments of the present disclosure can accurately predict antigen-antibody binding sites without being affected by differences in sequence lengths. In addition, a unified modeling approach is adopted for the antibody structural model, and a de novo folding approach is used for all difficult-to-predict regions, thereby avoiding the problem of template-free antibodies.

[0013] The method provided in the embodiments of the present disclosure combines the sequence information and structural information of the antibody to be predicted, determines the structural similarity and sequence similarity of the antibody to be predicted with other known antibodies based on the key regions of the antibody, and predicts the binding sites of the known antibodies and antigens that are most similar to the antibody to be predicted as the binding sites of the antibody to be predicted and the antigen, thereby achieving accurate prediction of the antibody binding sites. The method of the embodiments of the present disclosure can help antibody drugs targeting specific binding sites, and can quickly and accurately predict changes in binding sites caused by antigen mutations, thereby accelerating the development of vaccines or antibody drugs. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0015] Figure 1 is a schematic diagram illustrating determination of structural similarity between antibody sequences according to an embodiment of the present disclosure;

[0016] Figure 2 is a flow chart illustrating a method for determining an antigen-antibody binding site according to an embodiment of the present disclosure;

[0017] Figure 3 is a schematic flow chart illustrating the determination of sequence similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure;

[0018] Figure 4A is a schematic diagram showing partial numbering results of IMGT numbering of multiple antibody sequences according to an embodiment of the present disclosure;

[0019] Figure 4B is a schematic diagram illustrating extraction of CDR region sequences from antibody sequences based on IMGT numbering according to an embodiment of the present disclosure;

[0020] Figure 4C is a schematic diagram showing a multiple sequence alignment of CDRH3 region sequences according to an embodiment of the present disclosure;

[0021] Figure 5 is a flow chart illustrating a method for determining the structural similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure;

[0022] Figure 6 is a schematic flow chart illustrating the determination of the structural similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure;

[0023] Figure 7 is a schematic diagram illustrating an apparatus for determining an antigen-antibody binding site according to an embodiment of the present disclosure;

[0024] Figure 8 A schematic diagram of an apparatus for determining antigen-antibody binding sites according to an embodiment of the present disclosure is shown;

[0025] Figure 9 A schematic diagram illustrating the architecture of an exemplary computing device according to an embodiment of the present disclosure; and

[0026] Figure 10 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of the present disclosure more apparent, the exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0028] In this specification and the accompanying drawings, substantially the same or similar steps and elements are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance or ranking.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.

[0030] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.

[0031] The method for determining the antigen-antibody binding site disclosed herein may be based on artificial intelligence (AI). Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a manner similar to human intelligence. For example, for the method for determining the antigen-antibody binding site based on artificial intelligence, it can determine the known antibody that is most similar to the structure and sequence of the antibody to be predicted in a manner similar to how humans identify the similarity between the structure and sequence of the antibody to be predicted and the structure and sequence of known antibodies through the naked eye, thereby determining the binding site of the antibody to be predicted with the antigen. By studying the design principles and implementation methods of various intelligent machines, artificial intelligence enables the method for determining the antigen-antibody binding site disclosed herein to have the function of quickly and accurately predicting changes in the binding site caused by antigen mutations.

[0032] The antigen-antibody binding site determination method disclosed in the present invention can be based on the study of the interaction between antigen and antibody. Among them, there are unique groups on the surface of the antigen molecular structure that can specifically bind to and recognize antibody molecules and cell receptors, which are called antigenic determinants or epitopes. An epitope is an inherent structure on a protein molecule and is an inherent functional characteristic that is manifested when it is combined with a reactive substance. The epitope is the basis of the antigenicity of the viral molecule. Determining the sequence and conformational information of the epitope is the basis for revealing the mechanism of antigen-antibody interaction and is also of great guiding significance for the design of bioactive drugs and vaccines. The antibody molecule can be composed of a light chain and a heavy chain, and the light chain and the heavy chain each include a variable region and a constant region, wherein the variable region determines the specificity of binding to the antigen, and the constant region is related to the immune effect of the antibody. The complementary determining region (CDR) in the variable region is a potential region for the combination of the antibody and the antigen, and has a very strong flexible structure. It is a key structural region that determines the diversity of antibody recognition antigens and the specificity of the interaction. The B cells that produce different antibodies can change the amino acids in this region through VDJ gene recombination and somatic hypermutation, thereby enhancing the binding ability to the antigen.

[0033] Alternatively, the antigen-antibody binding site determination method of the present disclosure can be based on the antibody sequence numbering method. Standardized antibody numbering methods can be used to accurately define the complementary determining region and the binding affinity and specific light chain and heavy chain residues that affect the antibody-antigen interaction. Antibody numbering can include multiple common numbering schemes, such as the Kabat numbering scheme, the Chothia numbering scheme, and the international immunogenetics information system (IMGT, international ImMunoGeneTics information system) numbering scheme adopted in the present disclosure. IMGT is the main reference numbering scheme for immunogenetics and immunoinformatics, and it is numbered to the amino acid unit of antibody, and after numbering, it is ensured that the amino acid type of the highly conserved position of antibody is fixed, and it is convenient to compare the numbered antibody sequences simultaneously. Of course, the present disclosure only performs antibody sequence numbering as an example and not as a limitation with the above-mentioned IMGT numbering scheme, and therefore, other numbering schemes that can achieve similar effects are equally applicable to the antigen-antibody binding site determination method of the present disclosure.

[0034] Alternatively, the antigen-antibody binding site determination method disclosed herein can be based on multiple sequence alignment (MSA, Multiple Sequence Alignment). Multiple sequence alignment compares the amino acid sequences or nucleic acid sequences of multiple (3 or more) protein molecules with a phylogenetic relationship, and arranges the same bases or amino acid residues in the same column as much as possible so that the aligned bases or amino acid residues are homologous in evolution. Multiple sequence alignment is mainly for finding similar sequences. Similar sequences often originate from a common ancestral sequence, and they are likely to have similar spatial structures and biological functions. Therefore, for a protein with a known sequence but unknown structure and function, if the mechanism and function of certain proteins with similar sequences are known, the structure and function of the protein with unknown structure and function can be inferred. Therefore, in the antigen-antibody binding site determination method disclosed herein, the multiple sequence alignment method can be used to find known antibodies similar to the antibody sequence to be predicted, thereby being used to infer the unknown structure of the antibody to be predicted.

[0035] In summary, the solutions provided by the embodiments of the present disclosure involve technologies such as artificial intelligence, antibody numbering, and multiple sequence alignment. The embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings.

[0036] Proteins play an indispensable and important role in living organisms. They are not only the material basis of cells and tissues, but also participate in the composition of various functional reactions. The recognition process between proteins and ligand molecules, including the recognition of antigens and antibodies, enzymes and substrates, hormones and receptors, is involved in almost all important life activities in the body, such as immune responses, biochemical reactions, signal transduction, etc. The analysis of protein-ligand interactions is the basis for understanding the molecular mechanisms and regulatory processes of various biological functions in life activities. Among them, the study of the antigen-antibody recognition mechanism is helpful for the molecular design and mechanism explanation of vaccines, and helps guide the research of affinity-matured therapeutic antibodies. It is of great significance in disease prevention and clinical treatment and diagnosis applications.

[0037] As a key step in analyzing antibody function and discovering its potential as a therapeutic drug, there has been some research on identifying antibody-antigen binding sites. Existing techniques have demonstrated that sequence-based clustering can identify antibodies with similar binding sites. Currently, a commonly used method involves "cluster cloning" sequences, a form of lineage clustering. This method can be implemented in many ways, such as by mapping the V genes of an antibody's heavy and light chains to the V or J genes of the nearest neighboring immunoglobulin. Sequence similarity (e.g., including sequence length) of CDRH3 and CDRL3 (the third of the three CDR regions of the antibody's heavy (H) and light (L) chains, respectively) is then compared to identify antibodies of the same lineage. Other approaches have also attempted to ignore J gene segment information or consider only partial heavy chain sequences. Recent studies have also shown that rationally mining antibody structural information can also help effectively classify antibody-antigen binding sites. This method first performs homology modeling on the entire antibody Fv sequence. To ensure the quality of the overall modeling, this approach only retains structural models where the variable region can find a template and does not require de novo folding. After modeling is completed, structural clustering is performed separately according to the length of the antibody's CDR region (i.e., the three CDRH regions of the antibody's heavy chain (H) - CDRH1, CDRH2 and CDRH3, and the three CDRL regions of the antibody's light chain (L) - CDRL1, CDRL2 and CDRL3, a total of 6 CDR regions). Figure 1 is a schematic diagram illustrating the determination of structural similarity between antibody sequences according to an embodiment of the present disclosure. Each time the structural cluster similarity is obtained, after using the spatial alignment algorithm, such as Figure 1 As shown in FIG, the root mean square error (RMSD) of the distance between the corresponding α carbon atoms of the two sequences is calculated as the basis for judging the structural similarity of the two sequences. Figure 1In the lower right part, multiple example sequences are given for the third CDR region (CDRH3 and CDRL3) with the highest variability in the CDR region, wherein these amino acid sequences are classified by horizontal lines. For example, for the CDRH3 region, the 11 sequences shown are divided into four categories based on the amino acid types therein, wherein the amino acid sequences in each category can be considered to have a high probability of having similar spatial structures and biological functions.

[0038] However, for the above-mentioned technical solution of pedigree clustering by antibody sequences, in antibody databases (such as the new crown antibody database (CoV-AbDab)), there are often some antibodies that obviously belong to different lineages but have binding sites in the same area. Due to the existence of such antibody data, the accuracy of binding site prediction will be directly affected. In addition, since this method is usually only applicable to antibodies with dense pedigree clustering and more significant associations, it is not ideal for antibodies with large differences in sequence lengths. The existing technology of judging binding sites by mining antibody structural information also has major defects. First of all, in the antibody database, only a very small number of antibodies have obtained their true structure through experimental analysis methods, which results in the vast majority of cases not being able to directly use reliable antibody structure information, and can only be obtained by predicting the structure. In order to reduce the error in modeling accuracy, the antibody sequences without templates are filtered, which makes it impossible to effectively classify the antibody sequences without templates. In addition, in the structural clustering part, since the structural space alignment requires the antibody sequences to be of equal length, the number of antibodies that can be used for similarity comparison is also limited.

[0039] Based on this, the present disclosure provides a method for determining antigen-antibody binding sites, which combines the sequence information and structural information of the antibody to determine the known antibody most similar to the antibody to be predicted based on the structural similarity and sequence similarity between the antibody to be predicted and other known antibodies, thereby achieving accurate prediction of the antibody binding site.

[0040] Compared with existing methods of performing family tree clustering based on antibody sequences or technologies of determining binding sites by mining antibody structural information, the methods provided by the embodiments of the present disclosure can accurately predict antigen-antibody binding sites without being affected by differences in sequence lengths. In addition, a unified modeling approach is adopted for the antibody structural model, and a de novo folding approach is used for all difficult-to-predict regions, thereby avoiding the problem of template-free antibodies.

[0041] The method provided in the embodiments of the present disclosure combines the sequence information and structural information of the antibody to be predicted, determines the structural similarity and sequence similarity of the antibody to be predicted with other known antibodies based on the key regions of the antibody, and predicts the binding sites of the known antibodies and antigens that are most similar to the antibody to be predicted as the binding sites of the antibody to be predicted and the antigen, thereby achieving accurate prediction of the antibody binding sites. The method of the embodiments of the present disclosure can help antibody drugs targeting specific binding sites, and can quickly and accurately predict changes in binding sites caused by antigen mutations, thereby accelerating the development of vaccines or antibody drugs.

[0042] Figure 2 FIG. 2 is a flow chart illustrating a method 200 for determining an antigen-antibody binding site according to an embodiment of the present disclosure.

[0043] In step 201, an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted may be obtained.

[0044] As described above, the multiple known antibodies obtained and the antibody to be predicted bind to the same antigen via corresponding binding sites. The sequences and structures of these known antibodies, as well as their binding sites with the antigen, can be predetermined. Therefore, the binding site of the antibody to be predicted can be determined based on the similarity between the antibody to be predicted and these known antibodies. Furthermore, to predict the binding sites of different antibodies on the same antigen, the antibody light chain and heavy chain sequences can first be separated.

[0045] According to an embodiment of the present disclosure, at least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted may respectively include three complementary determining regions in the heavy chain of each of the plurality of known antibodies and the antibody to be predicted. Similarly, according to an embodiment of the present disclosure, at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted may respectively include three complementary determining regions in the light chain of each of the plurality of known antibodies and the antibody to be predicted.

[0046] Alternatively, taking a heavy chain and a light chain of an antibody as an example, the variable regions on the heavy chain and light chain may each have three complementary determining regions. For the heavy chain, its three CDR regions (i.e., CDRH regions) are CDRH1, CDRH2, and CDRH3, respectively, while for the light chain, its three CDR regions (i.e., CDRL regions) are CDRL1, CDRL2, and CDRL3, respectively. Among them, the first two complementary determining regions, CDR1 and CDR2 (i.e., CDRH1 and CDRH2, or CDRL1 and CDRL2), have lower variability than the third complementary determining region, CDR3 (CDRH3 or CDRL3). That is, for the antibody sequence, the CDRH3 region on the heavy chain and the CDRL3 region on the light chain have the highest variability on the corresponding peptide chains. Therefore, considering that the CDR region is a potential region for antibody-antigen binding, and that similar antibody sequences have approximately the same amino acid types in highly conserved regions, in subsequent similarity judgments, key regions in the antibody sequence (e.g., CDR1, CDR2, and CDR3 regions, or only the CDR3 region) can be intercepted for information comparison without having to compare the entire antibody sequence, so as to avoid the problem of large differences in comparison results due to different lengths between antibody sequences.

[0047] In step 202, the heavy chain sequence similarity and light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted can be determined based on at least a portion of the heavy chain and at least a portion of the light chain of each of the multiple known antibodies and the antibody to be predicted.

[0048] Optionally, for the heavy chain and light chain in the antibody, the heavy chain sequence similarity and light chain sequence similarity between the predicted antibody and the known antibody can be determined respectively, wherein since the heavy chain and the light chain have similar regional structures, the method for determining the heavy chain sequence similarity and the light chain sequence similarity can be similar.

[0049] Figure 3 is a schematic flow chart illustrating the determination of sequence similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure. Figure 3 As shown, Figure 3 The schematic process for determining sequence similarity between antibodies in can be performed separately for antibody heavy chains and antibody light chains.

[0050] According to an embodiment of the present disclosure, determining the sequence similarity of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted based on at least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted may include performing a sequence similarity comparison on the heavy chain of each of the plurality of known antibodies and the antibody to be predicted. Optionally, the sequence similarity comparison may be represented as follows: Figure 3 The boxed part of the schematic flow chart shown.

[0051] According to an embodiment of the present disclosure, the sequence similarity comparison may include: for each known antibody among the multiple known antibodies, extracting the three complementary determining regions of the respective heavy chains from the heavy chains of the known antibody and the antibody to be predicted; and performing a multiple sequence comparison on the three complementary determining regions of the heavy chain of the known antibody and the three complementary determining regions of the heavy chain of the antibody to be predicted to determine the sequence similarity of the heavy chains of the known antibody and the antibody to be predicted.

[0052] Alternatively, the determination of sequence similarity can be achieved by extracting key regions in the sequence as described above. For example, for the heavy chain portion, the three CDRH region sequences of the heavy chain can be extracted for multiple sequence alignment with the CDRH region sequences of the heavy chains of other antibodies to determine the sequence similarity between the predicted antibody and the known antibodies.

[0053] According to an embodiment of the present disclosure, extracting three complementary determining regions of each heavy chain from the respective heavy chains of the known antibody and the antibody to be predicted may include: performing antibody sequence numbering on the heavy chains of the known antibody and the antibody to be predicted, respectively, and extracting antibody sequences within a specific numbering range from the heavy chains of the known antibody and the antibody to be predicted that have been numbered by the antibody sequence, respectively, wherein the specific numbering range corresponds to the three complementary determining regions.

[0054] Optionally, the number of each position in the antibody sequence can be determined by first numbering the antibody sequence, and in this numbering process, the numbering range of the above-mentioned key region (CDR region) is fixed, so the amino acid sequence of the key region can be extracted from the antibody sequence by the number corresponding to the fixed numbering range.

[0055] As an example, in an embodiment of the present disclosure, the key regions are extracted from the antibody sequence using the IMGT numbering method, which counts residues continuously from 1 to 128 based on the germ-line V sequence alignment.

[0056] Figure 4A FIG1 is a schematic diagram showing a partial numbering result of IMGT numbering of multiple antibody sequences according to an embodiment of the present disclosure. Figure 4B FIG. 1 is a schematic diagram illustrating extraction of CDR region sequences from antibody sequences based on IMGT numbering according to an embodiment of the present disclosure.

[0057] Figure 4A The four antibody sequences SH1, SH2, SH3 and SH4 are shown in Figure 1. Figure 4A (a))、CDRH2( Figure 4A (b)) and CDRH3( Figure 4A (c)) Some example region sequences related to the above, where: Figure 4A The light grey regions in the correspond to the conserved regions in the antibody sequence, while the dark grey regions correspond to the CDR regions in the antibody sequence, e.g. Figure 4A The dark grey areas in (a) correspond to the CDRH1 regions of each of the four antibody sequences. Figure 4A The dark grey region in (b) corresponds to the CDRH2 region, while Figure 4A The dark grey region in (c) corresponds to the CDRH3 region. Figure 4A As shown, in the conserved regions of the antibody sequence, the amino acid types change less, while in the CDR regions, the changes are more. At the same time, CDRH1 and CDRH2 contain fewer amino acids and fewer amino acid changes than CDRH3, which has the highest variability.

[0058] like Figure 4A As shown in FIG, using the IMGT numbering method, the amino acids in the CDR regions of an antibody sequence are numbered within a fixed numbering range. For example, for the CDR1 region, the numbering range is [27, 38], for the CDR2 region, the numbering range is [56, 65], and for the CDR3 region, the numbering range is [105, 117]. This allows for the subsequent extraction of sequences for key regions to be quickly and conveniently extracted directly from the numbered sequences. For example, to extract the sequence of the CDR3 region, the amino acid sequence within the numbering range [105, 117] can be directly intercepted from the numbered antibody sequence.

[0059] Alternatively, for an antibody sequence comprising a CDR region containing more amino acids than the number in the number range corresponding to the CDR region, a corresponding number of additional numbers may be inserted into the number range corresponding to the CDR region so that the number in the number range equals the number of amino acids in the CDR region. Figure 4A As shown in (c), for the antibody sequences SH1 and SH2, since they include numbers in the CDRH3 region that are in the number range corresponding to the CDR region (e.g., Figure 4A (c) contains 13) more amino acids, so a corresponding number of additional numbers (e.g., 111A, 111B, 111C, 112D, 112C, 112B and 112A) are inserted in the middle part of the CDRH3 region (e.g., positions numbered 111 and 112), so that the amino acids in the CDRH3 region are numbered within the numbering range corresponding to the region.

[0060] Alternatively, for an antibody sequence including a CDR region having an amino acid number less than the number in the number range corresponding to the CDR region, a corresponding number of spaces may be inserted in the number range corresponding to the CDR region and each space may occupy a number, so that the number in the number range is equal to the number of amino acids in the CDR region. As described above, for the same purpose, Figure 4A For antibody sequences whose CDR regions contain fewer amino acids than the number in the number range corresponding to the CDR region, a corresponding number of spaces can be inserted in the middle of the number range and each space can occupy a number to better match the available structural data.

[0061] As described above, the numbering of each position in the antibody sequence can be determined by antibody sequence numbering. Furthermore, the positional alignment of the amino acids in the CDR regions of the antibody sequences using the antibody sequence numbering method can be achieved by inserting additional numbers or spaces, such that the antibody sequences within the numbering range corresponding to the CDR regions (including spaces) are of the same length. Therefore, antibody sequences of unequal lengths can be converted to equal lengths, and since the numbering range of the key regions (CDR regions) is fixed, the antibody sequences of the key regions can be extracted from the antibody sequence using the numbers corresponding to this fixed numbering range.

[0062] like Figure 4B As shown, for Figure 4A By parsing the aligned sequences numbered by IMGT, the four antibody sequences SH1, SH2, SH3, and SH4 of unequal length can be extracted from the three CDR regions corresponding to these antibody sequences. Among the extracted antibody sequence portions, the antibody sequence portions corresponding to the same CDR region have the same sequence length (including spaces). Therefore, based on the antibody sequences extracted from the key regions, a multiple sequence alignment operation can be performed to determine the sequence similarity between the antibody sequences.

[0063] Taking the CDRH3 region sequence as an example, Figure 4C Schematic diagram showing a multiple sequence alignment of CDRH3 region sequences according to an embodiment of the present disclosure.

[0064] Optionally, for Figure 4B The antibody sequence parts extracted from the CDRH3 region can be aligned by multiple sequences to ensure that as many columns as possible in these antibody sequence parts have the same (or similar) amino acids, that is, the sites of the same (or similar) residues are located in the same column. Figure 4CAs shown, the black part corresponds to the column that does not contain space elements, that is, all antibody sequences participating in the alignment have amino acid elements in this column. This is the result of multiple sequence alignment, which is obtained by moving the spaces in the antibody sequences so that the antibody sequences aligned by multiple sequences have the highest similarity.

[0065] As described above, for the heavy chain of an antibody sequence, a multiple sequence alignment can be performed on the antibody sequence portion of the three CDR regions extracted therefrom, thereby obtaining the heavy chain sequence similarity between multiple known antibodies and the antibody to be predicted. For example, the heavy chain sequence similarity can be determined by splicing the three CDR regions of the antibody sequence obtained through the multiple sequence alignment and using a specific calculation method (e.g., the BLOSUM62 protein sequence alignment scoring matrix).

[0066] According to an embodiment of the present disclosure, determining the light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted based on at least a portion of the light chain of each of the multiple known antibodies and the antibody to be predicted includes performing the sequence similarity alignment on the light chain of each of the multiple known antibodies and the antibody to be predicted.

[0067] For the light chain of the antibody sequence, the light chain sequence similarity of the antibody sequence can also be determined by determining the heavy chain sequence similarity with the heavy chain of the antibody sequence, such as referring to Figure 3 as well as Figures 4A-4C The described sequence similarity comparison operation can be applied to the light chain sequence of an antibody sequence to obtain the light chain sequence similarity between multiple known antibodies and the antibody to be predicted, which will not be described in detail herein.

[0068] Based on the above description, the heavy chains and light chains of the respective known antibodies and the antibodies to be predicted are respectively subjected to the following Figure 3 The operations shown can determine the heavy chain sequence similarity and light chain sequence similarity between multiple known antibodies and the antibody to be predicted.

[0069] Next, in step 203, the structural similarity between each of the plurality of known antibodies and the antibody to be predicted may be determined based on the heavy chains of the plurality of known antibodies and the antibody to be predicted.

[0070] As mentioned above, the structural information of antibodies is also helpful for effectively classifying the binding sites between antibodies and antigens. Therefore, it is also possible to find the known antibody that is most similar to the antibody to be predicted based on the structural similarity between multiple known antibodies and the antibody to be predicted.

[0071] Figure 5 FIG. 4 is a flowchart illustrating a method for determining the structural similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure. Figure 6is a schematic flowchart illustrating the determination of the structural similarity between a known antibody and an antibody to be predicted according to an embodiment of the present disclosure.

[0072] According to an embodiment of the present disclosure, step 203 may include: Figure 5 Steps 2031-2034 are shown.

[0073] In the embodiments of the present disclosure, since antibodies are composed of heavy chains and light chains, and the heavy chain has a larger number of amino acid residues than the light chain and thus has a more complex protein structure, when determining the structural similarity between antibodies, only the heavy chain portion of the antibody can be considered for structural similarity determination, such as Figure 6 shown.

[0074] In step 2031, at least one estimated structural model of the antibody to be predicted may be determined based on the heavy chain of the antibody to be predicted.

[0075] Alternatively, structural modeling can be performed on only the heavy chain partial sequence of the antibody to be predicted to obtain at least one estimated structural model. Each of the estimated structural models obtained can adopt a unified protein structure modeling method, and a de novo folding method can be adopted for key regions therein (e.g., CDR regions).

[0076] Alternatively, when performing structural modeling on the heavy chain partial sequence of the antibody to be predicted, an existing protein structure prediction tool can be selectively used to adopt a "de novo folding" protein structure prediction method to predict the three-dimensional structure of the protein from the amino acid sequence of the protein. For example, the AI tool "tFold" developed by Tencent can be used to generate at least one estimated structural model of the antibody to be predicted based on the heavy chain sequence of the antibody to be predicted, wherein the multi-data source fusion technology can first be used to mine the co-evolutionary information in multiple sets of multi-sequence alignments, and then a deep cross-attention residual network (DCARN) can be used to improve the prediction accuracy of some important protein two-dimensional structural information (e.g., residue pair distance and orientation matrix). Finally, a template-assisted free modeling (TBFM) method can be used to effectively fuse the structural information in the 3D model generated by free modeling (FM) and template modeling (TBM), thereby greatly improving the accuracy of the final three-dimensional modeling.

[0077] Of course, in addition to the above-mentioned method of predicting the three-dimensional spatial structure of the heavy chain sequence of the antibody to be predicted based on the heavy chain sequence of the antibody to be predicted through existing protein structure prediction tools such as tFold, other protein structure prediction methods that can achieve the same effect can also be applied to the antigen-antibody binding site prediction method disclosed in the present invention. The present disclosure only describes the above-mentioned tFold tool as an example and not as a limitation.

[0078] In step 2032, an estimated structural model may be selected from the at least one estimated structural model based on at least a portion of each of the at least one estimated structural model as the structural model of the antibody to be predicted.

[0079] After generating at least one estimated structural model of the antibody to be predicted as described above, quality assessment may be performed on these estimated structural models to select the estimated structural model with the best quality as the structural model of the antibody to be predicted.

[0080] According to an embodiment of the present disclosure, step 2032 may include: for each of the at least one estimated structural model, separately cutting out the partial structure corresponding to the third complementarity determining region from the estimated structural model, and performing protein structure quality assessment on the partial structure in the estimated structural model; and selecting the estimated structural model with the highest protein structure quality evaluated from the at least one estimated structural model as the structural model of the antibody to be predicted.

[0081] Optionally, a protein structure quality assessment can be performed on the structural portion corresponding to the key region in the generated estimated structural model. For example, in an embodiment of the present disclosure, a protein structure quality assessment can be performed only on the portion of the structure corresponding to the CDRH3 region in the generated estimated structural model. This is because the conserved regions of the aforementioned multiple known antibodies and the antibody to be predicted generally have similar amino acid sequences and even protein structures, while in the non-conserved regions, the CDRH3 region has the highest variability, while the variability of other CDR regions is lower. Therefore, the quality of the entire antibody structure can be estimated based on the quality of the partial structure corresponding to the CDRH3 region with the highest variability (hereinafter referred to as the CDRH3 heavy chain predicted structure).

[0082] Alternatively, the features of single-model evaluation and consensus evaluation methods can be combined, namely, information about a single predicted structure and the relationship between that predicted structure and other predicted conformations of the same sequence are simultaneously collected. The predicted two-dimensional protein structure information is then combined with the searched template information and the conformation of the CDRH3 heavy chain predicted structure to be evaluated to generate a graph with amino acid residues as vertices and residue-to-residue distance relationships as edges. Finally, the features of single-model evaluation and consensus evaluation methods can be integrated into the graph to predict the accuracy of the CDRH3 heavy chain predicted structure using a message passing network. Thus, the accuracy of each of at least one estimated structural model can be determined using the above method, and the estimated structural model with the highest accuracy among these estimated structural models can then be selected as the structural model of the antibody to be predicted.

[0083] Similarly, in addition to the above-mentioned protein structure quality assessment method, other protein structure quality assessment methods that can achieve the same effect can also be applied to the antigen-antibody binding site prediction method disclosed in the present disclosure. The present disclosure only describes the above-mentioned method as an example rather than a limitation.

[0084] After determining the structural model of the antibody to be predicted in step 2032, the similarity between the structural models of the antibody to be predicted and known antibodies can be analyzed. The similarity analysis of the structural models requires first aligning the structural models in three-dimensional space, and the structural alignment requires first aligning the antibody sequence portions corresponding to the structural models.

[0085] Therefore, in step 2033, a multiple sequence alignment may be performed on at least a portion of the heavy chain of the antibody to be predicted and at least a portion of the heavy chain of each of the plurality of known antibodies.

[0086] According to an embodiment of the present disclosure, the at least a portion of the heavy chain of the antibody to be predicted includes the third complementarity determining region of the three complementarity determining regions in the heavy chain of the antibody to be predicted, the at least a portion of each of the multiple known antibodies includes the third complementarity determining region of the three complementarity determining regions in the heavy chain of each of the multiple known antibodies, and the at least a portion of each of the at least one estimated structural model corresponds to the third complementarity determining region.

[0087] Alternatively, at least a portion of the heavy chain of the antibody to be predicted and at least a portion of the heavy chain of the known antibody may correspond to their CDRH3 regions, respectively, that is, a multiple sequence alignment is performed between the CDRH3 region sequence of the antibody to be predicted and the CDRH3 region sequence of the known antibody. Figure 3 The IMGT number described is extracted.

[0088] In step 2034, based on the result of the multiple sequence alignment, a structural alignment may be performed on the structural model of the antibody to be predicted and the structural model of each of the multiple known antibodies, and the structural similarity between each of the multiple known antibodies and the antibody to be predicted may be determined.

[0089] Through multiple sequence alignment, the sequence elements (including spaces) of the CDRH3 region sequences of the predicted antibody and the CDRH3 region sequences of known antibodies can be aligned one by one and have the same sequence length. Therefore, the structural model can be structurally aligned based on the antibody sequences obtained through multiple sequence alignment.

[0090] According to an embodiment of the present disclosure, performing a structural alignment on the structural model of the antibody to be predicted and the structural model of each of the multiple known antibodies based on the results of the multiple sequence alignment in step 2034 may include: performing a structural alignment on the portion of the structural model of the antibody to be predicted corresponding to the third complementarity determining region and the portion of the structural model of each of the multiple known antibodies corresponding to the third complementarity determining region based on the results of the multiple sequence alignment.

[0091] Alternatively, since the sequence elements (including spaces) of the CDRH3 region sequence of the antibody to be predicted and the CDRH3 region sequence of the known antibody can be aligned one by one, a structural alignment can be performed based on each alignment element pair, wherein the alignment element pairs can include space-amino acid residue pairs, space-space pairs, and amino acid residue-amino acid residue pairs.

[0092] Alternatively, because the absolute spatial positions of the structures of the predicted antibody and the known antibody may differ, for example, the structure of the predicted antibody may be at the origin, while the structure of the known antibody may be significantly deviated from the origin, before calculating the structural similarity between the predicted antibody and the known antibody, it is necessary to make the structures of the two antibodies overlap as much as possible to facilitate structural alignment. For example, the structure of one of the predicted antibody and the known antibody can be fixed, and the structure of the other can be arbitrarily rotated and / or translated. By searching for the degrees of freedom of rotation and translation, the sum of the distances between the alignment element pairs of the antibodies whose structural similarity is to be determined is minimized.

[0093] According to an embodiment of the present disclosure, determining the structural similarity between each of the multiple known antibodies and the antibody to be predicted may include: for each known antibody among the multiple known antibodies, determining the structural similarity between the structural model of the known antibody and the structural model of the antibody to be predicted based on the distance between the carbon atoms at the alignment position in the structural model of the antibody to be predicted and the structural model of the known antibody after structural alignment.

[0094] As described above, the structural alignment may be performed based on the distance in space of each alignment element pair, wherein the distance is the distance in space between the carbon atoms of the two elements in the alignment element pair.

[0095] Alternatively, the structural similarity between the structural model of the known antibody and the structural model of the antibody to be predicted can be determined based on the spatial distances between all alignment element pairs of the CDRH3 region sequence of the antibody to be predicted and the CDRH3 region sequence of each known antibody (e.g., based on the RMSD of the spatial distances between all alignment element pairs). For alignment element pairs that include at least one blank element, their spatial distances may be disregarded.

[0096] In step 204, based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted, a known antibody among the multiple known antibodies that is most similar to the antibody to be predicted can be determined, and the binding site of the known antibody with the antigen can be used as the binding site of the antibody to be predicted with the antigen.

[0097] In the embodiments of the present disclosure, the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between the antibody to be predicted and the known antibodies can be considered simultaneously, so as to jointly determine the known antibody that is most similar to the antibody to be predicted through the structural similarity and sequence similarity between the antibody to be predicted and the known antibodies, thereby achieving accurate prediction of the antigen-antibody binding site.

[0098] According to an embodiment of the present disclosure, determining in step 204 a known antibody among the multiple known antibodies that is most similar to the antibody to be predicted based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted may include: for each of the multiple known antibodies, determining the antibody similarity between the known antibody and the antibody to be predicted based on a weighted sum of the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between the known antibody and the antibody to be predicted; and selecting, from the multiple known antibodies, the known antibody with the highest antibody similarity to the antibody to be predicted as the known antibody most similar to the antibody to be predicted.

[0099] Optionally, determining the antibody similarity between the known antibody and the antibody to be predicted based on the weighted sum of the structural similarity, heavy chain sequence similarity and light chain sequence similarity between the known antibody and the antibody to be predicted may include determining the magnitude of influence (i.e., weight) of each of the structural similarity, heavy chain sequence similarity and light chain sequence similarity on the similarity between the antibodies.

[0100] Alternatively, the determination of the above-mentioned weights can be obtained by data training. For example, a training data set can be determined based on antibodies of known structures in the Antibody Structural Database (SAbDab), such as antibodies of known structures corresponding to the same antigen, and clustered according to the binding sites of these antibodies on the antigen (with the binding sites as the labels of the antibodies). Therefore, using the above-mentioned algorithm process, structural similarity, heavy chain sequence similarity and light chain sequence similarity can be trained as features, and by making the accuracy of all antibody data classification the highest, the weights of each of these features are searched to determine the weight of linear weighting.

[0101] As described above, the antigen-antibody binding site determination method disclosed herein uses a fusion of sequence and structural information. By using the antibody sequence number to obtain the antibody key region, the IMGT number is used to intercept the key region at the mapping position of the multiple sequence alignment as the benchmark comparison information to determine the sequence similarity between the antibody to be predicted and the known antibody based on the benchmark comparison information, thereby avoiding the large difference in sequence alignment results caused by the different lengths between the antibody sequences. In response to the problem of structural similarity, the antibody key region is modeled and screened and multiple sequence alignments and structural alignments are performed with the key regions of other known antibodies to determine its structural similarity with other known antibodies. In the structural modeling part, a unified modeling method is adopted, and the template information is only used as an optional feature item. All key regions are folded from scratch, thereby avoiding the problem of different lengths between antibody sequences and no template antibodies. Among them, in order to ensure the quality of antibody modeling, the CDRH3 region with the largest change in the antibody structure is separately intercepted to evaluate its quality using protein structure model evaluation technology, and the best structural model is selected as the modeling result.

[0102] Figure 7 FIG. 7 is a schematic diagram illustrating an apparatus 700 for determining an antigen-antibody binding site according to an embodiment of the present disclosure.

[0103] According to an embodiment of the present disclosure, the antigen-antibody binding site determination device 700 may include an antibody acquisition module 701 , a sequence alignment module 702 , a structure alignment module 703 and a site determination module 704 .

[0104] The antibody acquisition module 701 may be configured to acquire an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted.

[0105] Since the acquired multiple known antibodies and the above-mentioned antibody to be predicted act on the same antigen via corresponding binding sites, the sequences and structures of these known antibodies and their binding sites with the antigen can be predetermined. Therefore, the binding site of the antibody to be predicted and the antigen can be determined based on the similarity between the antibody to be predicted and these known antibodies. Optionally, the antibody acquisition module 701 can perform the operations described above with reference to step 201.

[0106] Alternatively, at least a portion of the heavy chain (or light chain) of each of the plurality of known antibodies and the antibody to be predicted may each include three complementarity determining regions of the heavy chain (or light chain) of each of the plurality of known antibodies and the antibody to be predicted. Taking a heavy chain and a light chain of an antibody as an example, the variable regions on the heavy chain and the light chain may each have three complementarity determining regions. For the heavy chain, the three CDR regions (i.e., CDRH regions) are CDRH1, CDRH2, and CDRH3, respectively, and for the light chain, the three CDR regions (i.e., CDRL regions) are CDRL1, CDRL2, and CDRL3, respectively.

[0107] The sequence alignment module 702 can be configured to determine the heavy chain sequence similarity and light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted based on at least a portion of the heavy chain and at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted. Optionally, the sequence alignment module 702 can perform the operations described above with reference to step 202.

[0108] For example, for the heavy chain and light chain in an antibody, the heavy chain sequence similarity and light chain sequence similarity between the predicted antibody and the known antibody can be determined respectively. Since the heavy chain and the light chain have similar regional structures, the methods for determining the heavy chain sequence similarity and the light chain sequence similarity can be similar. For example, for the heavy chain and the light chain of the antibody sequence, the heavy chain sequence similarity and the light chain sequence similarity can be determined respectively by referring to Figure 3 as well as Figures 4A-4C The described sequence similarity alignment operation determines the heavy chain sequence similarity and light chain sequence similarity of the antibody sequences.

[0109] Alternatively, the determination of sequence similarity can be achieved by intercepting key regions in the sequence. For example, for the heavy chain portion, the three CDRH region sequences of the heavy chain can be extracted for multiple sequence alignment with the CDRH region sequences of the heavy chains of other antibodies to determine the sequence similarity between the predicted antibody and the known antibody. For example, the number of each position in the antibody sequence can be determined by first numbering the antibody sequence (e.g., IMGT numbering), and in this numbering process, the numbering range of the above-mentioned key region (CDR region) is fixed, so the amino acid sequence of the key region can be intercepted from the antibody sequence by the number corresponding to the fixed numbering range.

[0110] Optionally, for the extracted antibody sequence portions, multiple sequence alignment can be performed so that as many columns as possible in these antibody sequence portions have the same (or similar) amino acids, that is, the sites of the same (or similar) residues are located in the same column, thereby obtaining the heavy chain (or light chain) sequence similarity between multiple known antibodies and the antibody to be predicted.

[0111] As described above, the structural information of antibodies also helps to effectively classify the binding sites of antibodies and antigens. Therefore, it is also possible to find the known antibodies that are most similar to the antibody to be predicted based on the structural similarity between multiple known antibodies and the antibody to be predicted. Therefore, the structure comparison module 703 can be configured to determine the structural similarity between each of the multiple known antibodies and the antibody to be predicted based on the heavy chains of the multiple known antibodies and the antibody to be predicted. Optionally, the structure comparison module 703 can perform the operations described above with reference to step 203.

[0112] For example, structural modeling can be performed on only the heavy chain partial sequence of the antibody to be predicted to obtain at least one estimated structural model. Each of the estimated structural models obtained can adopt a unified protein structure modeling method, and a de novo folding method can be adopted for key regions therein (e.g., CDR regions).

[0113] Optionally, after generating at least one estimated structural model of the antibody to be predicted as described above, these estimated structural models can be quality assessed to select the estimated structural model with the best quality as the structural model of the antibody to be predicted. For example, the protein structure quality assessment can be performed on the structural portion corresponding to the key region in the generated estimated structural model. After determining the structural model of the antibody to be predicted, the similarity between the structural models of the antibody to be predicted and the known antibodies can be analyzed. For the similarity analysis of the structural models, it is necessary to first perform structural alignment of the structural models in three-dimensional space, and the structural alignment requires first performing sequence alignment of the antibody sequence portion corresponding to the structural model, as described above with reference to steps 2033 and 2034.

[0114] In an embodiment of the present disclosure, the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between the antibody to be predicted and the known antibodies can be considered simultaneously to determine the known antibody most similar to the antibody to be predicted by combining the structural similarity and sequence similarity between the antibody to be predicted and the known antibodies. The site determination module 704 can be configured to determine a known antibody that is most similar to the antibody to be predicted among the multiple known antibodies based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity of each of the multiple known antibodies to the antibody to be predicted, and use the binding site of the known antibody to the antigen as the binding site of the antibody to be predicted to the antigen. Optionally, the site determination module 704 can perform the operations described above with reference to step 204.

[0115] For example, before calculating the weighted sum of the structural similarity, heavy chain sequence similarity and light chain sequence similarity between the known antibody and the antibody to be predicted, the influence of each of the structural similarity, heavy chain sequence similarity and light chain sequence similarity on the similarity between antibodies (i.e., weight) can be determined, which can be obtained by, for example, data training.

[0116] Therefore, by determining the known antibody that is most similar to the antibody to be predicted among multiple known antibodies, the binding site of the known antibody with the antigen can be used as the binding site of the antibody to be predicted, thereby achieving accurate prediction of the antigen-antibody binding site.

[0117] According to yet another aspect of the present disclosure, an apparatus for determining an antigen-antibody binding site is provided. Figure 8 FIG2 shows a schematic diagram of an apparatus 2000 for determining an antigen-antibody binding site according to an embodiment of the present disclosure.

[0118] like Figure 8 As shown, the antigen-antibody binding site determination device 2000 may include one or more processors 2010 and one or more memories 2020. The memories 2020 may store computer-readable code, which, when executed by the one or more processors 2010, may execute the antigen-antibody binding site determination method described above.

[0119] The processor in the embodiments of the present disclosure may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor, and may be an X86 architecture or an ARM architecture.

[0120] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0121] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 9 The architecture of the computing device 3000 shown in FIG. Figure 9 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the antigen-antibody binding site determination method provided by the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 8 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 9 One or more components of a computing device are shown.

[0122] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided. Figure 10 A schematic diagram 4000 of a storage medium according to the present disclosure is shown.

[0123] like Figure 10As shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are executed by a processor, the antigen-antibody binding site determination method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. Volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRRAM). It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0124] Embodiments of the present disclosure also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for determining an antigen-antibody binding site according to an embodiment of the present disclosure.

[0125] Embodiments of the present disclosure provide a method, apparatus, device, and computer-readable storage medium for determining an antigen-antibody binding site.

[0126] Compared with existing methods of performing family tree clustering based on antibody sequences or technologies of determining binding sites by mining antibody structural information, the methods provided by the embodiments of the present disclosure can accurately predict antigen-antibody binding sites without being affected by differences in sequence lengths. In addition, a unified modeling approach is adopted for the antibody structural model, and a de novo folding approach is used for all difficult-to-predict regions, thereby avoiding the problem of template-free antibodies.

[0127] The method provided in the embodiments of the present disclosure combines the sequence information and structural information of the antibody to be predicted, determines the structural similarity and sequence similarity of the antibody to be predicted with other known antibodies based on the key regions of the antibody, and predicts the binding sites of the known antibodies and antigens that are most similar to the antibody to be predicted as the binding sites of the antibody to be predicted and the antigen, thereby achieving accurate prediction of the antibody binding sites. The method of the embodiments of the present disclosure can help antibody drugs targeting specific binding sites, and can quickly and accurately predict changes in binding sites caused by antigen mutations, thereby accelerating the development of vaccines or antibody drugs.

[0128] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the module, program segment, or a part of the code contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0129] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0130] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for determining antigen-antibody binding sites, comprising: Obtaining an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted; Determining the heavy chain sequence similarity and light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted based on at least a portion of the heavy chain and at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted; Determining, based on the multiple known antibodies and the heavy chain of the antibody to be predicted, a structural similarity between each of the multiple known antibodies and the antibody to be predicted, wherein, for each of the multiple known antibodies, the structural similarity between the known antibody and the antibody to be predicted is determined by performing a structural comparison between a structural model of the antibody to be predicted determined based on the heavy chain of the antibody to be predicted and the structural model of the known antibodies; as well as Based on the structural similarity, heavy chain sequence similarity and light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted, a known antibody among the multiple known antibodies that is most similar to the antibody to be predicted is determined, and the binding site of the known antibody with the antigen is used as the binding site of the antibody to be predicted with the antigen.

2. The method according to claim 1, wherein At least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted respectively includes three complementarity determining regions in the heavy chain of each of the plurality of known antibodies and the antibody to be predicted; At least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted respectively includes three complementarity determining regions in the light chain of each of the plurality of known antibodies and the antibody to be predicted; Wherein, determining the sequence similarity of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted based on at least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted comprises performing a sequence similarity comparison on the heavy chain of each of the plurality of known antibodies and the antibody to be predicted, wherein the sequence similarity comparison comprises: For each known antibody in the plurality of known antibodies, extracting three complementary determining regions of the respective heavy chains from the respective heavy chains of the known antibody and the antibody to be predicted; and Performing a multiple sequence alignment of the three complementarity determining regions of the heavy chain of the known antibody and the three complementarity determining regions of the heavy chain of the antibody to be predicted to determine the sequence similarity of the heavy chains of the known antibody and the antibody to be predicted; Wherein, determining the light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted based on at least a portion of the light chain of each of the multiple known antibodies and the antibody to be predicted comprises performing the sequence similarity alignment on the light chain of each of the multiple known antibodies and the antibody to be predicted.

3. The method according to claim 2, wherein: Extracting three complementary determining regions of the respective heavy chains from the respective heavy chains of the known antibody and the antibody to be predicted comprises: The heavy chains of the known antibody and the antibody to be predicted are respectively numbered by antibody sequence, and antibody sequences within a specific numbering range are respectively extracted from the heavy chains of the known antibody and the antibody to be predicted that have been numbered by antibody sequence, and the specific numbering range corresponds to the three complementarity determining regions.

4. The method according to claim 2, wherein: Determining the structural similarity between each of the plurality of known antibodies and the antibody to be predicted based on the heavy chains of the plurality of known antibodies and the antibody to be predicted comprises: determining at least one estimated structural model of the antibody to be predicted based on the heavy chain of the antibody to be predicted; selecting, based on at least a portion of each of the at least one estimated structural model, an estimated structural model from the at least one estimated structural model as the structural model of the antibody to be predicted; performing a multiple sequence alignment of at least a portion of the heavy chain of the antibody to be predicted with at least a portion of the heavy chain of each of the plurality of known antibodies; and Based on the result of the multiple sequence alignment, a structural alignment is performed on the structural model of the antibody to be predicted and the structural model of each of the multiple known antibodies, and the structural similarity between each of the multiple known antibodies and the antibody to be predicted is determined.

5. The method according to claim 4, wherein: The at least a portion of the heavy chain of the antibody to be predicted includes a third complementarity determining region of the three complementarity determining regions in the heavy chain of the antibody to be predicted, the at least a portion of each of the plurality of known antibodies includes a third complementarity determining region of the three complementarity determining regions in the heavy chain of each of the plurality of known antibodies, and the at least a portion of each of the at least one estimated structural model corresponds to the third complementarity determining region; Wherein, based on the result of the multiple sequence alignment, performing a structural alignment on the structural model of the antibody to be predicted and the structural model of each of the plurality of known antibodies comprises: Based on the result of the multiple sequence alignment, a structural alignment is performed on the portion of the structural model of the antibody to be predicted corresponding to the third complementarity determining region and the portion of the structural model of each of the multiple known antibodies corresponding to the third complementarity determining region.

6. The method according to claim 5, wherein: Determining the structural similarity between each of the plurality of known antibodies and the antibody to be predicted comprises: For each known antibody among the multiple known antibodies, the structural similarity between the structural model of the known antibody and the structural model of the antibody to be predicted is determined based on the distance between the carbon atoms at the alignment position in the structural model of the antibody to be predicted and the structural model of the known antibody.

7. The method according to claim 5, wherein: Based on at least a portion of each of the at least one estimated structural model, selecting an estimated structural model from the at least one estimated structural model as the structural model of the antibody to be predicted comprises: For each of the at least one estimated structural model, separately cutting out the partial structure corresponding to the third complementarity determining region from the estimated structural model, and performing protein structure quality assessment on the partial structure in the estimated structural model; and The estimated structural model having the highest protein structure quality evaluated is selected from the at least one estimated structural model as the structural model of the antibody to be predicted.

8. The method of claim 1, wherein: Determining, based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity of each of the plurality of known antibodies to the antibody to be predicted, a known antibody among the plurality of known antibodies that is most similar to the antibody to be predicted comprises: For each of the plurality of known antibodies, determining the antibody similarity between the known antibody and the antibody to be predicted based on a weighted sum of the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between the known antibody and the antibody to be predicted; and From the plurality of known antibodies, the known antibody with the highest antibody similarity to the antibody to be predicted is selected as the known antibody most similar to the antibody to be predicted.

9. An apparatus for determining antigen-antibody binding sites, comprising: an antibody acquisition module, configured to acquire an antibody to be predicted and a plurality of known antibodies corresponding to the same antigen as the antibody to be predicted; a sequence alignment module configured to determine, based on at least a portion of the heavy chain and at least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted, a heavy chain sequence similarity and a light chain sequence similarity between each of the plurality of known antibodies and the antibody to be predicted; a structural comparison module, configured to determine, based on the multiple known antibodies and the heavy chain of the antibody to be predicted, a structural similarity between each of the multiple known antibodies and the antibody to be predicted, wherein, for each of the multiple known antibodies, the structural similarity between the known antibody and the antibody to be predicted is determined by performing a structural comparison between a structural model of the antibody to be predicted determined based on the heavy chain of the antibody to be predicted and the structural model of the known antibodies; as well as The site determination module is configured to determine a known antibody among the multiple known antibodies that is most similar to the antibody to be predicted based on the structural similarity, heavy chain sequence similarity and light chain sequence similarity of each of the multiple known antibodies with the antibody to be predicted, and use the binding site of the known antibody with the antigen as the binding site of the antibody to be predicted with the antigen.

10. The device according to claim 9, wherein At least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted respectively includes three complementarity determining regions in the heavy chain of each of the plurality of known antibodies and the antibody to be predicted; At least a portion of the light chain of each of the plurality of known antibodies and the antibody to be predicted respectively includes three complementarity determining regions in the light chain of each of the plurality of known antibodies and the antibody to be predicted; Wherein, determining the sequence similarity of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted based on at least a portion of the heavy chain of each of the plurality of known antibodies and the antibody to be predicted comprises performing a sequence similarity comparison on the heavy chain of each of the plurality of known antibodies and the antibody to be predicted, wherein the sequence similarity comparison comprises: For each known antibody in the plurality of known antibodies, extracting three complementary determining regions of the respective heavy chains from the respective heavy chains of the known antibody and the antibody to be predicted; and Performing a multiple sequence alignment of the three complementarity determining regions of the heavy chain of the known antibody and the three complementarity determining regions of the heavy chain of the antibody to be predicted to determine the sequence similarity of the heavy chains of the known antibody and the antibody to be predicted; Wherein, determining the light chain sequence similarity between each of the multiple known antibodies and the antibody to be predicted based on at least a portion of the light chain of each of the multiple known antibodies and the antibody to be predicted comprises performing the sequence similarity alignment on the light chain of each of the multiple known antibodies and the antibody to be predicted.

11. The device according to claim 10, wherein Determining the structural similarity between each of the plurality of known antibodies and the antibody to be predicted based on the heavy chains of the plurality of known antibodies and the antibody to be predicted comprises: determining at least one estimated structural model of the antibody to be predicted based on the heavy chain of the antibody to be predicted; selecting, based on at least a portion of each of the at least one estimated structural model, an estimated structural model from the at least one estimated structural model as the structural model of the antibody to be predicted; performing a multiple sequence alignment of at least a portion of the heavy chain of the antibody to be predicted with at least a portion of the heavy chain of each of the plurality of known antibodies; and Based on the result of the multiple sequence alignment, a structural alignment is performed on the structural model of the antibody to be predicted and the structural model of each of the multiple known antibodies, and the structural similarity between each of the multiple known antibodies and the antibody to be predicted is determined.

12. The device according to claim 11, wherein Determining, based on the structural similarity, heavy chain sequence similarity, and light chain sequence similarity of each of the plurality of known antibodies to the antibody to be predicted, a known antibody among the plurality of known antibodies that is most similar to the antibody to be predicted comprises: For each of the plurality of known antibodies, determining the antibody similarity between the known antibody and the antibody to be predicted based on a weighted sum of the structural similarity, heavy chain sequence similarity, and light chain sequence similarity between the known antibody and the antibody to be predicted; and From the plurality of known antibodies, the known antibody with the highest antibody similarity to the antibody to be predicted is selected as the known antibody most similar to the antibody to be predicted.

13. An apparatus for determining antigen-antibody binding sites, comprising: one or more processors; as well as One or more memories storing a computer executable program, wherein when the computer executable program is executed by the processor, the method according to any one of claims 1 to 8 is performed.

14. A computer program product stored on a computer-readable storage medium and comprising computer instructions which, when executed by a processor, cause a computer device to perform the method of any one of claims 1 to 8.

15. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the instructions are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Deep learning method for predicting binding site on antibody through sequence

    CN112397139A

  • Method for the prediction of an epitope

    US20050026215A1

  • Prediction of side-chain degradation in polymers through physics based simulations

    WO2021025940A1