Multi-modal antigen-antibody affinity prediction method based on antibody structure
Through multimodal information encoding and extraction methods, combining heavy and light chain information of antibodies and structural information of antigens, the problem of ignoring 3D structure and single mode modeling in the prior art is solved, and the accuracy of antigen-antibody affinity prediction and model scalability are improved.
Patent Information
- Application Number
- CN202510276155.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-29
AI Technical Summary
When predicting antigen-antibody affinity, the prior art ignores the 3D structural information and single-modal modeling of the antibody, resulting in insufficient information utilization and inaccurate capture of complex multimodal interactions.
Multimodal information encoding and extraction methods are used to extract heavy and light chain information of antibodies through the Roformer network and the GearNet network respectively, and antigen information is extracted in combination with the protein language model ESM2, and the characteristics of network fusion antibodies and antigens are synergistically extracted through cross attention mechanism and multi-scale features to generate robust affinity prediction results.
It improves the accuracy of antigen-antibody affinity prediction and the scalability of the model, enhances the ability to identify key features, and provides a more accurate and comprehensive feature basis.
Smart Images

Figure CN120388619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and more particularly to a method for predicting the affinity of multimodal antigen-antibody based on antibody structure. Background Art
[0002] Determining the affinity of antibody-antigen interaction is an important step in antibody development. High-affinity antibodies are crucial for the rapid and accurate diagnosis of diseases. They can specifically bind to pathogens or abnormal molecules in the body, providing sensitive and specific detection signals. Moreover, the effectiveness of vaccines depends to a large extent on their ability to induce the production of high-affinity antibodies. By optimizing the antigen structure or using adjuvants, the immune response can be enhanced, prompting the body to produce a stronger antibody response. In the research and development of therapeutic antibodies, affinity is one of the key indicators for evaluating antibody quality. High-affinity antibodies can bind to target antigens more effectively, improve the therapeutic effect, and reduce side effects at the same time.
[0003] In the study of antibody-antigen interaction, scholars have also made a series of important progress. Previous AI methods have utilized antibody structure information to predict the affinity between antibodies and antigens. Based on features such as solvent accessible surface area, residue depth, secondary structure information, etc., key features that can effectively capture antigen-antibody interaction have been extracted. However, although this method uses structural information, it ignores the crucial 3D structure in antigen-antibody binding. Techniques such as multi-relational graph construction, multi-level geometric message passing, and contrastive pre-training of large-scale unlabeled protein structure data have been adopted to effectively extract the geometric features of the antigen-antibody complex interface. They also developed Gearbind to predict antigen-antibody affinity. However, this method relies too much on the information of the complex and ignores the information of the antigen and antibody sequences. Some scientists have also focused on sequence information and created models for affinity prediction. A novel model, DG-Affinity, encodes the features of proteins and antibodies separately and fuses them before inputting into the neural network to predict the affinity between antigens and antibodies. Although this method fuses the features, it only makes a simple splicing and ignores the interaction between antigens and antibodies. By pre-training a large language model on a large antibody sequence dataset and pairing the antigen epitope sequence with the antibody sequence, a high-precision antigen-antibody binding prediction model, A2binder, is designed to predict binding specificity and affinity. Although the model makes full use of the information encoding of the antibody and antigen sequences, it ignores the importance of structural information.
[0004] Despite the commendable progress made in these works, they only focus on single-modal modeling. Sequence and structural information, especially the 3D structure of antibodies, is crucial for predicting antigen-antibody affinity. In addition, these studies only combine feature information without considering complex multi-modal interactions. A single binding process cannot determine which information is more important for affinity prediction. Summary of the Invention
[0005] The present invention proposes a multi-modal antigen-antibody affinity prediction method based on antibody structure, which improves the feature capture ability of the antibody-antigen binding mode through multi-modal information encoding and extraction. And it enhances the interaction between the two modalities of antibody sequence and structure to capture the consistency and complementary information between different modalities.
[0006] To achieve the above invention objectives, the technical solutions adopted by the present invention are specifically as follows: A multi-modal antigen-antibody affinity prediction method based on antibody structure, comprising the following steps:
[0007] S10: Collect the Biomap dataset, which contains 1706 antigen-antibody pairing data, including the heavy chain sequence, light chain sequence, antigen sequence, and affinity value of the antibody. The 14H and 14L affinity datasets are also used. The heavy chains of the 14H dataset are different, while the light chains are constant, and the 14L dataset is the opposite. And the affinity datasets are divided into training, validation, and test parts;
[0008] S20: Construct a multi-modal antibody information mining module. In the multi-modal antibody information mining module, the Roformer network is introduced to extract the heavy chain and light chain information of the antibody respectively, and the GearNet network is used to mine the structural information of the antibody. Subsequently, the sequence and structural information are adaptively fused by means of a cross-attention mechanism;
[0009] S30: Construct an antigen information mining module to achieve antigen information extraction by introducing the protein large language model ESM2;
[0010] S40: Input the extracted features into the fusion prediction module and construct a multi-scale feature collaborative extraction network based on CNN. Input the multi-scale antibody and antigen representations into the fusion layer to generate a robust affinity prediction result;
[0011] S50: Evaluate the prediction performance by comparing the error between the affinity value predicted by the model and the true value, and obtain the best model parameters for prediction performance on the training set and validation set;
[0012] S60: The test set data goes through the above steps S1 to S5 to obtain the final affinity value;
[0013] As a multi-modal antigen-antibody affinity prediction method based on the antibody structure provided by the present invention, the specific steps of step S20 are as follows:
[0014] Step S21: Divide the antibody sequence into heavy and light chains and encode them through two Roformer networks pre-trained on a large number of heavy and light chain sequences of antibodies respectively. Let the heavy chain sequence be represented as and the light chain sequence be represented as where N H and N L represent the lengths of the heavy and light chain sequences respectively. The encoded representations of the heavy and light chains are as follows:
[0015] E H = Roformer H (H), E L = Roformer L (L) (1)
[0016] In the formula, Roformer H and Roformer L are Roformer encoders pre-trained on heavy and light chain sequences respectively. After independently encoding the two chains, their representations are concatenated as:
[0017] E = Concat(E H , E L ) (2)
[0018] In the formula l s is the sequence length, and d s is the embedding dimension. This method retains the independent information of the heavy and light chains and simultaneously utilizes their collaborative contributions to binding specificity and affinity prediction through a unified representation;
[0019] Step S22: Mine antibody structure information. First, predict the antibody structure through the antibody structure prediction tool IgFold. Secondly, construct an antibody graph. The antibody structure is represented as a residue-level relational graph where and ε represent the sets of nodes and edges, represents the set of edge types. Three types of edges are constructed for the residue-level antibody graph: If the sequential distance between nodes i and j is lower than the predefined threshold d seq , then add a sequential edge. If the Euclidean distance between nodes i and j is less than the threshold , then add an edge between nodes i and j. Finally, to explain the different spatial scales between proteins, each node is connected to its k nearest neighbors based on the Euclidean distance;
[0020] Step S23: Extract the structural information from the constructed antibody graph, construct the node features of the antibody graph using one-hot encoding, and construct the edge feature f(i,j,r) by concatenating the one-hot encoded node features of nodes i and j, the edge type, and their sequence and spatial distances. Through the GearNet-Edge encoder, an edge graph is constructed based on the antibody graph On this basis In the antibody graph If the edge (i,j) shares a node with the edge (k,l), then an edge is added from (i,j) to (k,l) in Subsequently, the features of the edge (i,j) in the constructed edge graph are updated through the edge message passing mechanism. The message passing formula is defined as:
[0021]
[0022] In the formula represents the feature of the edge (i,j) at layer l, σ is the non-linear activation function ReLU, BN represents batch normalization, represents the neighborhood set of the edge (i,j), Wr is the learnable weight matrix of edge type r. Through multi-layer message passing, the features of the edge are gradually updated to capture the multi-scale relationships between the edges in the antibody structure. After completing the edge message passing, the updated edge features are re-introduced into the antibody graph to enhance the representation ability of the nodes. By combining the edge messages, the feature update formula of node i is re-defined as:
[0023]
[0024] In the formula represents the feature of node i at layer l, represents the neighbor set of node i, is the feature of neighbor node j from the previous layer, and FC refers to the fully connected layer that maps the edge feature to the dimension matching the node feature. Then, all the antibody node features are combined to form the structural embedding representation, denoted as where l t is the number of nodes in the structural embedding, d t is the dimension of the structural feature;
[0025] Step S24: After obtaining the antibody sequence and structural information, through the cross-attention module, the sequence representation and the structural representation are integrated interactively to extract the key relevant features from the two modalities. Then, the multi-modal output X is calculated as follows a :
[0026] Q = EWQ , K = UW K , V = UW V (5)
[0027]
[0028] In the formula, Q is the query matrix, K and V are the key and value matrices respectively, and dk represents the dimension of the key vector. attention(Q, K, V) represents the calculation process of the attention mechanism. softmax(·) represents the softmax function, which is used to calculate the attention weights.
[0029] As a multi-modal antigen-antibody affinity prediction method based on the antibody structure provided by the present invention, the specific steps of step S30 are as follows:
[0030] Step S31: Extract antigen information through the pre-trained protein large prediction model ESM2. Assume that the antigen sequence is represented as S = {a1, a2,..., a L}, where L is the length of the sequence, and a i represents the i-th amino acid residue. First, preprocess the sequence by converting each amino acid a i into its one-hot encoded representation. Add special markers " <cls>”and" <eos>”, and pad the sequence to a fixed length to meet the model input requirements, obtaining the processed sequence S′:
[0031] S′ = {<cls>, a1, a2,..., a L , <eos>, <pad>,...} (7)
[0032] Step S32: Then input the processed sequence into the ESM2 model. ESM2 adopts 33 layers of Transformer encoders, using the multi-head attention mechanism to capture the dependencies between amino acid residues, thereby learning rich features of the antigen sequence. After passing through the encoding layer, the final output H (33) is a tensor of size (L + 2, 1280), where the vector at each position contains the depth feature information of the residue. To extract the overall features of the antigen, a pooling operation is performed on all amino acid residue positions to obtain the final antigen feature vector X p :
[0033]
[0034] In the formula, L represents the length of the antigen sequence, denotes taking the average of the summation result, and H i 33 represents the feature vector of the i-th amino acid residue.
[0035] As a method for predicting the affinity of multi-modal antigen-antibody based on the antibody structure provided by the present invention, the specific steps of step S40 are as follows:
[0036] Step S41: Input the antibody feature X a extracted by the encoder in Sections S2 and S3 and the antigen feature X p into the prediction network to predict the affinity value. Use the multi-scale feature extraction convolutional neural network (MF_CNN) to integrate the information of the antibody and the antigen. MF_CNN uses a 3-layer CNN backbone composed of convolution, pooling, and ReLU operations to extract multi-scale features. Subsequently, use the fully connected layer (FC) and residual operations to further combine these features to obtain the final output:
[0037] F a = Residual(FC(X a )) = FC(X a ) + X a (8)
[0038] F p = Residual(FC(X p )) = FC(X p ) + X p (9)
[0039] In the formula, Fa and Fp are the characteristics of the antibody and antigen respectively, Residual represents the residual operation, and FC represents the fully connected layer.
[0040] Step S42: Transmit the integrated characteristics of the antibody and antigen to the multi-layer perceptron to finally generate the affinity value:
[0041] Affinity = MLP([Fa, Fp]) (10)
[0042] In the formula, MLP represents the multi-layer perceptron.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. Richness of multi-modal features: The present invention constructs a multi-modal antibody information extraction module and an antigen information extraction module simultaneously. For the first time, the present invention applies the structural information and sequence information together to the prediction of antigen-antibody affinity. The multi-modal method improves the information utilization efficiency, obtains a richer feature expression, and provides a more accurate and comprehensive feature basis for subsequent prediction tasks.
[0045] 2. Multi-modal interaction ability: The model in the present invention fuses the features in the antibody structure and sequence spaces through the use of the cross-attention mechanism, and learns the overall and complementary information between the features. The multi-modal interaction can effectively capture the complex relationship between the antibody sequence and structure, enhance the model's ability to identify key features, and thus improve the accuracy of antigen-antibody affinity prediction.
[0046] 3. Improvement of model scalability: When facing new prediction tasks, the present method can quickly adapt to and integrate new information sources without large-scale modification of the overall framework. And because the present method is based on modular design, each part can be optimized and updated independently, further improving the scalability of the model. Description of the Drawings
[0047] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention;
[0048] Figure 1 It is a flowchart of the multi-modal antigen-antibody affinity prediction method based on antibody structure in the embodiment of the present application;
[0049] Figure 2 It is a model framework diagram of the multi-modal antigen-antibody affinity prediction method based on antibody structure in the embodiment of the present application. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention;
[0051] Embodiment 1
[0052] See Figure 1 As shown, the technical solution provided in this embodiment is a multi-modal antigen-antibody affinity prediction method based on the antibody structure, including the following steps:
[0053] S10: In this example, the Biomap dataset and the 14H and 14L affinity datasets are used respectively. The three datasets are divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and five-fold cross-validation is performed on these three datasets. Among them, the BioMap dataset comes from the antigen-antibody affinity prediction competition organized by BioMap, containing 1706 antigen-antibody pair data, with Delta G as the label. It includes 638 unique antigens and 1277 unique antibodies, most of which are from humans and mice, and a small number of components are from hamsters, chimpanzees, rhesus monkeys, rabbits, rats, and alpacas. The 14H and 14L datasets come from the LL-SARS-CoV-2 database, and each dataset contains 1 to 3 mutations in the CDR region. The original data lacking affinity information is deleted from the 14H and 14L datasets. Then, the affinity measurement values of all sequences with the same antibody sequence are averaged. Therefore, there are 13922 unique heavy chain data entries in the 14H dataset and 18708 unique light chain data entries in the 14L dataset. In the 14H dataset, only the heavy chain is variable, and the light chain and antigen remain unchanged in each entry. Similarly, in the 14L dataset, only the light chain changes, while the heavy chain and antigen remain unchanged in all entries.
[0054] S20: A multi-modal antibody information mining module is constructed. In the multi-modal antibody information mining module, the Roformer network is introduced to extract the heavy chain and light chain information of the antibody respectively, and the GearNet network is used to mine the structural information of the antibody. Subsequently, the sequence and structural information are adaptively fused by means of a cross-attention mechanism;
[0055] Step S21: The antibody sequence is divided into a heavy chain and a light chain and encoded by two Roformer networks pre-trained on a large number of antibody heavy chain sequences and light chain sequences respectively. Let the heavy chain sequence be represented as The light chain sequence is represented as where N H and N L represent the lengths of the heavy chain and light chain sequences respectively. The encoded representations of the heavy chain and light chain are as follows:
[0056] E H = Roformer H (H), E L = Roformer L (L) (1)
[0057] In the formula, Roformer H and Roformer L are Roformer encoders pre-trained on the heavy and light chain sequences respectively. After independently encoding the two chains, their representations are concatenated as:
[0058] E = Concat(E H , E L ) (2)
[0059] In the formula l s is the sequence length and d s is the embedding dimension. This method retains the independent information of the heavy and light chains while leveraging their collaborative contributions to binding specificity and affinity prediction through a unified representation.
[0060] Step S22: Mine antibody structure information. Predict the antibody structure using the antibody structure prediction tool IgFold. Secondly, an antibody graph is constructed, and the antibody structure is represented as a residue-level relational graph where and represent the sets of nodes and edges, and represents the set of edge types. Three types of edges are constructed for the residue-level antibody graph: If the sequential distance between nodes i and j is below a predefined threshold d seq , a sequential edge is added. If the Euclidean distance between nodes i and j is less than the threshold , an edge is added between nodes i and j. Finally, to account for different spatial scales between proteins, each node is connected to its k nearest neighbors based on the Euclidean distance.
[0061] Step S23: Extract structure information from the constructed antibody graph. Use one-hot encoding to construct the node features of the antibody graph, and construct the edge feature f(i, j, r) by concatenating the one-hot encoded node features of nodes i and j, the edge type, and their sequence and spatial distances. Based on the antibody graph , an edge graph is constructed using the GearNet-Edge encoder In the antibody graph , if edge (i, j) and edge (k, l) share a node, then in Add an edge from (i, j) to (k, l). Subsequently, the features of the edges in the constructed edge graph are updated through the edge message passing mechanism. The features of edge (i, j) in the graph are updated, and the message passing formula is defined as:
[0062]
[0063] In the formula represents the features of edge (i, j) at layer l, σ is the non-linear activation function ReLU, BN represents batch normalization, represents the neighborhood set of edge (i, j), Wr is the learnable weight matrix for edge type r. Through multi-layer message passing, the features of the edges are gradually updated to capture the multi-scale relationships between the edges in the antibody structure. After completing the edge message passing, the updated edge features are re-introduced into the antibody graph to enhance the representation ability of the nodes. By combining the edge messages, the feature update formula for node i is re-defined as:
[0064]
[0065] In the formula represents the features of node i at layer l, represents the neighbor set of node i, is the feature of neighbor node j from the previous layer, and FC refers to the fully connected layer that maps the edge features to the dimension matching the node features. Then, all the antibody node features are combined to form the structural embedding representation, denoted as where l t is the number of nodes in the structural embedding, and d t is the dimension of the structural features.
[0066] Step S24: After obtaining the antibody sequence and structure information, through the cross-attention module, the sequence representation and the structure representation are integrated interactively to extract key relevant features from both modalities. Then, the multi-modal output X is calculated as follows a :
[0067] Q = EW Q , K = UW K , V = UW V (5)
[0068]
[0069] In the formula, Q is the query matrix, K and V are the key and value matrices respectively, dk represents the dimension of the key vector. attention(Q, K, V) represents the calculation process of the attention mechanism. softmax(·) represents the softmax function, which is used to calculate the attention weights.
[0070] S30: Construct an antigen information mining module to extract antigen information by introducing the protein large language model ESM2;
[0071] Step S31: Extract antigen information through the pre-trained protein large prediction model ESM2. Assume the antigen sequence is represented as S = {a1, a2,..., a L}, where L is the length of the sequence, and a i represents the i-th amino acid residue. First, preprocess the sequence by converting each amino acid a i into its one-hot encoding representation. Add special markers at the beginning and end of the sequence respectively " <cls>” and " <eos>”, and padding the sequence to a fixed length to meet the model input requirements, obtaining the processed sequence S′:
[0072] S′ = {<cls>, a1, a2,..., a L , <eos>, <pad>,...} (7)
[0073] Step S32: Then input the processed sequence into the ESM2 model. ESM2 adopts 33 layers of Transformer encoders, using the multi-head attention mechanism to capture the dependencies between amino acid residues, thereby learning rich features of the antigen sequence. After passing through the encoding layer, the final output H (33) is a tensor of size (L + 2, 1280), where the vector at each position contains the depth feature information of the residue. To extract the overall feature of the antigen, a pooling operation is performed on all amino acid residue positions to obtain the final antigen feature vector X p :
[0074]
[0075] In the formula, L represents the length of the antigen sequence, denotes taking the average of the summation result, and H i 33 represents the feature vector of the i-th amino acid residue.
[0076] S40: Input the extracted features into the fusion prediction module and construct a multi-scale feature collaborative extraction network based on CNN. Input the multi-scale representations of antibodies and antigens into the fusion layer to generate a robust affinity prediction result;
[0077] Step S41: Input the antibody feature X a and antigen feature X p extracted by the encoder in Sections S2 and S3 into the prediction network to predict the affinity value. Use the multi-scale feature extraction convolutional neural network (MF_CNN) to integrate the information of antibodies and antigens. MF_CNN uses a 3-layer CNN backbone composed of convolution, pooling, and ReLU operations to extract multi-scale features. Subsequently, use the fully connected layer (FC) and residual operations to further combine these features to obtain the final output:
[0078] F a = Residual(FC(X a )) = FC(X a ) + X a (8)
[0079] F p = Residual(FC(X p )) = FC(X p )+X p (9)
[0080] In the formula, Fa and Fp are the characteristics of the antibody and antigen respectively, Residual represents the residual operation, and FC represents the fully connected layer.
[0081] Step S42: Transfer the integrated characteristics of the antibody and antigen to the multi-layer perceptron to finally generate the affinity value:
[0082] Affinity = MLP([Fa,Fp]) (10)
[0083] In the formula, MLP represents the multi-layer perceptron.
[0084] S50: Evaluate the prediction performance by comparing the error between the affinity value predicted by the model and the true value, and obtain the model parameters with the best prediction performance on the training set and the validation set;
[0085] S60: The data in the test set goes through the above steps S1 to S5 to obtain the final affinity value.
[0086] As can be seen from Tables 1, 2, and 3 below, the antigen-antibody affinity prediction method of this embodiment has better performance compared with other baseline models;
[0087] Table 1 Comparison of related work on the Biomap dataset. The values in parentheses represent the standard errors of Pearson and Spearman correlations.
[0088]
[0089]
[0090] Table 2 Comparison of related work on the 14H dataset. The values in parentheses represent the standard errors of Pearson and Spearman correlations.
[0091]
[0092] Table 3 Comparison of related work on the 14L dataset. The values in parentheses represent the standard errors of Pearson and Spearman correlations.
[0093]
[0094] The above embodiments merely represent the preferred embodiments of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.< / eos> < / cls> < / eos> < / cls>
Claims
1. A multi-modal antigen-antibody affinity prediction method based on antibody structure, characterized in that, It includes the following steps: S10: Collect the Biomap dataset, which contains 1706 antigen-antibody pairing data, including the heavy chain sequence, light chain sequence, antigen sequence, and affinity value of the antibody. Use the 14H and 14L affinity datasets. The heavy chains of the 14H dataset are different, and the light chains are constant. The 14L dataset is the opposite. And divide the affinity dataset into training, validation, and test parts; S20: Construct a multi-modal antibody information mining module. In the multi-modal antibody information mining module, introduce the Roformer network to extract the heavy chain and light chain information of the antibody respectively. At the same time, adopt the GearNet network to mine the structural information of the antibody, and use the cross-attention mechanism to adaptively fuse the sequence and structural information; S30: Construct an antigen information mining module to achieve antigen information extraction by introducing the protein large language model ESM2; S40: Input the extracted features into the fusion prediction module, and construct a multi-scale feature collaborative extraction network based on CNN. Input the multi-scale antibody and antigen representations into the fusion layer to generate a robust affinity prediction result; S50: Evaluate the prediction performance by comparing the error between the affinity value predicted by the model and the real value, and obtain the best model parameters on the training set and validation set; S60: The test set data goes through the above steps S1 to S5 to obtain the final affinity value.
2. The multimodal antigen-antibody affinity prediction method based on the antibody structure according to claim 1, wherein The step S2 includes the following steps: Step S21: Divide the antibody sequence into heavy and light chains, and encode them through two Roformer networks pre-trained on a large number of heavy and light chain sequences of antibodies respectively. Let the heavy chain sequence be represented as and the light chain sequence be represented as where N H and N L represent the lengths of the heavy and light chain sequences respectively. The encoded representations of the heavy and light chains are as follows: E H = Roformer H (H), E L = Roformer L (L) (1) Roformer in the formula H and Roformer L are Roformer encoders pre-trained on the heavy-chain and light-chain sequences respectively. After independently encoding the two chains, their representations are concatenated as follows: E = Concat(E H , E L ) (2) In the formula l s is the sequence length, and d s is the embedding dimension. This method retains the independent information of the heavy and light chains while leveraging their synergistic contributions to binding specificity and affinity prediction through a unified representation; Step S22: Predict the antibody structure using the antibody structure prediction tool IgFold. Secondly, an antibody graph is constructed, and the antibody structure is represented as a residue-level relationship graph where and ε represent the sets of nodes and edges, represents the set of edge types. Three types of edges are constructed for the residue-level antibody graph: if the sequential distance between nodes i and j is lower than the predefined threshold d seq , then add a sequential edge. If the Euclidean distance between nodes i and j is less than the threshold , then add an edge between nodes i and j, and each node is connected to its k nearest neighbors based on the Euclidean distance; Step S23: Extract structural information from the constructed antibody graph, construct node features of the antibody graph using one-hot encoding, and construct edge features f(i, j, r) by concatenating the one-hot encoded node features of nodes i and j, the edge type, and their sequence and spatial distances. Based on the antibody graph an edge graph is constructed In the antibody graph if edge (i, j) and edge (k, l) share a common node, then in an edge is added from (i, j) to (k, l). Subsequently, the features of edge (i, j) in the constructed edge graph are updated through the edge message passing mechanism. The message passing formula is defined as: In the formula represents the feature of edge (i, j) at layer l, σ is the non-linear activation function ReLU, and BN represents batch normalization. represents the neighborhood set of edge (i, j). Wr is the learnable weight matrix for edge type r. Through multi-layer message passing, the features of the edge are gradually updated to capture the multi-scale relationships between edges in the antibody structure. After completing the edge message passing, the updated edge features are re-introduced into the antibody graph. Enhance the representation ability of nodes. By combining edge messages, the feature update formula for node i is redefined as: In the formula represents the feature of node i in layer l, represents the neighbor set of node i, is the feature of neighbor node j from the previous layer, and FC refers to the fully connected layer that maps the edge feature to the dimension matching the node feature. Then, all the antibody node features are combined to form the structural embedding representation, denoted as where l t is the number of nodes in the structural embedding, and d t is the dimension of the structural feature; Step S24: After obtaining the antibody sequence and structural information, through the cross-attention module, the sequence representation and the structural representation are integrated interactively to extract key relevant features from both modalities, and the multimodal output X is calculated as follows a : Q = EW Q , K = UW K , V = UW V (5) In the formula, Q is the query matrix, K and V are the key and value matrices respectively, dk represents the dimension of the key vector, attention(Q, K, V) represents the calculation process of the attention mechanism, and softmax(·) represents the softmax function, which is used to calculate the attention weight.
3. A method for predicting the affinity of a multimodal antigen-antibody based on the antibody structure according to claim 1, characterized in that, The step S3 includes the following steps: Step S31: Extract antigen information through the pre-trained large protein prediction model ESM2. Assume the antigen sequence is represented as S = {a1, a2,..., a L}, where L is the length of the sequence, and a i represents the i-th amino acid residue. The sequence is preprocessed by converting each amino acid a i into its one-hot encoded representation, and special markers are added at the beginning and end of the sequence, respectively, " <cls>”and" <eos>”, and pad the sequence to a fixed length to meet the model input requirements to obtain the processed sequence S′:< / eos> < / cls> S′ = {<cls>, a1, a2,..., a L , <eos>, <pad>,...} (7) Step S32: Then, the processed sequence is input into the ESM2 model. ESM2 adopts 33 layers of Transformer encoders, uses the multi-head attention mechanism to capture the dependencies between amino acid residues, learns rich features of the antigen sequence, and finally outputs H after passing through the encoding layer. (33) is a tensor of size (L + 2, 1280), where the vector at each position contains the depth feature information of the residue. To extract the overall features of the antigen, a pooling operation is performed on all amino acid residue positions to obtain the final antigen feature vector X. p : In the formula, L represents the length of the antigen sequence, represents taking the average of the summation result, represents the feature vector of the i-th amino acid residue.
4. A method for predicting the affinity of a multimodal antigen-antibody based on the antibody structure according to claim 1, characterized in that, The step S4 includes the following steps: Step S41: Input the antibody feature X a extracted by the encoder in Sections S2 and S3 p and the antigen feature X into the prediction network to predict the affinity value. Integrate the information of the antibody and the antigen using a multi-scale feature extraction convolutional neural network (MF_CNN). MF_CNN uses a 3-layer CNN backbone composed of convolutional, pooling, and ReLU operations to extract multi-scale features, and combines these features using a fully connected layer and residual operations to obtain the final output: F a = Residual(FC(X a )) = FC(X a ) + X a (8) F p = Residual(FC(X p )) = FC(X p ) + X p (9) In the formula, Fa and Fp are the features of the antibody and antigen respectively, Residual represents the residual operation, and FC represents the fully connected layer; Step S42: Transmit the integrated features of the antibody and antigen to the multi-layer perceptron to finally generate the affinity value: Affinity = MLP([Fa, Fp]) (10) In the formula, MLP represents the multi-layer perceptron.