Method and device for predicting interaction between nano antibody and antigen

By using the protein big language model ESM-2 and the antigen binding site prediction model DeepNano-site, combined with feature fusion technology of cue encoder, the problem of insufficient feature utilization in the prediction of nanobody-antigen interaction was solved, the prediction performance and robustness were improved, and the computer-aided design and screening of nanobody drug design was promoted.

CN121999855APending Publication Date: 2026-05-08TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2024-11-01
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing sequence-based protein interaction prediction methods suffer from negative transfer problems in predicting nanobody-antigen interactions and fail to effectively utilize the features of unsupervised training protein big language models, neglecting the importance of amino acid sites at the binding interface, leading to a decline in prediction performance.

Method used

The protein big language model ESM-2 was used to extract sequence features of nanobodies and antigens. Initial interaction probabilities were generated through minimum pooling, average pooling and maximum pooling strategies. Combined with the antigen binding site prediction model DeepNano-site, attention representations of antigen binding sites were extracted using a cue encoder and fused with sequence features to update the interaction probabilities to obtain the final prediction results.

Benefits of technology

It significantly improves the predictive performance of nanobody-antigen interactions, enhances the robustness and generalization ability of the model, shortens the development cycle of nanobody new drugs, and reduces development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999855A_ABST
    Figure CN121999855A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information, in particular to a nanometer antibody and antigen interaction prediction method and device.The method comprises the steps that sequence characteristics of nanometer antibodies and antigens are obtained on the basis of a protein large language model, so that the initial nanometer antibody and antigen interaction probability is generated, and an antigen binding site prediction model is used for predicting the nanometer antibody and antigen interaction probability. Determining an estimated antigen binding site of the sequence features of the nano antibody and the antigen, extracting a corresponding target feature, fusing the target feature with the obtained sequence features, and updating the initial interaction probability of the nano antibody and the antigen by using the fused feature, so as to obtain a final prediction result of the interaction of the nano antibody and the antigen. According to the method provided by the invention, the importance of amino acid sites on a binding interface is highlighted when the interaction between the nano antibody and the antigen is predicted, and the prediction performance of the interaction between the nano antibody and the antigen is remarkably improved by the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of bioinformatics, artificial intelligence, and biomedicine, and in particular to a method and apparatus for predicting the interaction between nanobodies and antigens. Background Technology

[0002] Several studies have applied molecular dynamics or machine learning methods to predict nanobody-antigen interactions (NAIs). However, most of these methods require precise structures of both nanobodies and antigens, limiting their application in predicting NAI interactions. With the development of next-generation sequencing technology, protein sequence information can now be obtained rapidly and at low cost, making the development of efficient sequence-based NAI prediction methods more urgent. Technically, both nanobodies and antigens are essentially proteins, and existing sequence-based protein interaction prediction methods can be helpful for NAI prediction. However, applying conventional protein interaction prediction methods to NAI prediction often results in negative transfer, possibly due to significant pattern differences between nanobody sequences and general protein sequences.

[0003] The sequence-structure-function paradigm posits that a protein's amino acid sequence determines its spatial structure, which in turn determines its function. Sequence-based protein interaction prediction methods first require protein sequence characterization. Past research has shown that protein characterization obtained through large language models has achieved superior performance in various tasks, including protein-protein binding affinity prediction. Current state-of-the-art sequence-based protein interaction prediction methods, such as D-SCRIPT and Topsy-Turvy, both utilize large protein language models in their model design.

[0004] However, most studies in related technologies directly use protein embedding features pre-trained by language models as input to protein interaction prediction models, without making better use of these unsupervised training protein features. In addition, current sequence-based protein interaction prediction methods always treat all sites on the protein sequence as equally important in the model input, which reduces the prediction performance of sequence-based protein interaction prediction methods. Summary of the Invention

[0005] This application is based on the inventor's understanding and insights into the following issues:

[0006] Nanobodies are protein fragments extracted from the variable domains of heavy-chain antibodies unique to camels and sharks. Compared to traditional monoclonal antibodies, nanobodies not only retain the ability to specifically bind antigens but also have smaller molecular weights, lower immunogenicity, and stronger tissue penetration. Nanobodies can form various non-covalent bonds with antigens, and the nanobody-antigen interaction (NAI) problem is an important branch of the protein-protein interaction (PPI) problem. Its research is of great significance for elucidating immune mechanisms and designing nanobodies de novo. In recent years, the development of nanobodies has been rapid, and they have been widely used in detection and treatment. Public databases related to nanobodies are also continuously being released, which promotes research on related methods. However, existing methodological research mainly focuses on the structural prediction, naturalness assessment, or binding site prediction of nanobodies. Currently, few studies apply deep learning methods to NAI prediction, although this research direction has begun to gain attention in the fields of biomedicine, bioinformatics, and artificial intelligence in recent years.

[0007] The most crucial characteristic of nanobodies, as targeted protein drugs, is their ability to specifically bind to target antigens. The regions where antibodies and antigens interact are called paratopes and epitopes, respectively. Whether nanobodies or traditional IgG (Immunoglobulin G), the paratopes are generally located in the CDRs (complementarity determining regions) of the V region. The difference lies in that traditional IgG recognizes antigens through the combined recognition of heavy and light chain CDRs, while nanobodies only have a heavy chain and therefore rely solely on heavy chain CDRs for antigen recognition. To compensate for the reduced sequence diversity caused by the absence of the light chain, the CDR3 region of nanobodies is longer than that of traditional IgG. Therefore, although nanobodies are smaller than traditional IgG, they can still specifically bind to various types of antigens. In recent years, the prediction of interactions between monoclonal antibodies and antigens has received some attention and research, but the prediction of interactions between nanobodies and antigens is currently lacking. Studying the binding modes of nanobodies to antigens and constructing interaction prediction models is a key technological gap in the current design of anticancer nanobody drugs. Addressing this technological gap can directly advance the computer-aided design and screening of nanobodies, which will greatly reduce the R&D costs of new nanobodies and significantly shorten the R&D cycle.

[0008] This application provides a method and apparatus for predicting nanobody-antigen interactions. When predicting nanobody-antigen interactions, the method highlights the importance of amino acid sites at the binding interface. In contrast, existing protein interaction prediction methods generally treat all sites on the protein sequence as equally important. The nanobody-antigen interaction prediction method proposed in this application significantly improves the predictive performance of nanobody-antigen interactions.

[0009] The first aspect of this application provides a method for predicting the interaction between nanobody and antigen, comprising the following steps: extracting features of target amino acid sequences based on a protein big language model to obtain sequence features of nanobody and antigen that meet preset conditions; generating an initial nanobody-antigen interaction probability using the sequence features of the nanobody and antigen, and predicting antigen binding sites based on the sequence features of the nanobody and antigen using an antigen binding site prediction model to obtain an estimated binding site of the antigen; extracting target features of the estimated binding site of the antigen using a preset prompt encoder, fusing the target features with the sequence features of the nanobody and antigen to obtain fused features, and updating the initial nanobody-antigen interaction probability using the fused features to obtain the final prediction result of the nanobody-antigen interaction.

[0010] Optionally, in one embodiment of this application, the step of extracting amino acid sequence features based on the protein big language model to obtain sequence features of nanobodies and antigens that meet preset conditions includes: representing each amino acid in the sequence using the embedding vector output in the last hidden layer of the protein big language model; and performing targeted pooling processing on the features of each amino acid based on the minimum pooling strategy, average pooling strategy, and maximum pooling strategy in the target pooling strategy to obtain sequence features of the nanobodies and antigens that meet preset conditions.

[0011] Optionally, in one embodiment of this application, generating the initial nanobody-antigen interaction probability using the sequence features of the nanobody and the antigen includes: constructing a first predicted interaction probability based on a minimum pooling strategy, a second predicted interaction probability based on an average pooling strategy, and a third predicted interaction probability based on a maximum pooling strategy, respectively, based on the sequence features of the nanobody and the antigen; and generating the initial nanobody-antigen interaction probability using the first predicted interaction probability, the second predicted interaction probability, and the third predicted interaction probability.

[0012] Optionally, in one embodiment of this application, the step of using an antigen binding site prediction model to predict the antigen binding site based on the sequence features of the nanobody and the antigen, and obtaining the estimated binding site of the antigen, includes: performing average pooling processing on the features of each residue of the nanobody based on the antigen binding site prediction model to obtain the final features of the nanobody; splicing the final features of the nanobody with the features of each antigen residue to obtain the spliced ​​features; and inputting the spliced ​​features into a target neural network composed of multiple residual blocks to obtain the estimated binding site of the antigen.

[0013] Optionally, in one embodiment of this application, the step of extracting target features of the estimated binding site of the antigen using a preset cue encoder and fusing the target features with the sequence features of the nanobody and the antigen to obtain fused features includes: obtaining an attention representation of the estimated binding site of the antigen using the preset cue encoder; determining an attention embedding of the antigen binding site based on the attention representation; and fusing the attention embedding into the sequence features of the nanobody and the antigen to obtain the fused features.

[0014] Optionally, in one embodiment of this application, the initial formula for calculating the interaction probability between the nanobody and the antigen is:

[0015]

[0016] Wherein, X1 and X2 represent the amino acid sequences of the nanobody and the antigen, respectively; Represents the feature mapping function of ESM-2; P represents the i-th multilayer perceptron; min P mean and P max These represent three pooling strategies (min pooling, average pooling, and max pooling).

[0017] A second aspect of this application provides a device for predicting nanobody-antigen interactions, comprising: an extraction module for extracting features of a target amino acid sequence based on a protein big language model to obtain sequence features of nanobody and antigen that meet preset conditions; a generation module for generating an initial nanobody-antigen interaction probability using the sequence features of the nanobody and antigen, and predicting an antigen binding site based on the sequence features of the nanobody and antigen using an antigen binding site prediction model to obtain an estimated binding site of the antigen; and a prediction module for extracting target features of the estimated binding site of the antigen using a preset prompt encoder, fusing the target features with the sequence features of the nanobody and antigen to obtain fused features, and updating the initial nanobody-antigen interaction probability using the fused features to obtain a final prediction result of the nanobody-antigen interaction.

[0018] Optionally, in one embodiment of this application, the extraction module includes: a characterization unit, used to characterize each amino acid on the sequence using the embedding vector output from the last hidden layer in the protein big language model; and an acquisition unit, used to perform target pooling processing on the features of each amino acid based on the minimum pooling strategy, average pooling strategy, and maximum pooling strategy of the target pooling strategy, so as to obtain the sequence features of the nanobody and antigen that meet the preset conditions.

[0019] Optionally, in one embodiment of this application, the generation module includes: a construction unit, configured to construct, based on the sequence characteristics of the nanobody and the antigen, a first predicted interaction probability based on a minimum pooling strategy, a second predicted interaction probability based on an average pooling strategy, and a third predicted interaction probability based on a maximum pooling strategy; and a generation unit, configured to generate the initial nanobody and antigen interaction probability using the first predicted interaction probability, the second predicted interaction probability, and the third predicted interaction probability.

[0020] Optionally, in one embodiment of this application, the generation module includes: a processing unit, configured to perform average pooling processing on the features of each residue of the nanobody based on the antigen binding site prediction model to obtain the final features of the nanobody; a splicing unit, configured to splice the final features of the nanobody with the features of each antigen residue to obtain the spliced ​​features; and a first determining unit, configured to input the spliced ​​features into a target neural network composed of multiple residual blocks to obtain the estimated binding site of the antigen.

[0021] Optionally, in one embodiment of this application, the prediction module includes: a first acquisition unit, configured to acquire the attention representation of the estimated antigen binding site using the preset cue encoder; a second determination unit, configured to determine the attention embedding of the antigen binding site based on the attention representation; and a second acquisition unit, configured to fuse the attention embedding into the sequence features of the nanobody and the antigen to obtain the fusion feature.

[0022] Optionally, in one embodiment of this application, the initial formula for calculating the interaction probability between the nanobody and the antigen is:

[0023]

[0024] Wherein, X1 and X2 represent the amino acid sequences of the nanobody and the antigen, respectively; Represents the feature mapping function of ESM-2; P represents the i-th multilayer perceptron; min P mean and P max These represent three pooling strategies (min pooling, average pooling, and max pooling).

[0025] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a method for predicting the interaction between a nanobody and an antigen as described in the above embodiments.

[0026] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the interaction between a nanobody and an antigen.

[0027] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, is used to implement the above-described method for predicting the interaction between a nanobody and an antigen.

[0028] This application's embodiments can extract features from the target amino acid sequence based on a protein big language model to obtain the sequence features of nanobodies and antigens, thereby generating an initial nanobodies-antigen interaction probability. Then, using an antigen binding site prediction model, the estimated binding sites of the antigen based on the sequence features of the nanobodies and antigens are determined, thereby extracting the corresponding target features. These features are then fused with the sequence features of the nanobodies and antigens to obtain fused features. Finally, the initial nanobodies-antigen interaction probability is updated using these fused features to obtain the final prediction result of the nanobodies-antigen interaction. This solves the problem in protein interaction-related technologies where unsupervised pre-trained protein big language models lack efficient methods for utilizing protein sequence representations, and where all sites on the protein sequence are treated as equally important while ignoring the importance of amino acid sites at the binding interface, thus reducing the prediction performance of nanobodies-antigen interactions.

[0029] The additional aspects and advantages of this application will be further described below. Attached Figure Description

[0030] The above and / or additional aspects and advantages of this application can be further explained in conjunction with the accompanying drawings and description of the embodiments, wherein:

[0031] Figure 1 This is a flowchart of a method for predicting the interaction between a nanobody and an antigen, based on an embodiment of this application.

[0032] Figure 2 This is a framework diagram of a nanobody and antigen interaction prediction method in a specific embodiment of this application;

[0033] Figure 3 This is a schematic diagram of a model for predicting antigen binding sites based on antigen and nanobody sequences in a specific embodiment of this application;

[0034] Figure 4 This is a schematic diagram illustrating the testing of NAI data used in a study using models (D-SCRIPT, Topsy-Turvy, and DeepNano-seq) trained on human PPI data, as described in a specific embodiment of this application.

[0035] Figure 5 This is a schematic diagram illustrating the analysis of the full-length sequence similarity of most nanobodies in a specific embodiment of this application;

[0036] Figure 6 This is a schematic diagram illustrating the filtering operation of 2422 nanobody-antigen binding pairs downloaded from the SAbDab (Single-domain Antibody Database) nano database in a specific embodiment of this application.

[0037] Figure 7 This is a schematic diagram comparing the NAI prediction performance of two DeepNano-seq prediction models trained on the SAbDab-nano dataset and the human PPI dataset in a specific embodiment of this application.

[0038] Figure 8 This is a quantitative schematic diagram showing the proportion of amino acids to the full-length sequence at the binding interface of a nanobody-antigen in a specific embodiment of this application.

[0039] Figure 9 This is a quantitative schematic diagram of the amino acid types at the binding interface of a nanobody-antigen in a specific embodiment of this application.

[0040] Figure 10 This diagram illustrates a comparison of the NAI prediction metrics AUPRC (Area Under the Precision-Recall Curve), AUROC (Area Under the Receiver Operating Characteristic), accuracy, recall, precision, and F1-score for DeepNano in one specific embodiment of this application and DeepNano-seq in another specific embodiment.

[0041] Figure 11 This is a distribution of prediction scores for positive nanobodies and one million background nanobodies using the DeepNano-seq model in a specific embodiment of this application.

[0042] Figure 12 This is a distribution graph of the predicted scores and experimentally measured ELISA (Enzyme-Linked Immunosorbent Assay) values ​​of 59 anti-GST (Glutathione S-transferase) nanobodies in a specific embodiment of this application.

[0043] Figure 13 This is a schematic diagram of a nanobody-antigen interaction prediction device provided according to an embodiment of this application;

[0044] Figure 14 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0045] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0046] The following description, with reference to the accompanying drawings, illustrates a method and apparatus for predicting nanobody-antigen interactions according to embodiments of this application. Addressing the issues mentioned in the background section regarding the lack of efficient utilization of protein sequence characterization obtained from unsupervised pre-trained protein large language models, and the problem that all sites on the protein sequence are treated as equally important while neglecting the importance of amino acid sites at the binding interface, thus reducing the predictive performance of nanobody-antigen interactions, this application provides a method for predicting nanobody-antigen interactions. In this method, features of the target amino acid sequence can be extracted based on a protein large language model to obtain the sequence features of the nanobody and antigen, thereby generating an initial probability of nanobody-antigen interaction. An antigen binding site prediction model is then used to determine the estimated antigen binding sites of the sequence features of the nanobody and antigen, thereby extracting the corresponding target features. These features are then fused with the sequence features of the nanobody and antigen to obtain fused features. Finally, the initial probability of nanobody-antigen interaction is updated using the fused features to obtain the final prediction result of the nanobody-antigen interaction. This addresses the problem that unsupervised pre-trained protein large language models lack efficient methods for utilizing protein sequence representations, and that treating all sites on the protein sequence as equally important while neglecting the importance of amino acid sites at the binding interface reduces the predictive performance of nanobody-antigen interactions.

[0047] Specifically, Figure 1 This is a flowchart illustrating a method for predicting the interaction between a nanobody and an antigen, as provided in an embodiment of this application.

[0048] like Figure 1 As shown, this method for predicting the interaction between a nanobody and an antigen includes the following steps:

[0049] In step S101, based on the protein big language model, the features of the target amino acid sequence are extracted to obtain the sequence features of nanobodies and antigens that meet the preset conditions.

[0050] In this embodiment of the application, the protein big language model is the protein big language model ESM-2; the preset conditions are the conditions for implementing three different pooling strategies.

[0051] It is understood that embodiments of this application can extract features of the target amino acid sequence based on a protein big language model, for example, such as... Figure 2As shown in part a, embodiments of this application can set up a deep ensemble learning framework, DeepNano-seq. DeepNano-seq uses the Protein Large Language Model (ESM-2) to extract features from amino acid sequences. It employs three different pooling strategies (min pooling, average pooling, and max pooling) to obtain sequence features of nanobodies and antigens, effectively improving the feasibility of predicting the interaction between nanobodies and antigens.

[0052] In one embodiment of this application, features of amino acid sequences are extracted based on a protein big language model to obtain sequence features of nanobodies and antigens that meet preset conditions. This includes: representing each amino acid in the sequence using the embedding vector output from the last hidden layer in the protein big language model; and performing target pooling processing on the features of each amino acid based on the minimum pooling strategy, average pooling strategy, and maximum pooling strategy of the target pooling strategy to obtain sequence features of nanobodies and antigens that meet preset conditions.

[0053] In this embodiment of the application, the powerful protein big language model ESM-2 is used to obtain protein features. ESM-2 was first applied to protein structure prediction and has recently been tried for some downstream applications, such as protein-ligand complex structure prediction, protein-nucleic acid binding site prediction and drug target interaction prediction.

[0054] In actual implementation, this application uses the embedding vector output by the last hidden layer of ESM-2 to characterize each amino acid in the sequence. In order to obtain protein characterization of the same dimension, this application performs pooling operation on all amino acid features in the sequence direction. For example, three pooling strategies (average pooling, max pooling and min pooling) are used to obtain protein characterization with more information, thereby obtaining the sequence features of nanobodies and antigens.

[0055] In step S102, the initial interaction probability between the nanobody and the antigen is generated using the sequence characteristics of the nanobody and the antigen, and the antigen binding site prediction model is used to predict the antigen binding site based on the sequence characteristics of the nanobody and the antigen, so as to obtain the estimated binding site of the antigen.

[0056] In the embodiments of this application, the antigen binding site prediction model is a model used for predicting the interaction between nanobodies and antigens, such as the DeepNano-site model.

[0057] It is understood that the embodiments of this application can utilize the sequence features of nanobodies and antigens to generate initial nanobodies-antigen interaction probabilities. For example, based on the sequence features of nanobodies and antigens, the embodiments of this application can construct three independent branches to predict the interaction probabilities using different features obtained through three pooling strategies via DeepNano-seq, and use the average of the prediction results of the three branches to obtain the initial nanobodies-antigen interaction probabilities. Then, an antigen binding site prediction model, such as the DeepNano-site model, is used to predict the antigen binding sites in the sequence features of nanobodies and antigens to obtain the estimated binding sites of the antigen, effectively improving the robustness of the prediction.

[0058] For example, in this embodiment, considering the small molecular weight of nanobodies (approximately 15 kDa), their antigen-binding region is much smaller than the entire antigen length, especially for large molecular antigens. If the model could identify the sites on the full-length antigen sequence that directly influence the interaction, it might achieve more robust predictive performance. Therefore, this embodiment designs a new computational pipeline, such as... Figure 2 As shown in section b, in this embodiment of the application, the DeepNano-site model is used to predict antigen binding sites from antigen and nanobody sequences, thereby obtaining the estimated binding sites of the antigen.

[0059] Optionally, in one embodiment of this application, generating initial nanobody-antigen interaction probabilities using the sequence features of the nanobody and the antigen includes: based on the sequence features of the nanobody and the antigen, generating a first predicted interaction probability based on a minimum pooling strategy, a second predicted interaction probability based on an average pooling strategy, and a third predicted interaction probability based on a maximum pooling strategy; and generating initial nanobody-antigen interaction probabilities using the first predicted interaction probability, the second predicted interaction probability, and the third predicted interaction probability.

[0060] In some embodiments, this application can utilize the sequence characteristics of nanobodies and antigens, and then use DeepNano-seq to obtain different features through three pooling strategies. Next, a first predicted interaction probability based on a minimum pooling strategy, a second predicted interaction probability based on an average pooling strategy, and a third predicted interaction probability based on a maximum pooling strategy are constructed respectively. Then, the average value of the first predicted interaction probability, the second predicted interaction probability, and the third predicted interaction probability is calculated to obtain the initial interaction probability between the nanobodies and antigens.

[0061] In one embodiment of this application, the initial formula for calculating the interaction probability between the nanobody and the antigen is as follows:

[0062]

[0063] Wherein, X1 and X2 represent the amino acid sequences of the nanobody and the antigen, respectively; Represents the feature mapping function of ESM-2; P represents the i-th multilayer perceptron; min P mean and P max These represent three pooling strategies (min pooling, average pooling, and max pooling).

[0064] The formula for calculating the loss function during training of the DeepNano-site model is as follows:

[0065]

[0066] Where g represents the truth value and CE represents the cross-entropy loss function.

[0067] For example, embodiments of this application may employ a decision-side fusion method, that is, using protein representations obtained from different pooling strategies to predict interactions between proteins, rather than directly splicing the three representations together. Figure 2 (Part a of the document). In end-to-end learning, this application fine-tunes the last hidden layer of ESM-2 and freezes the remaining layers. Due to GPU (Graphics Processing Unit) memory limitations, this application does not perform full parameter fine-tuning of ESM-2. Currently, there are six publicly available ESM-2 models with different parameter scales (ranging from 8 million parameters to 15 billion parameters). This application uses the lightest 8M parameter model to demonstrate the advanced nature of the model structure design in this application.

[0068] Optionally, in one embodiment of this application, an antigen binding site prediction model is used to predict the antigen binding site based on the sequence features of the nanobody and the antigen, thereby obtaining the estimated binding site of the antigen. This includes: using the antigen binding site prediction model to perform average pooling processing on the features of each residue of the nanobody to obtain the final features of the nanobody; splicing the final features of the nanobody with the features of each antigen residue to obtain the spliced ​​features; and inputting the spliced ​​features into a target neural network composed of multiple residual blocks to obtain the estimated binding site of the antigen.

[0069] As one possible approach, the antigen-binding interface is the region in the antigen sequence that contributes the most to the interaction. However, obtaining the precise location of the antigen-binding interface is both expensive and time-consuming. Therefore, embodiments of this application design a model for predicting antigen-binding sites based on antigen and nanobody sequences, such as... Figure 3As shown, this embodiment still uses ESM-2 to encode the amino acid sequence. Unlike DeepNano-seq, this embodiment only performs average pooling on the features of each residue of the nanobody, without pooling on the characterization of each residue of the antigen. The final features of the nanobody are combined with the features of each antigen residue to obtain the spliced ​​features. The spliced ​​features are then input into a neural network composed of multiple residual blocks to predict whether each antigen residue belongs to a binding site. This allows for the estimation of the antigen binding site, effectively improving the feature expression ability and enhancing the model's generalization ability and prediction accuracy.

[0070] In step S103, the target features of the estimated binding site of the antigen are extracted using a preset prompt encoder, and the target features are fused with the sequence features of the nanobody and the antigen to obtain fused features. The initial interaction probability of the nanobody and the antigen is updated using the fused features to obtain the final prediction result of the interaction between the nanobody and the antigen.

[0071] It is understood that embodiments of this application can utilize a cue encoder to extract target features of the estimated binding site of the antigen, and fuse these target features with the sequence features of the nanobody and the antigen to obtain fused features. For example, embodiments of this application can extract features from the estimated antigen binding site and incorporate them into DeepNano-seq, thereby enhancing DeepNano-seq's focus on the antigen binding interface. Figure 2 In part c, DeepNano characterizes antigen binding sites by constructing a transformer-based encoder and adds it to DeepNano-seq. It then uses fusion features to update the initial nanobody-antigen interaction probabilities to obtain the final prediction results of nanobody-antigen interactions, thereby enhancing the prediction of NAI.

[0072] Optionally, in one embodiment of this application, a preset cue encoder is used to extract target features of the estimated binding site of the antigen, and the target features are fused with the sequence features of the nanobody and the antigen to obtain fused features. This includes: using the preset cue encoder to obtain attention representation of the estimated antigen binding site; determining attention embedding of the antigen binding site based on the attention representation; and fusing the attention embedding into the sequence features of the nanobody and the antigen to obtain fused features.

[0073] For example, such as Figure 2As shown in section c, this embodiment of the application designs a cue encoder to obtain the attention representation of the antigen-binding interface, reflecting which part of the antigen contributes the most to the interaction. The backbone of the cue encoder is a transformer encoder layer. The cue encoder takes the 0 / 1 vectors representing antigen-binding sites as input. Each token (0 / 1) is embedded in an 8-dimensional vector, which is optimized during training. After the token embedding, this embodiment of the application adds a learnable position encoding layer to introduce information on the relative positions of all tokens in the sequence. The transformer encoder layer of the cue encoder takes the sum of position features and label features as input and finally outputs the representation of all sites on the antigen.

[0074] In other words, the embodiments of this application adopt an average pooling strategy to obtain attentional representation of the entire antigen-binding interface. Based on this, a cue encoder is used to obtain attentional embedding of the antigen-binding site and add it to the sequence features output by ESM-2, that is, integrate it into the sequence features of the nanobody and the antigen to obtain fusion features, thereby obtaining additional knowledge of the antigen-binding site. As a result, DeepNano can make more reliable interaction predictions.

[0075] For example, this application collected a human PPI dataset and four NAI datasets. The human PPI dataset came from research on the D-SCRIPT method. It should be noted that in the D-SCRIPT research, due to GPU memory limitations, all PPIs with a protein length exceeding 800 amino acids were removed. This application obtained 3481 NAI datasets from a recent nanobody study. To fairly test the PPI method, this application removed NAI datasets with a sequence length exceeding 800 amino acids, ultimately obtaining 1800 NAIs as an independent test dataset. The training NAI data came from the SAbDab-nano database, a sub-database of SAbDab containing the structures of all nanobody-antigen complexes in the PDB.

[0076] Next, this application filtered out some obvious false positive instances, leaving 1184 NAIs as positive data, resulting in a final NAI training dataset containing 11209 entries. In addition to the NAI training and testing datasets, this application also collected 33 experimentally validated anti-HSA (human serum albumin) nanobodies and 59 anti-GST (glutathione S-transferase) nanobodies from a recent study. The sequences for GST and HSA were obtained from the UniProt database. To simulate a virtual screening process, this application randomly extracted one million natural nanobody sequences from the INDI database as background. These background nanobodies were then mixed with the anti-HSA and anti-GST nanobodies to construct two additional independent test datasets. Generally, the vast majority of background nanobodies cannot bind to HSA or GST; therefore, the background nanobodies were labeled negative when calculating model performance metrics.

[0077] In this study, multiple models were trained sequentially in this application. First, DeepNano-seq was trained on human PPIs and compared with other PPI methods. DeepNano-seq is a deep ensemble network with three branches that independently output predicted scores of interactions. During training, the losses of these three predicted scores were calculated separately and summed before backpropagation. Second, DeepNano-seq was retrained on prepared NAI data and compared with the version previously trained on human PPIs to verify the performance improvement in predicting NAIs. Third, DeepNano-site was trained on NAIs, and this model was used as only one component of DeepNano. Finally, the weights of DeepNano-site were frozen, and DeepNano was trained on NAIs. DeepNano was then compared with the two versions of DeepNano-seq and all state-of-the-art PPI methods.

[0078] Finally, training the DeepNano-seq model based on the lightest 8M parameter ESM-2 took approximately 12 hours for ten epochs on 421,792 human PPIs, approximately 10 minutes for ten epochs on 11,209 NAIs, and approximately 40 minutes for ten epochs on NAIs. Of these, 20 minutes were used to train the DeepNano-site with a learning rate set to 0.00005. In this embodiment, the checkpoint that performs best on the validation set can be selected as the final model, and all experiments were conducted on a Linux system configured with a 40GB GPU.

[0079] Those skilled in the art should understand that nanobodies are a special type of antibody obtained from camels, and are small proteins, typically composed of no more than 149 amino acids. Nanobodies can specifically bind to antigens, and since both antigens and nanobodies are proteins, the interaction between nanobodies and antigens may have some undiscovered connection to general protein-protein interactions. To evaluate the generalization performance of PPI methods for NAI prediction, this application uses models trained on human PPI data (D-SCRIPT, Topsy-Turvy, and DeepNano-seq) to test the NAI data used in the study. Figure 4 As can be seen, all three methods have poor predictive performance for NAI data.

[0080] Models trained on human PPIs predicted almost all true NAI pairs as negative, resulting in near-zero recall metrics (D-SCRIPT: 0.02; Topsy-Turvy: 0.02; DeepNano-seq: 0). One possible reason for this is the significant difference between camel nanobodies and human proteins. Specifically, such as... Figure 5 As shown, most nanobodies exhibit full-length sequence similarity exceeding 70% because most nanobodies are highly conserved in the framework regions (FRs). The main sequence differences between nanobodies lie in the complementarity-determining regions (CDRs), particularly CDR3. However, studies using the D-SCRIPT method employed a lower sequence similarity threshold (40%) to remove redundancy when processing human PPI training data. This resulted in models trained on low-resolution human PPI data exhibiting poor resolution of nanobodies. Therefore, more NAI data should be collected to establish nanobodies-specific interaction prediction models.

[0081] A recent study applying machine learning to NAI prediction used NAI data collected from the sdAd-DB database released in 2018; this application also uses these data as independent test datasets. To obtain more NAI data for model training, this application downloaded 2422 nanobody-antigen binding pairs from the SAbDab-nano database, such as... Figure 6As shown, the SAbDab-nano database is a subset of the SAbDab structural antibody database updated in 2021. After a series of sample filtering operations to ensure data accuracy, this application ultimately obtained 1184 pairs of positive NAI data. Analysis revealed that these NAI pairs did not have the same sequences as the independent test set. This application randomly divided these binding pairs into training and validation sets. The validation set was used for an early stopping strategy to prevent model overfitting. This application retrained DeepNano-seq using the newly collected NAI data and tested it again on the independent NAI test set. Figure 7 In comparison with the previous version trained using human PPI data, the new training version of DeepNano-seq showed a significant performance improvement. The AUROC (Area Under the Receiver Operating Characteristic Curve) of DeepNano-seq increased from 0.5542 to 0.6596, and the AUPRC increased from 0.3571 to 0.6343. In addition, the recall value of DeepNano-seq also increased from 0 to 0.3810.

[0082] The SAbDab-nano database provides not only sequences of nanobodies and antigens, but also structural information on nanobody-antigen complexes. Combined with... Figure 8 and Figure 9 As shown, this application quantifies the quantity and type of amino acids at the nanobody-antigen binding interface, revealing different amino acid type preferences at the nanobody-antigen binding interface. Tyr and Ser are very common at the nanobody binding interface, while they appear less frequently at the antigen binding interface. Figure 9 Given that any protein can potentially become an antigen, this difference also highlights the distinction between PPI and NAI. Figure 8 In the study, it was observed that most antigen-binding interfaces comprised no more than 30% of the entire antigen sequence. For nanobodies, the proportion of binding interfaces was also relatively small. However, existing analyses show that most nanobodies' binding sites are distributed across only three core receptors (CDRs), which was also directly observed in Example 5F9D. Given the conservation of nanobody sequences, data-driven models are more likely to focus on CDRs and nanobody binding interfaces. While the antigen-binding interface contributes the most to the interaction, its proportion in the input is relatively small. Therefore, prioritizing the antigen-binding interface in modeling is crucial; however, DeepNano-seq and four other PPI methods have not considered this issue.

[0083] Antigen binding sites can be precisely determined using X-ray crystallography or cryo-electron microscopy to define the structure of nanobody-antigen complexes. However, obtaining accurate antigen binding sites at high throughput requires significant manpower and resources. Therefore, this application employs a DeepNano-site model to predict antigen binding sites from pure sequence information. The antigen binding sites predicted by the DeepNano-site model are considered key cues, and a cue encoder is designed to extract their features. The DeepNano model combines sequence features and antigen-binding interface features, demonstrating significant performance improvements on independent NAI test sets, such as... Figure 10 As shown, the results indicate that DeepNano's AUROC (0.7941) and AUPRC (0.7689) are both superior to the pure sequence DeepNano-seq (AUROC: 0.6596, AUPRC: 0.6343). When using a threshold of 0.5 to distinguish between positive and negative NAIs, DeepNano has the highest accuracy, reaching 0.9831.

[0084] The following examples illustrate the ability of this application to screen target nanobodies from a large-scale natural nanobody library. Two test cases were constructed. The first case study identified 33 experimentally validated nanobodies that could bind to HSA from a library of one million nanobodies; the second case study identified 59 nanobodies that could bind to GST from a library of one million nanobodies. This application compares three models: DeepNano and DeepNano-seq trained on nanobodies, and DeepNano-seq trained on human PPIs. It should be noted that this application did not test D-SCRIPT and Topsy-Turvy because the file for characterizing one million nanobody-antigen pairs was too large to store on a hard drive.

[0085] To visually represent the results, such as Figure 11 As shown, this application plotted the prediction score distribution of the model for positive nanobodies (33 anti-HSA or 59 anti-GST) and one million background nanobodies. It can be seen that for almost all positive nanobodies, the prediction score of the DeepNano-seq model trained on human PPIs tends to zero, which is consistent with previous results. In Case 1, which screened nanobodies against HSA, the AUROC of DeepNano-seq trained on human PPIs was 0.3, and the FRANK score was 33.63%. Figure 11(Part a in the text). This FRANK value indicates that if DeepNano-seq trained on human PPIs is used for virtual screening of anti-HSA nanobodies, at least 336,400 top-scoring nanobodies should be selected for wet experimental testing to find positive anti-HSA nanobodies. However, without high-throughput technology, this experimental workload is enormous, such as... Figure 11 As shown in section b, compared to the version trained on human PPI data, DeepNano-seq trained on NAI data significantly improved the FRANK metric, reducing the FRANK value for screening anti-HSA to 0.62%. This means that only 6184 nanobodies need to be selected from the top-ranked sequences to find true anti-HSA nanobodies. For DeepNano, the required number of nanobodies is further reduced to 1032 (e.g., ...). Figure 11 (part c in the text).

[0086] In Case Study 2, which screened nanobodies targeting GST, DeepNano's results were largely similar to those in Case Study 1. Furthermore, Figure 12 A distribution chart of predicted scores and experimentally measured ELISA values ​​for 59 anti-GST nanobodies was plotted. The embodiments of this application found that the Pearson correlation coefficient (PCC) between DeepNano's predicted scores and ELISA (logarithmic) values ​​was 0.2491, with a p-value of 0.0571. Notably, during training, DeepNano only learned to distinguish between positive and negative NAIs. Although the PCC value of 0.2491 is not high, it still indicates that DeepNano has learned some capabilities beyond its established knowledge.

[0087] Furthermore, it can be seen that the FRANK-all index of all three models is not very satisfactory, indicating that identifying all known positive nanobodies from a million background sequences remains challenging. Nevertheless, in the virtual screening of anti-HSA and anti-GST, this application still observed improvements in the FRANK-all and AUROC indices of DeepNano compared to the other two DeepNano-seq models, demonstrating the superiority of DeepNano.

[0088] A method for predicting nanobody-antigen interactions, proposed in this application, can extract features of the target amino acid sequence based on a protein big language model to obtain the sequence features of the nanobody and antigen, thereby generating an initial probability of nanobody-antigen interaction. Then, using an antigen binding site prediction model, the estimated antigen binding sites of the sequence features of the nanobody and antigen are determined, and the corresponding target features are extracted and fused with the sequence features of the nanobody and antigen to obtain fused features. These fused features are then used to update the initial probability of nanobody-antigen interaction to obtain the final prediction result of the nanobody-antigen interaction. This solves the problems in related technologies where unsupervised pre-trained protein big language models lack efficient methods for utilizing protein sequence representations, and where all sites on the protein sequence are treated as equally important while ignoring the importance of amino acid sites at the binding interface, thus reducing the prediction performance of nanobody-antigen interactions.

[0089] Next, referring to the accompanying drawings, a nanobody-antigen interaction prediction device according to an embodiment of this application is described.

[0090] Figure 13 This is a block diagram of a nanobody-antigen interaction prediction device according to an embodiment of this application.

[0091] like Figure 13 As shown, the nanobody-antigen interaction prediction device 10 includes: an extraction module 100, a generation module 200, and a prediction module 300.

[0092] Specifically, the extraction module 100 is used to extract features of the target amino acid sequence based on the protein big language model in order to obtain the sequence features of nanobodies and antigens that meet preset conditions.

[0093] The generation module 200 is used to generate an initial nanobody-antigen interaction probability using the sequence characteristics of nanobodies and antigens, and to predict the antigen binding site based on the sequence characteristics of nanobodies and antigens using an antigen binding site prediction model, so as to obtain the estimated binding site of the antigen.

[0094] The prediction module 300 is used to extract the target features of the estimated binding site of the antigen using a preset prompt encoder, and fuse the target features with the sequence features of the nanobody and the antigen to obtain fused features. The fused features are then used to update the initial interaction probability of the nanobody and the antigen to obtain the final prediction result of the interaction between the nanobody and the antigen.

[0095] Optionally, in one embodiment of this application, the extraction module 100 includes a characterization unit and an acquisition unit.

[0096] The representation unit is used to represent each amino acid in the sequence using the embedding vector output from the last hidden layer in the protein big language model.

[0097] The acquisition unit is used to perform target pooling processing on the features of each amino acid based on the minimum pooling strategy, average pooling strategy, and maximum pooling strategy of the target pooling strategy, so as to obtain the sequence features of nanobodies and antigens that meet the preset conditions.

[0098] Optionally, in one embodiment of this application, the generation module 200 includes a construction unit and a generation unit.

[0099] The building unit is used to construct a first predicted interaction probability based on a minimum pooling strategy, a second predicted interaction probability based on an average pooling strategy, and a third predicted interaction probability based on a maximum pooling strategy, based on the sequence characteristics of nanobodies and antigens, respectively.

[0100] The generation unit is used to generate initial nanobody-antigen interaction probabilities using a first predicted interaction probability, a second predicted interaction probability, and a third predicted interaction probability.

[0101] Optionally, in one embodiment of this application, the generation module 300 includes: a processing unit, a splicing unit, and a first determining unit.

[0102] The processing unit is used to perform average pooling of the characteristics of each residue of the nanobody using an antigen binding site prediction model to obtain the final characteristics of the nanobody.

[0103] The splicing unit is used to splice the final feature of the nanobody with the feature of each antigen residue to obtain the spliced ​​feature.

[0104] The first determining unit is used to input the spliced ​​features into a target neural network composed of multiple residual blocks to obtain the predicted binding site of the antigen.

[0105] Optionally, in one embodiment of this application, the prediction module 300 includes: a first acquisition unit, a second determination unit, and a second acquisition unit.

[0106] The first acquisition unit is used to acquire the attention representation of the estimated antigen binding site using a preset prompt encoder.

[0107] The second determining unit is used to determine the attentional embedding of the antigen binding site based on attentional characterization.

[0108] The second acquisition unit is used to embed attention into the sequence features of the nanobody and the antigen to obtain fusion features.

[0109] Optionally, in one embodiment of this application, the initial formula for calculating the interaction probability between the nanobody and the antigen is:

[0110]

[0111] Wherein, X1 and X2 represent the amino acid sequences of the nanobody and the antigen, respectively; Represents the feature mapping function of ESM-2; P represents the i-th multilayer perceptron; min P mean and P max These represent three pooling strategies (min pooling, average pooling, and max pooling).

[0112] It should be noted that the foregoing explanation of an embodiment of a nanobody-antigen interaction prediction method also applies to a nanobody-antigen interaction prediction device of this embodiment, and will not be repeated here.

[0113] A nanobody-antigen interaction prediction device proposed in this application can extract the features of the target amino acid sequence based on a protein big language model to obtain the sequence features of the nanobody and antigen, thereby generating an initial nanobody-antigen interaction probability. Then, using an antigen binding site prediction model, it determines the estimated antigen binding sites of the sequence features of the nanobody and antigen, extracts the corresponding target features, and fuses them with the sequence features of the nanobody and antigen to obtain fused features. These fused features are then used to update the initial nanobody-antigen interaction probability to obtain the final prediction result of the nanobody-antigen interaction. This solves the problem in related technologies where unsupervised pre-trained protein big language models lack efficient methods for utilizing protein sequence representations, and where all sites on the protein sequence are treated as equally important while ignoring the importance of amino acid sites at the binding interface, thus reducing the prediction performance of nanobody-antigen interactions.

[0114] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:

[0115] The memory 1401, the processor 1402, and the computer program stored on the memory 1401 and executable on the processor 1402.

[0116] When the processor 1402 executes the program, it implements a method for predicting the interaction between nanobody and antigen provided in the above embodiments.

[0117] Furthermore, electronic devices also include:

[0118] Communication interface 1403 is used for communication between memory 1401 and processor 1402.

[0119] The memory 1401 is used to store computer programs that can run on the processor 1402.

[0120] The memory 1401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0121] If the memory 1401, processor 1402, and communication interface 1403 are implemented independently, then the communication interface 1403, memory 1401, and processor 1402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 14 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0122] Optionally, in a specific implementation, if the memory 1401, processor 1402, and communication interface 1403 are integrated on a single chip, then the memory 1401, processor 1402, and communication interface 1403 can communicate with each other through an internal interface.

[0123] The processor 1402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0124] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting the interaction between a nanobody and an antigen.

[0125] This embodiment also provides a computer program product, including a computer program that, when executed, is used to implement the above-described method for predicting the interaction between a nanobody and an antigen.

[0126] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0127] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0128] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0129] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0130] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0131] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0132] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0133] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for predicting the interaction between nanobodies and antigens, characterized in that, Includes the following steps: Based on the protein big language model, the features of the target amino acid sequence are extracted to obtain the sequence features of nanobodies and antigens that meet the preset conditions. The initial interaction probability between the nanobody and the antigen is generated using the sequence characteristics of the nanobody and the antigen. Then, the antigen binding site prediction model is used to predict the antigen binding site based on the sequence characteristics of the nanobody and the antigen, so as to obtain the estimated binding site of the antigen. The target features of the estimated binding site of the antigen are extracted using a preset prompt encoder, and the target features are fused with the sequence features of the nanobody and the antigen to obtain fused features. The initial interaction probability of the nanobody and the antigen is updated using the fused features to obtain the final prediction result of the interaction between the nanobody and the antigen.

2. The method according to claim 1, characterized in that, The method of extracting amino acid sequence features based on a protein big language model to obtain sequence features of nanobodies and antigens that meet preset conditions includes: Each amino acid in the sequence is represented by the embedding vector output from the last hidden layer in the protein big language model; Based on the minimum pooling strategy, average pooling strategy, and maximum pooling strategy of the target pooling strategy, the features of each amino acid are subjected to target pooling processing to obtain the sequence features of the nanobody and antigen that meet the preset conditions.

3. The method according to claim 2, characterized in that, The generation of the initial nanobody-antigen interaction probability using the sequence characteristics of the nanobody and the antigen includes: Based on the sequence characteristics of the nanobody and the antigen, a first predicted interaction probability based on the minimum pooling strategy, a second predicted interaction probability based on the average pooling strategy, and a third predicted interaction probability based on the maximum pooling strategy are constructed respectively. The initial nanobody-antigen interaction probability is generated using the first predicted interaction probability, the second predicted interaction probability, and the third predicted interaction probability.

4. The method according to claim 1, characterized in that, The method of using an antigen binding site prediction model to predict antigen binding sites based on the sequence characteristics of the nanobody and the antigen, and obtaining the estimated antigen binding site, includes: Based on the antigen binding site prediction model, the characteristics of each residue of the nanobody are averaged and pooled to obtain the final characteristics of the nanobody. The final feature of the nanobody is spliced ​​with the feature of each antigen residue to obtain the spliced ​​feature; The spliced ​​features are input into a target neural network composed of multiple residual blocks to obtain the estimated binding sites of the antigen.

5. The method according to claim 1, characterized in that, The step of extracting target features of the estimated binding site of the antigen using a preset prompt encoder and fusing the target features with the sequence features of the nanobody and the antigen to obtain fused features includes: The attention representation of the estimated binding site of the antigen is obtained using the preset prompt encoder. Based on the attentional characterization, the attentional embedding of the predicted binding site of the antigen is determined; The attention is embedded into the sequence features of the nanobody and the antigen to obtain the fusion feature.

6. The method according to claim 1, characterized in that, The initial formula for calculating the interaction probability between the nanobody and the antigen is as follows: Wherein, X1 and X2 represent the amino acid sequences of the nanobody and the antigen, respectively; The feature mapping function represents the protein-based large language model ESM-2; P represents the i-th multilayer perceptron; min P mean and P max These represent three pooling strategies (min pooling, average pooling, and max pooling).

7. A device for predicting the interaction between nanobodies and antigens, characterized in that, include: The extraction module is used to extract features of target amino acid sequences based on a protein big language model in order to obtain sequence features of nanobodies and antigens that meet preset conditions. The generation module is used to generate an initial nanobody-antigen interaction probability using the sequence characteristics of the nanobody and the antigen, and to predict the antigen binding site based on the sequence characteristics of the nanobody and the antigen using an antigen binding site prediction model, so as to obtain the estimated binding site of the antigen. The prediction module is used to extract the target features of the estimated binding site of the antigen using a preset prompt encoder, and fuse the target features with the sequence features of the nanobody and the antigen to obtain fused features. The fused features are then used to update the initial interaction probability of the nanobody and the antigen to obtain the final prediction result of the interaction between the nanobody and the antigen.

8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement a method for predicting nanobody-antigen interactions as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement a method for predicting the interaction between a nanobody and an antigen as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to implement a method for predicting the interaction between a nanobody and an antigen as described in any one of claims 1-6.