Autoimmune antibody antigen interaction prediction method and system based on deep learning

CN122551869APending Publication Date: 2026-08-11BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]针对现有技术过度依赖一维序列,缺乏立体视角,从而导致自身免疫抗体抗原互作预测结果可靠性较低的不足,本发明提出一种基于深度学习的自身免疫抗体抗原互作预测方法及系统,从而解决现有技术存在的问题

Benefits of technology

本发明通过采集与自身免疫病相关的抗原序列和抗体序列,在特征提取阶段同时提取抗体CDR区及抗原表位的序列理化特征与基于深度学习预测的三维空间结构特征(包括残基接触图、表面暴露特征矩阵及空间物理化学特征),并在融合阶段采用图神经网络聚合空间邻域信息后再经加权融合生成多模态融合向量,最终通过多层感知机输出互作置信度得分,该方法突破了传统方法仅依赖序列相似性的局限,将蛋白质的空间折叠构象信息引入预测模型,解决了因结构数据匮乏导致的预测准确率低的问题;同时,通过序列特征与空间结构特征的深度融合,将原本孤立的一维信息和三维信息统一建模,克服了数据维度单一的技术瓶颈,有助于提高自身免疫病抗体-抗原互作预测的准确性,为自身免疫病的生物标志物发现及靶向药物研发提供了可靠的技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551869A_ABST
    Figure CN122551869A_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based method and system for predicting autoimmune antibody-antigen interactions, relating to the fields of bioinformatics, database technology, and immunology. The method includes: collecting antigen and antibody sequences related to autoimmune diseases; extracting one-dimensional sequence features and spatial structure features from the antibody and antigen sequences; performing feature extraction on the one-dimensional sequence features to obtain a sequence feature vector; inputting the spatial structure features into a graph neural network to output a structure feature matrix; weighted fusion of the sequence feature vector and the structure feature matrix to first generate antibody-side fusion features and antigen-side fusion features, and then generating a pairing fusion vector through an antibody-antigen pairing layer; inputting the fusion vector into a multilayer perceptron to output the confidence score or affinity prediction value of the antibody-antigen pair interaction. This method overcomes the technical bottleneck of traditional methods with their single data dimension, and helps improve the accuracy of predicting antibody-antigen interactions in autoimmune diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics, database technology, and immunology, specifically to a method and system for predicting autoimmune antibody-antigen interactions based on deep learning. Background Technology

[0002] In existing bioinformatics research, the binding mechanism of antibodies and antigens is the core of drug development and disease diagnosis.

[0003] Currently, most existing antigen-antibody databases are general-purpose, such as sequence-based homology sequence alignment tools or traditional structure matching databases. By inputting the one-dimensional amino acid sequence of the antigen or antibody, sequence prediction tools are used to infer the interaction relationship, or a general antibody database is used for retrieval in the absence of a specific classification for autoimmune diseases. However, the existing schemes have the following defects: (1) Insufficient accuracy: The interaction between antibody and antigen is highly dependent on the three-dimensional structure (epitopes and complementarity-determining regions CDRs). Existing prediction tools are mostly based on sequences, ignoring the influence of complex spatial folding on affinity, resulting in insufficient accuracy; (2) Data silos and uniformity: Existing databases lack automated systems that deeply integrate antibody / antigen sequence characteristics and three-dimensional spatial structure characteristics.

[0004] In summary, current antigen-antibody databases lack proprietary high-throughput data integration for autoimmune diseases (such as lupus, rheumatoid arthritis, and multiple sclerosis). Furthermore, most databases only operate at a single "one-dimensional sequence" level, lacking a three-dimensional perspective, which leads to low reliability of autoimmune antibody-antigen interaction prediction results. Summary of the Invention

[0005] To address the shortcomings of existing technologies that rely excessively on one-dimensional sequences and lack a three-dimensional perspective, resulting in low reliability of autoimmune antibody antigen interaction prediction results, this invention proposes a deep learning-based method and system for predicting autoimmune antibody antigen interactions, thereby solving the problems existing in the prior art.

[0006] A deep learning-based method for predicting autoimmune antibody-antigen interactions includes the following steps: Obtain antigen and antibody sequences associated with autoimmune diseases; One-dimensional sequence features are generated by extracting antibody CDR region sequences from antibody sequences, linear epitope candidate fragments and conformational epitope candidate regions from antigen sequences; The amino acid sequences of the antigen and antibody sequences are input into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue. Based on the three-dimensional spatial coordinate matrix, spatial structural features including residue contact diagrams, surface exposure feature matrices, and spatial physicochemical features are calculated. One-dimensional sequence features are extracted using a sequence feature extraction network to obtain local sequence feature vectors. Spatial structural features are input into a graph neural network, and information about each residue in its three-dimensional neighborhood is aggregated through a message passing mechanism to output a structural feature matrix containing antibody and antigen structural features. The sequence feature vector and structural feature matrix are weighted and fused to generate antibody-side fusion features and antigen-side fusion features first, and then a pairing fusion vector is generated through an antibody-antigen pairing layer. Based on the antibody-antigen pairing fusion vector, the interaction prediction results of the autoimmune antibody antigen containing the interaction confidence score or affinity prediction value are output.

[0007] Furthermore, the step of inputting the amino acid sequences of the antigen and antibody sequences into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue specifically includes the following steps: The input amino acid sequence is used as the target sequence. Depending on the type of structure prediction model used, co-evolutionary information is extracted using multiple sequence alignment encoding, or sequence context features are extracted using single sequence encoding. The co-evolutionary information or sequence context features are jointly encoded with the target sequence features as structure prediction input features. The predicted structural features are fed into a Transformer-based deep learning network, which iteratively processes the data using a self-attention mechanism, outputting each amino acid residue in the sequence. C α This represents the Cartesian coordinates of atoms in three-dimensional space, which in turn generates a three-dimensional coordinate matrix. ,in This represents the length of the amino acid sequence.

[0008] Furthermore, a residue contact diagram is calculated based on a three-dimensional spatial coordinate matrix, specifically by calculating the contact diagram between any two amino acid residues. Euclidean distance between atoms By setting a distance threshold, if If the distance is less than the threshold, the corresponding position in the matrix is ​​assigned a value of 1, indicating contact; otherwise, it is assigned a value of 0, thus generating a residue contact map adjacency matrix.

[0009] Furthermore, the surface exposure feature matrix is ​​calculated based on the three-dimensional spatial coordinate matrix, specifically including: When the three-dimensional spatial coordinate matrix contains only C α When representing atomic coordinates, by using... C αThe atomic coordinates are used to calculate the number of neighboring residues, the average distance between neighboring residues, the local spatial density, and the residue exposure status of the target residue within a preset spatial radius, and to form a surface exposure feature matrix according to the residue sequence. ,in The length of the amino acid sequence, k For feature dimensions; When the structure prediction model outputs the set of all atomic coordinates of the protein and can determine the corresponding atom type and atomic van der Waals radius, the probe rolling ball algorithm is used to simulate water molecules rolling on the surface of the protein's three-dimensional conformation with a solvent probe of a preset radius. The solvent-accessible surface area value and relative solvent-accessible surface area value of each residue are calculated as a supplementary dimension of the surface exposure feature matrix.

[0010] Furthermore, spatial physicochemical features are calculated based on a three-dimensional spatial coordinate matrix. Specifically, the local charge distribution, local hydrophobicity distribution, proportion of polar residues, proportion of aromatic residues, proportion of charged residues, and spatial neighborhood distance statistics are calculated based on the three-dimensional coordinates of amino acids, residue types, atom types, preset charge parameters, protonation state, pH conditions, and spatial neighborhood coordinates, in order to construct a spatial physicochemical feature matrix.

[0011] Furthermore, the weighted fusion of sequence feature vectors and structural feature matrices to first generate antibody-side fusion features and antigen-side fusion features, and then generating paired fusion vectors through an antibody-antigen pairing layer, specifically includes the following steps: A cross-attention mechanism is used to map the structural feature matrix to a query vector and the sequence feature vector to a key vector and a value vector. The attention weight matrix is ​​obtained by calculating the dot product of the query vector and the key vector and then normalizing it. The value vector is weighted using an attention weight matrix to output a cross-modal fusion feature representation after aligning sequence features with structural features; The cross-modal fusion feature representation and the corresponding structural feature matrix are concatenated along the channel dimension, and after being flattened by the pooling layer, antibody-side fusion features and antigen-side fusion features are generated respectively. An antibody-antigen pairing layer is constructed based on antibody-side fusion characteristics and antigen-side fusion characteristics. The pairing layer takes antibody CDR residue characteristics and antigen epitope candidate residue characteristics as inputs, and calculates the correlation weight between antibody residues and antigen residues through cross-protein cross attention to generate antibody-antigen interface fusion characteristics. After weighted fusion of antibody-side fusion features, antigen-side fusion features, and antibody-antigen interface fusion features, the mixture is compressed through a pooling layer and a fully connected layer to output an antibody-antigen pairing fusion vector.

[0012] Furthermore, the step of outputting the interaction prediction result, which includes an interaction confidence score or affinity prediction value for the autoimmune antibody-antigen pairing fusion vector, is specifically achieved by inputting the antibody-antigen pairing fusion vector into a multilayer perceptron. After undergoing multiple nonlinear transformations, the vector is calculated using a Sigmoid or Softmax activation function in the last layer of the multilayer perceptron to output an interaction confidence score between [0,1] or an output used to characterize K. d IC 50 Or, an indicator related to the affinity or activity of ΔG.

[0013] Furthermore, the deep learning-based structure prediction model is AlphaFold2, AlphaFold-Multimer, ESMFold, IgFold, ABodyBuilder, or a similar deep learning structure prediction model trained with antibody structure data.

[0014] Furthermore, it also includes preprocessing the antigen and antibody sequence data after collecting antigen and antibody sequences related to autoimmune diseases; the preprocessing includes sequence deduplication, merging of synonymous disease names, labeling of antibody heavy and light chain sources, merging of antigen homology entries, and labeling of interaction evidence sources; wherein, when there are multiple interaction evidence sources for the same antibody-antigen pair, the multiple evidence sources are merged into an evidence tag list under the same candidate pair index.

[0015] This invention also includes a deep learning-based autoimmune antibody antigen interaction prediction system, comprising: The acquisition module is used to acquire antigen and antibody sequences related to autoimmune diseases; The feature extraction module is used to extract antibody CDR region sequences from antibody sequences, linear epitope candidate fragments and conformational epitope candidate regions from antigen sequences, and encode one-dimensional sequence features; the amino acid sequences of antigen and antibody sequences are input into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue, and spatial structural features including residue contact maps, surface exposure feature matrices and spatial physicochemical features are calculated based on the three-dimensional spatial coordinate matrix. The feature fusion module is used to extract features from one-dimensional sequence features through a sequence feature extraction network to obtain local sequence feature vectors. Spatial structural features are input into a graph neural network, and information of each residue in its three-dimensional spatial neighborhood is aggregated through a message passing mechanism to output a structural feature matrix containing antibody structural features and antigen structural features. The sequence feature vector and the structural feature matrix are weighted and fused to generate antibody-side fusion features and antigen-side fusion features first, and then a pairing fusion vector is generated through an antibody-antigen pairing layer. The prediction module is used to output the interaction prediction results of autoimmune antibody antigens, including interaction confidence scores or affinity prediction values, based on the antibody-antigen pairing fusion vector.

[0016] This invention provides a deep learning-based method for predicting autoimmune antibody-antigen interactions, which has the following advantages: This invention collects antigen and antibody sequences related to autoimmune diseases. During the feature extraction stage, it simultaneously extracts the sequence physicochemical features of antibody CDR regions and antigen epitopes, along with three-dimensional spatial structural features (including residue contact maps, surface exposure feature matrices, and spatial physicochemical features) predicted based on deep learning. In the fusion stage, a graph neural network aggregates spatial neighborhood information, followed by weighted fusion to generate a multimodal fusion vector. Finally, an interaction confidence score is output through a multilayer perceptron. This method overcomes the limitations of traditional methods that rely solely on sequence similarity, introducing protein spatial folding conformation information into the prediction model, thus solving the problem of low prediction accuracy due to a lack of structural data. Furthermore, through the deep fusion of sequence features and spatial structural features, it unifies the previously isolated one-dimensional and three-dimensional information into a unified model, overcoming the technical bottleneck of single data dimensions. This helps improve the accuracy of predicting antibody-antigen interactions in autoimmune diseases, providing reliable technical support for the discovery of biomarkers and the development of targeted drugs for autoimmune diseases. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the deep learning-based method for predicting autoimmune antibody-antigen interactions in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] This invention proposes a deep learning-based method for predicting autoimmune antibody-antigen interactions. This method breaks through the limitations of traditional methods that rely solely on sequence similarity by introducing three-dimensional structures and fusing sequences with structures, thus solving the problem of inaccurate predictions caused by a lack of structural data. Furthermore, it addresses the issue of poor specificity of general databases by constructing a multimodal database for predicting autoimmune antibody-antigen interactions. Finally, it establishes an automated data processing pipeline to overcome the technical bottleneck of cleaning and integrating multi-source heterogeneous data.

[0020] like Figure 1 As shown, the method specifically includes the following steps: S1. Acquire heterogeneous data from multiple sources and perform data cleaning.

[0021] S1.1 Collect information on antigen sequences, antibody (autoantibody) sequences, interaction evidence, disease origin, and sample origin related to autoimmune diseases through web crawlers or API interfaces.

[0022] S1.2. Perform data cleaning on multi-source heterogeneous data to remove redundant and non-standardized data, and construct an underlying dataset containing an antibody-antigen candidate pairing index. The data cleaning process specifically includes: establishing an original data index table for the original antibody, antigen, and disease information, and unifying the antibody number, antigen number, disease name, and sequence format; removing data entries containing non-standard amino acid characters, missing sequence lengths, or missing pairing relationships; merging multiple names for the same antibody, antigen, or disease under a unified number and retaining the original name as a traceability field; and merging multiple sources of evidence for the same antibody-antigen pairing into an evidence label list under the same candidate pairing index when multiple sources of evidence exist for the same antibody-antigen pairing.

[0023] The antibody-antigen candidate pairing index specifically includes: establishing an antibody-antigen candidate pairing index based on a standardized sequence data table, according to disease origin, antigen target, and antibody number; using reported interaction pairs as positive or evidence samples, and unreported but unpredictable pairs as unknown candidate samples; introducing rule-screened non-interaction pairs as negative samples during the training phase, and avoiding the simultaneous appearance of homologous antigens or highly similar antibody sequences in different datasets when dividing the training set, validation set, and test set.

[0024] S2. By extracting antibody CDR region sequences, antigen linear epitopes, and physicochemical properties from antibody and antigen sequences, one-dimensional sequence features are encoded.

[0025] S2.1. One-dimensional sequence features are generated by extracting antibody CDR region sequences from antibody sequences, linear epitope candidate fragments from antigen sequences, and conformational epitope candidate regions.

[0026] On the antibody side, the complementarity determination region (CDR) is located according to a preset antibody numbering rule, which includes one or more of the IMGT, Chothia, or Kabat numbering rules. When the data contains heavy and light chains, the VH and VL sources are labeled respectively, and their pairing relationship is preserved.

[0027] The antigen side includes linear epitope candidate fragments and conformational epitope candidate regions; the conformational epitope candidate regions are generated based on known epitope annotations, sliding windows, residue physicochemical properties, and spatial neighborhood relationships.

[0028] S2.2 Encoding method: Perform amino acid type encoding, position encoding, hydrophobicity encoding, charge encoding, hydrophilicity encoding, and sequence context embedding encoding on antibody CDR fragments, antigen epitope candidate fragments, and full-length context sequences, and align them according to the antibody-antigen candidate pairing index to form a sequence feature encoding set that can be input into the model.

[0029] S3. Input the amino acid sequences of the antigen and antibody sequences into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue. Based on the three-dimensional spatial coordinate matrix, calculate the spatial structural features including residue contact diagrams, surface exposure feature matrices, and spatial physicochemical features.

[0030] S3.1, Input Feature Extraction and Encoding for Structural Prediction: The input amino acid sequence is used as the target sequence. Depending on the type of structure prediction model used, co-evolutionary information is extracted using multiple sequence alignment encoding, or sequence context features are extracted using single sequence encoding. The co-evolutionary information or sequence context features are jointly encoded with the target sequence features as the structure prediction input features.

[0031] S3.2 Generation of three-dimensional spatial coordinate matrix: Structural prediction models (such as deep learning networks derived from AlphaFold) iteratively process the input feature matrix to analyze the spatial proximity between amino acid residues; each amino acid residue in the network output sequence... C α This represents the Cartesian coordinates of atoms in three-dimensional space, which in turn generates a three-dimensional coordinate matrix. ,in This represents the length of the amino acid sequence.

[0032] To ensure consistency in matrix dimensions, the three-dimensional spatial coordinate matrix... S Defined as each amino acid residue C α The set of three-dimensional Cartesian coordinates representing atoms, i.e. ,in N The length is the amino acid sequence length; if the full atomic structure prediction result is used, the full atomic coordinates are also represented as the set of three-dimensional coordinates of each heavy atom under each residue.

[0033] The structural prediction model can be AlphaFold2, AlphaFold-Multimer, ESMFold, IgFold, ABodyBuilder, or a similar deep learning structural prediction model trained on antibody structural data. During implementation, the model name, version, input sequence range, whether multiple sequence alignment is used, and output coordinate type are recorded. This structural prediction model is used to generate structural features.

[0034] S3.3 Calculate spatial structural features based on a three-dimensional spatial coordinate matrix, including residue contact maps, surface exposure feature matrices, and spatial physicochemical features. Specifically, this includes:

[0035] ① Residue contact diagram: Calculate the contact diagram of any two amino acid residues. Euclidean distance between atoms Set a distance threshold (e.g., ...). ),like If the residue is less than the threshold, the corresponding position in the matrix is ​​assigned a value of 1 (indicating contact); otherwise, it is assigned a value of 0, thus generating a residue contact map adjacency matrix. .

[0036] ② Surface exposure feature matrix: When the three-dimensional spatial coordinate matrix contains only C α When representing atomic coordinates, by using... C α The atomic coordinates are used to calculate the number of neighboring residues, the average distance between neighboring residues, the local spatial density, and the residue exposure status of the target residue within a preset spatial radius, and to form a one-dimensional surface exposure feature matrix according to the residue sequence. ,in The length of the amino acid sequence, k For feature dimensions.

[0037] When the structural prediction model is excluded In addition to the coordinate matrix, the algorithm outputs the set of all atomic coordinates of the protein. When the corresponding atom type and atomic van der Waals radius can be determined, the probe rolling ball algorithm is used to simulate water molecules rolling on the surface of the protein's three-dimensional conformation with a solvent probe of a preset radius. The solvent-accessible surface area value and relative solvent-accessible surface area value of each residue are calculated as a supplementary dimension to the surface exposure feature matrix.

[0038] ③ Spatial physicochemical characteristics: Based on the three-dimensional coordinates of amino acids, residue type, atom type, preset charge parameters, protonation state, pH conditions, and spatial neighborhood coordinates, the local charge distribution, local hydrophobicity distribution, proportion of polar residues, proportion of aromatic residues, proportion of charged residues, and spatial neighborhood distance statistics are calculated to construct a spatial physicochemical characteristic matrix.

[0039] S4. Extract features from one-dimensional sequence features using a sequence feature extraction network to obtain local sequence feature vectors; input spatial structure features into a graph neural network, aggregate information of each residue in its three-dimensional spatial neighborhood through a message passing mechanism, and output a structural feature matrix containing spatial folding information; perform weighted fusion of sequence feature vectors and structural feature matrices to first generate antibody-side fusion features and antigen-side fusion features, and then generate pairing fusion vectors through an antibody-antigen pairing layer.

[0040] S4.1 Independent feature extraction from dual-branch networks: Sequence branching: The physicochemical properties of a one-dimensional sequence are input into a one-dimensional convolutional neural network (1D-CNN) or a sequence self-attention model (such as ProtBERT) to extract local motif features and output a sequence feature vector. .

[0041] Structural branch: The extracted spatial structural features are input into a graph neural network (GNN, graph convolutional network GCN, or graph attention network GAT); the GNN aggregates information about each residue in its three-dimensional neighborhood through a message passing mechanism, and outputs a structural feature matrix containing antibody structural features and antigen structural features. .

[0042] S4.2 Multimodal Data Fusion Based on Cross-Attention Mechanism: To make sequence and structural features complementary, a cross-attention mechanism layer was designed for model fusion. ①The structural feature matrix Mapping to a query vector (Query, Q), the sequence feature vector The mapping is a key vector (Key, K) and a value vector (Value, V).

[0043] ② By calculating the dot product of Q and K and performing Softmax normalization, the attention weight matrix is ​​obtained; then, the value vector is weighted and summed using the attention weight matrix to obtain the cross-modal fusion feature representation after the sequence features and spatial structure features are aligned.

[0044] ③ The cross-modal fusion feature representation and the corresponding structural feature matrix are concatenated in the channel dimension and compressed into a fixed-length vector through a pooling layer and a fully connected layer to generate antibody-side fusion features and antigen-side fusion features, respectively.

[0045] Based on antibody-side fusion features and antigen-side fusion features, an antibody-antigen pairing layer is constructed. This pairing layer takes antibody CDR residue features and antigen epitope candidate residue features as inputs, and calculates the correlation weight between antibody residues and antigen residues through cross-protein cross attention to generate antibody-antigen interface fusion features.

[0046] Antibody-side fusion features, antigen-side fusion features, and antibody-antigen interface fusion features are spliced ​​or weighted and then compressed through pooling and fully connected layers to generate antibody-antigen pairing fusion vectors. .

[0047] S5. Based on the antibody-antigen pairing fusion vector, output the autoimmune antibody-antigen interaction prediction results; the interaction prediction results include the interaction confidence score or affinity prediction value: Antibody-antigen pairing fusion vector The input is a multilayer perceptron (MLP, i.e., a fully connected classification network). After multiple nonlinear transformations, the last layer of the MLP is calculated using either a sigmoid or softmax activation function. The final output is the interaction confidence score between [0,1], or an output used to represent K. d IC 50 Alternatively, it can be an affinity or activity-related indicator of ΔG. When the score exceeds a preset threshold, a predictive interaction mapping between the two is established in the database, and a result table is output for the front-end system to call.

[0048] Specifically, when the training data only contains interaction evidence labels, the model uses a Sigmoid or Softmax classification head to output interaction confidence scores between 0 and 1; when the training data contains K... d IC 50 When using affinity or activity-related tags such as ΔG, the model sets up a separate regression head to output the affinity prediction value.

[0049] The interaction confidence score is compared with a preset screening threshold. Antibody-antigen candidate pairs with scores higher than the threshold are sorted in descending order and associated with disease origin, antigen target number, antibody number, evidence label, model version and prediction time to generate a prediction result table that can be called by the database and front-end system.

[0050] S6. Structured storage and visualization of the database: The multidimensional features and interactions obtained from the above processing are associated and stored to establish a relational or graph database model; at the same time, a front-end display system with retrieval, sequence alignment, and three-dimensional structure visualization interactive interface is constructed.

[0051] The database must store at least the antibody sequence list, antigen sequence list, disease label list, candidate pairing index list, interaction evidence list, structural feature index list, model prediction result list, and version record list.

[0052] The front-end system provides functions for searching, sorting, filtering, and viewing three-dimensional structures by disease, antigen target, antibody number, and interaction confidence.

[0053] Based on the above methods, the present invention proposes an embodiment, which specifically includes the following steps: S1: Collect and clean data related to autoimmune diseases to generate a set of candidate autoimmune interaction samples.

[0054] S101: Collect data related to autoimmune diseases, including amino acid sequences of autoantibody variable regions, amino acid sequences of antigens, records of reported antibody-antigen interactions, disease names, and sample source identifiers. The collected data will be compiled into a raw data index table. Autoimmune diseases include, but are not limited to, systemic lupus erythematosus, rheumatoid arthritis, multiple sclerosis, Sjögren's syndrome, and autoimmune thyroid diseases.

[0055] S102: Call the original data index table, unify the antibody number, antigen number, disease name and sequence format, and remove data entries containing non-standard amino acid characters, missing sequence lengths or missing pairing relationships; for cases where there are multiple synonyms for the same antigen or antibody, merge them under a unified number and retain the original name as a traceability field to generate a standardized sequence data table.

[0056] Data standardization and cleaning includes sequence deduplication, merging of synonymous disease names, labeling of antibody heavy and light chain sources, merging of antigen homology entries, and labeling of interaction evidence sources. When there are multiple interaction evidence sources for the same antibody-antigen pair, the multiple evidence sources are merged into a list of evidence labels under the same candidate pair index.

[0057] S103: Based on the standardized sequence data table, an antibody-antigen candidate pairing index is established according to disease origin, antigen target, and antibody number. If a pair has a reported interaction record, the record is written into the corresponding entry as an interaction evidence label; if a pair is a candidate for prediction, a candidate prediction label is written into it, ultimately generating an autoimmune candidate interaction sample set.

[0058] S2: Extract one-dimensional sequence features of antibodies and antigens to generate a sequence feature coding set.

[0059] S201: Call the antibody sequence from the autoimmune candidate interaction sample set, locate the complementarity determination region (CDR) according to the preset antibody numbering rules, which include one or more of the IMGT, Chothia, or Kabat numbering rules; when the data contains heavy and light chains, label the VH and VL sources respectively and retain their pairing relationship.

[0060] S202: This function retrieves antigen sequences from the autoimmune candidate interaction sample set and filters candidate epitope regions based on known epitope annotations, a sliding window segment with a set step size, and selected residue physicochemical properties (such as hydrophilicity and antigenicity index). The candidate epitope regions include both linear epitope candidate fragments and conformational epitope candidate regions that can be subsequently associated with spatial neighborhood information, generating a list of candidate epitope fragments.

[0061] S203: Perform amino acid type encoding, position encoding, hydrophobicity encoding, charge encoding, and sequence context embedding encoding on the antibody key fragment list and the antigen epitope candidate fragment list, respectively. Sequence context embedding encoding can be obtained by a one-dimensional convolutional neural network, a recurrent neural network, or a sequence self-attention encoder. All encodings are aligned according to the antibody-antigen pairing index to generate a sequence feature encoding set.

[0062] S3: Construct a spatial structure map set.

[0063] S301: Input the antibody sequence and antigen sequence into the deep learning structure prediction model respectively to obtain the three-dimensional spatial coordinate matrix of antibody residues and the three-dimensional spatial coordinate matrix of antigen residues. ,in This represents the length of the corresponding amino acid sequence. Deep learning structure prediction models can be protein structure prediction models based on the AlphaFold architecture, or structure prediction models trained with antibody structure data.

[0064] S302: Based on the three-dimensional spatial coordinate matrix of antibody residues and the three-dimensional spatial coordinate matrix of antigen residues, calculate the atoms representing any two residues. Euclidean distance between .like Less than a preset distance threshold (e.g.) Then, in the residue contact adjacency matrix Write the contact identifier in the residue field; otherwise, write the non-contact identifier. Simultaneously record the continuous distance values ​​to form a residue distance matrix.

[0065] S303: Based on the three-dimensional spatial coordinate matrix of residues Calculate the surface exposure characteristics, local charge distribution, local hydrophobicity characteristics, and spatial neighborhood properties of each residue. Where S is C α When the coordinate matrix is ​​obtained, the surface exposure features are used to approximately describe the exposure degree of residues in the three-dimensional conformation; when full atomic coordinate information is obtained, the surface exposure features include solvent-accessible surface area or relative solvent-accessible surface area. Local charge distribution and local hydrophobicity features are used to describe the physicochemical environment around the residues, and spatial neighborhood attributes are used to describe the adjacency relationship of residues in the three-dimensional conformation. These features collectively serve as node attributes of the graph neural network, generating a spatial structure map set.

[0066] S4: Align and fuse sequence features and spatial structure features to generate a multimodal fusion vector set.

[0067] S401: Invoke the sequence feature encoding set, input the antibody complementarity determination region encoding, antigen epitope candidate fragment encoding, and sequence context embedding vector into the sequence feature extraction network to obtain a sequence feature vector containing antibody sequence feature vectors and antigen sequence feature vectors. .

[0068] S402: This function calls upon a spatial structure map set, using the residue contact adjacency matrix as the edge connections in a graph neural network (Graph Neural Network), and surface exposure features, local charge distribution, local hydrophobicity features, and spatial neighborhood attributes as node attributes to construct antibody and antigen structure maps. It's important to note that the Graph Neural Network is input to the adjacency matrix and node attribute matrix from the spatial structure map set, not just a single residue contact adjacency matrix or a single physicochemical attribute.

[0069] S403: Message passing is performed on the antibody and antigen structure maps using a graph convolutional network or graph attention network, enabling each residue node to aggregate the structural information of its spatial neighborhood residue nodes, generating a structure feature matrix containing both the antibody and antigen structure feature matrices. .

[0070] S404: Employs a cross-attention mechanism to align and fuse the sequence feature vector and the structural feature matrix. Specifically, the structural feature matrix... Mapping to query vectors, converting sequence feature vectors The query vector is mapped to a key vector and a value vector. The similarity between the query vector and the key vector is calculated and normalized to obtain the cross-modal attention weight matrix. Then, the value vector is weighted and summed using the cross-modal attention weight matrix to obtain the cross-modal fusion feature representation after the sequence features and spatial structure features are aligned.

[0071] S405: The cross-modal fusion feature representation and the corresponding structural feature matrix are concatenated along the channel dimension and compressed into a fixed-length vector through a pooling layer and a fully connected layer to generate antibody-side fusion features and antigen-side fusion features, respectively. Based on the antibody-side fusion features and antigen-side fusion features, an antibody-antigen pairing layer is constructed. The pairing layer takes antibody CDR residue features and antigen epitope candidate residue features as input, and calculates the correlation weight between antibody residues and antigen residues through cross-protein cross-attention to generate antibody-antigen interface fusion features. After weighted fusion of antibody-side fusion features, antigen-side fusion features and antibody-antigen interface fusion features, the vector is compressed through a pooling layer and a fully connected layer to output the antibody-antigen pairing fusion vector.

[0072] S5: By inputting the antibody-antigen pairing fusion vector into a multilayer perceptron, and after multiple nonlinear transformations, the final layer of the multilayer perceptron is calculated using a sigmoid or softmax activation function to output an interaction confidence score between [0,1] or an output used to characterize K. d IC 50 Or, an indicator related to the affinity or activity of ΔG.

[0073] Specifically, when the training data only contains interaction evidence labels, the model uses a Sigmoid or Softmax classification head to output interaction confidence scores between 0 and 1; when the training data contains K... d IC 50 When using affinity or activity-related tags such as ΔG, the model sets up a separate regression head to output the affinity prediction value.

[0074] The interaction confidence score is compared with a preset screening threshold. Antibody-antigen candidate pairs with scores higher than the threshold are sorted in descending order and associated with disease origin, antigen target number, antibody number, evidence label, model version and prediction time to generate a prediction result table that can be called by the database and front-end system.

[0075] The interaction confidence score is compared with a preset screening threshold. Antibody-antigen pairs with scores higher than the preset screening threshold are sorted in descending order and associated with disease source labels, antigen target numbers and evidence labels to generate a set of autoimmune antibody-antigen interaction prediction results.

[0076] The present invention has the following beneficial effects: (1) Multidimensional perspective improves accuracy: It breaks through the limitations of relying solely on sequence similarity in the traditional approach. By introducing sequence features, structural features, and interface pairing features, it helps to improve the reliability of predicting antibody-antigen interactions in autoimmune diseases.

[0077] (2) Field expertise and application value: A multimodal database for predicting antibody-antigen interactions in autoimmune diseases has been constructed, providing a solid data foundation for the discovery of early screening biomarkers and the development of targeted drugs for autoimmune diseases.

[0078] (3) Automation of data processing: It realizes the construction of an automated pipeline from raw sequence to multidimensional data of "sequence-structure-interaction", which improves the efficiency of database update.

[0079] Based on the same inventive concept, this invention also proposes a deep learning-based autoimmune antibody antigen interaction prediction system, comprising: The acquisition module is used to acquire antigen and antibody sequences related to autoimmune diseases.

[0080] The feature extraction module is used to extract antibody CDR region sequences from antibody sequences, linear epitope candidate fragments and conformational epitope candidate regions from antigen sequences to encode one-dimensional sequence features; the amino acid sequences of antigen and antibody sequences are input into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue, and spatial structural features including residue contact maps, surface exposure feature matrices and spatial physicochemical features are calculated based on the three-dimensional spatial coordinate matrix.

[0081] The feature fusion module is used to extract features from one-dimensional sequence features through a sequence feature extraction network to obtain local sequence feature vectors. Spatial structural features are input into a graph neural network, and information of each residue in its three-dimensional spatial neighborhood is aggregated through a message passing mechanism to output a structural feature matrix containing antibody structural features and antigen structural features. The sequence feature vector and the structural feature matrix are weighted and fused to generate antibody-side fusion features and antigen-side fusion features first, and then a pairing fusion vector is generated through an antibody-antigen pairing layer.

[0082] The prediction module is used to output the interaction prediction results of autoimmune antibody antigens, including interaction confidence scores or affinity prediction values, based on the antibody-antigen pairing fusion vector.

[0083] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A deep learning-based method for predicting autoimmune antibody-antigen interactions, characterized in that, Includes the following steps: Obtain antigen and antibody sequences associated with autoimmune diseases; One-dimensional sequence features are generated by extracting antibody CDR region sequences from antibody sequences, linear epitope candidate fragments and conformational epitope candidate regions from antigen sequences; The amino acid sequences of the antigen and antibody sequences are input into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue. Based on the three-dimensional spatial coordinate matrix, spatial structural features including residue contact diagrams, surface exposure feature matrices, and spatial physicochemical features are calculated. One-dimensional sequence features are extracted using a sequence feature extraction network to obtain local sequence feature vectors. Spatial structural features are input into a graph neural network, and information about each residue in its three-dimensional neighborhood is aggregated through a message passing mechanism to output a structural feature matrix containing antibody and antigen structural features. The sequence feature vector and structural feature matrix are weighted and fused to generate antibody-side fusion features and antigen-side fusion features first, and then a pairing fusion vector is generated through an antibody-antigen pairing layer. Based on the antibody-antigen pairing fusion vector, the interaction prediction results of the autoimmune antibody antigen containing the interaction confidence score or affinity prediction value are output.

2. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The step of inputting the amino acid sequences of the antigen and antibody sequences into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue specifically includes the following steps: The input amino acid sequence is used as the target sequence. Depending on the type of structure prediction model used, co-evolutionary information is extracted using multiple sequence alignment encoding, or sequence context features are extracted using single sequence encoding. The co-evolutionary information or sequence context features are jointly encoded with the target sequence features as structure prediction input features. The predicted structural features are fed into a Transformer-based deep learning network, which iteratively processes the data using a self-attention mechanism, outputting each amino acid residue in the sequence. C α This represents the Cartesian coordinates of atoms in three-dimensional space, which in turn generates a three-dimensional coordinate matrix. ,in This represents the length of the amino acid sequence.

3. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The residue contact diagram is calculated based on a three-dimensional spatial coordinate matrix, specifically by calculating any two amino acid residues. Euclidean distance between atoms ; By setting a distance threshold, if If the distance is less than the threshold, the corresponding position in the matrix is ​​assigned a value of 1, indicating contact; otherwise, it is assigned a value of 0, thus generating a residue contact map adjacency matrix.

4. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The surface exposure feature matrix is ​​calculated based on a three-dimensional spatial coordinate matrix, specifically including: When the three-dimensional spatial coordinate matrix contains only C α When representing atomic coordinates, by using... C α The atomic coordinates are used to calculate the number of neighboring residues, the average distance between neighboring residues, the local spatial density, and the residue exposure status of the target residue within a preset spatial radius, and to form a surface exposure feature matrix according to the residue sequence. ,in The length of the amino acid sequence, k For feature dimensions; When the structure prediction model outputs the set of all atomic coordinates of the protein and can determine the corresponding atom type and atomic van der Waals radius, the probe rolling ball algorithm is used to simulate water molecules rolling on the surface of the protein's three-dimensional conformation with a solvent probe of a preset radius. The solvent-accessible surface area value and relative solvent-accessible surface area value of each residue are calculated as a supplementary dimension of the surface exposure feature matrix.

5. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, Based on the calculation of spatial physicochemical features using a three-dimensional spatial coordinate matrix, specifically by calculating the local charge distribution, local hydrophobicity distribution, proportion of polar residues, proportion of aromatic residues, proportion of charged residues, and spatial neighborhood distance statistics according to the three-dimensional coordinates of amino acids, residue type, atom type, preset charge parameters, protonation state, pH conditions, and spatial neighborhood coordinates, a spatial physicochemical feature matrix is ​​constructed.

6. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The weighted fusion of sequence feature vectors and structural feature matrices to generate antibody-side fusion features and antigen-side fusion features, followed by the generation of paired fusion vectors through an antibody-antigen pairing layer, specifically includes the following steps: A cross-attention mechanism is used to map the structural feature matrix to a query vector and the sequence feature vector to a key vector and a value vector. The attention weight matrix is ​​obtained by calculating the dot product of the query vector and the key vector and then normalizing it. The value vector is weighted using an attention weight matrix to output a cross-modal fusion feature representation after aligning sequence features with structural features; The cross-modal fusion feature representation and the corresponding structural feature matrix are concatenated along the channel dimension, and after being flattened by the pooling layer, antibody-side fusion features and antigen-side fusion features are generated respectively. An antibody-antigen pairing layer is constructed based on antibody-side fusion characteristics and antigen-side fusion characteristics. The pairing layer takes antibody CDR residue characteristics and antigen epitope candidate residue characteristics as inputs, and calculates the correlation weight between antibody residues and antigen residues through cross-protein cross attention to generate antibody-antigen interface fusion characteristics. After weighted fusion of antibody-side fusion features, antigen-side fusion features, and antibody-antigen interface fusion features, the mixture is compressed through a pooling layer and a fully connected layer to output an antibody-antigen pairing fusion vector.

7. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The interaction prediction result, which outputs an interaction confidence score or affinity prediction value for autoimmune antibody-antigen based on the antibody-antigen pairing fusion vector, is specifically achieved by inputting the antibody-antigen pairing fusion vector into a multilayer perceptron. After undergoing multiple nonlinear transformations, the final layer of the multilayer perceptron is calculated using a sigmoid or softmax activation function to output an interaction confidence score between [0,1] or an output used to characterize K. d IC 50 Or, an indicator related to the affinity or activity of ΔG.

8. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, The deep learning-based structure prediction model is AlphaFold2, AlphaFold-Multimer, ESMFold, IgFold, ABodyBuilder, or a similar deep learning structure prediction model trained on antibody structure data.

9. The method for predicting autoimmune antibody-antigen interactions based on deep learning according to claim 1, characterized in that, It also includes preprocessing the antigen and antibody sequence data after collecting antigen and antibody sequences related to autoimmune diseases; the preprocessing includes sequence deduplication, merging of synonymous disease names, labeling of antibody heavy and light chain sources, merging of antigen homology entries, and labeling of interaction evidence sources; wherein, when there are multiple interaction evidence sources for the same antibody or antigen pair, the multiple evidence sources are merged into an evidence label list under the same candidate pair index.

10. A deep learning-based autoimmune antibody antigen interaction prediction system, characterized in that, include: The acquisition module is used to acquire antigen and antibody sequences related to autoimmune diseases; The feature extraction module is used to extract antibody CDR region sequences from antibody sequences, linear epitope candidate fragments and conformational epitope candidate regions from antigen sequences, and encode one-dimensional sequence features. The amino acid sequences of the antigen and antibody sequences are input into a deep learning-based structure prediction model to generate a three-dimensional spatial coordinate matrix for each amino acid residue. Based on the three-dimensional spatial coordinate matrix, spatial structural features including residue contact diagrams, surface exposure feature matrices, and spatial physicochemical features are calculated. The feature fusion module is used to extract features from one-dimensional sequence features through a sequence feature extraction network to obtain local sequence feature vectors. Spatial structural features are input into a graph neural network, and information of each residue in its three-dimensional spatial neighborhood is aggregated through a message passing mechanism to output a structural feature matrix containing antibody structural features and antigen structural features. The sequence feature vector and the structural feature matrix are weighted and fused to generate antibody-side fusion features and antigen-side fusion features first, and then a pairing fusion vector is generated through an antibody-antigen pairing layer. The prediction module is used to output the interaction prediction results of autoimmune antibody antigens, including interaction confidence scores or affinity prediction values, based on the antibody-antigen pairing fusion vector.