Geometric graph network-based B cell epitope prediction method and equipment, and storage medium

By extracting antigen sequences and three-dimensional structural features, and using the target neural network to process the data set, the problem of insufficient multi-level feature capture in B cell epitope prediction is solved, improving the accuracy and reliability of the prediction.

CN120388604APending Publication Date: 2025-07-29SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510348324.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, B cell epitope prediction is difficult to fully capture the multi-level characteristics of protein interactions, resulting in poor prediction performance.

Method used

By extracting the sequence characteristics of the antigen sequence and the structural characteristics of the three-dimensional structural data of the protein, an antigen map is constructed, and the antigen map data set is processed based on the pre-trained target neural network to determine the predicted epitope.

Benefits of technology

It significantly improves the accuracy and reliability of epitope prediction, accurately captures the key geometric features of protein surface and its interactions, and solves the problem of predicting complex epitope conformations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388604A_ABST
    Figure CN120388604A_ABST
Patent Text Reader

Abstract

The invention discloses a geometric graph network-based B cell epitope prediction method and device and a storage medium, and relates to the technical field of machine learning, the method comprises the following steps: extracting sequence features corresponding to an antigen sequence, and extracting structural features corresponding to three-dimensional structural data of protein; determining geometric features of the antigen according to an antigen map constructed by the three-dimensional structure data; splicing the geometric features, the sequence features and the structural features into an antigen image data set; and processing the antigen map data set based on a pre-trained target neural network, and determining a predicted epitope. The technical problem that the B cell epitope prediction performance is low due to the fact that multi-level characteristics of protein interaction are difficult to comprehensively capture for linear epitope prediction in the prior art is solved, and the epitope prediction accuracy and reliability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and particularly relates to a B-cell epitope prediction method, device, and storage medium based on a geometric graph network. Background Art

[0002] B cells are a major component of the adaptive immune system. Therefore, the recognition of B-cell epitopes is of great significance in some biotechnological and clinical applications, such as vaccine design, disease diagnosis, and therapeutic antibody development.

[0003] In related technologies, the random forest algorithm is mainly used for prediction, and multiple sequence features such as amino acid composition, secondary structure, and solvent accessibility are combined to represent the epitope region. However, B-cell epitopes are composed of multiple discontinuous protein fragments, and the related technologies for linear epitope prediction are difficult to comprehensively capture the multi-level features of protein interactions, thereby resulting in low performance of B-cell epitope prediction.

[0004] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of the present application is to provide a B-cell epitope prediction method, device, and storage medium based on a geometric graph network, aiming to solve the technical problem that in related technologies for linear epitope prediction, it is difficult to comprehensively capture the multi-level features of protein interactions, thereby resulting in low performance of B-cell epitope prediction.

[0006] To achieve the above purpose, the present application proposes a B-cell epitope prediction method based on a geometric graph network, and the method includes:

[0007] Extracting the sequence features corresponding to the antigen sequence, and extracting the structural features corresponding to the three-dimensional structure data of the protein;

[0008] Determining the geometric features of the antigen according to the antigen graph constructed from the three-dimensional structure data;

[0009] Concatenating the geometric features, the sequence features, and the structural features into an antigen graph data set;

[0010] Processing the antigen graph data set based on a pre-trained target neural network to determine the predicted epitope.

[0011] In one embodiment, the antigen sequence is converted into a vector representation by a protein language model, and the sequence features are gradually extracted based on the vector representation; and,

[0012] Extract the relative solvent accessibility of amino acids and the secondary structure information of the antigen from the three-dimensional structure data, and convert the secondary structure information into a one-hot encoded vector;

[0013] Concatenate the relative solvent accessibility and the one-hot encoded vector according to the dimension to obtain the structural feature.

[0014] In one embodiment, based on the coordinate matrix of each antigen and combining the spatial structure and shape information of the antigen in the three-dimensional structure data, model the antigen as the antigen graph;

[0015] Based on the local coordinate system, compare the distance, direction, and angle characteristics between the internal atoms of the nodes, and obtain the geometric node features by calling the node function;

[0016] Based on each antigen graph, determine the geometric edge features by comparing the Euclidean distance between each node and the distance threshold and calling the edge function;

[0017] Merge the geometric node features and the geometric edge features into the geometric feature.

[0018] In one embodiment, concatenate the antigen graph with the corresponding sequence feature, structural feature, and geometric feature of the antigen graph to obtain the antigen graph data;

[0019] Integrate each antigen graph data to obtain the antigen graph data set.

[0020] In one embodiment, based on the coordinate matrix of the antigen, determine the coordinates of each atom in the antigen and the central coordinates of each side chain atom;

[0021] Based on the coordinates of each atom and the central coordinates of each side chain atom, call the vector function to calculate and obtain the vector feature;

[0022] According to the vector feature and the coordinate matrix, obtain the target antigen graph data set;

[0023] Determine the predicted epitope based on the target antigen graph data set.

[0024] In one embodiment, based on the vector feature, call the dihedral angle calculation formula to calculate the dihedral angle between the normal vector of two residues and the direction vector connecting the residues, and input the dihedral angle into the mapping formula to obtain the input feature;

[0025] Based on the input feature and the coordinate matrix, perform convolution on the node feature and the edge feature, and perform message passing through the spherical harmonic function to integrate and obtain the target node feature;

[0026] Concatenate the edge features and the target node features, and process them through linear transformation and activation function to obtain target edge features;

[0027] Update the node features and the edge features to the target node features and the target edge features in the antigen graph dataset to obtain the target antigen graph dataset.

[0028] In one embodiment, the target neural network extracts antigen representations based on the target antigen graph dataset;

[0029] Based on the antigen representations, the weight matrix, and the bias term, and invoking a logic function, determine the predicted epitopes.

[0030] In one embodiment, divide the antigen graph dataset by cross-validation, use one fold of the antigen graph dataset as the validation set, and use the remaining antigen graph datasets as the training datasets;

[0031] Train the basic neural network by using the training datasets to obtain an output result;

[0032] Adjust the parameters of the basic neural network by backpropagation according to the deviation between the output result and the true epitope labels;

[0033] Evaluate the performance of each basic neural network according to the validation set, and set the basic neural network with the highest performance as the target neural network.

[0034] In addition, to achieve the above object, the present application also proposes an epitope prediction device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the B cell epitope prediction method based on the geometric graph network as described above.

[0035] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the B cell epitope prediction method based on the geometric graph network as described above.

[0036] The present application provides a method for predicting B-cell epitopes based on a geometric graph network, including extracting sequence features corresponding to an antigen sequence and extracting structural features corresponding to three-dimensional structure data of a protein; determining geometric features of the antigen according to an antigen graph constructed based on the three-dimensional structure data; splicing the geometric features, the sequence features, and the structural features into an antigen graph data set; and processing the antigen graph data set based on a pre-trained target neural network to determine predicted epitopes. By fusing the sequence information and three-dimensional geometric features of the antigen, the deficiencies of sequence-based methods in dealing with conformational epitopes and the limitations of structure-based methods in not fully utilizing sequence information and geometric information are made up for, and the accuracy of epitope prediction is improved.

[0037] In summary, the present application effectively solves the problem of predicting complex epitope conformations by accurately capturing the key geometric features on the protein surface and its interactions, overcomes the technical problem that related technologies are difficult to comprehensively capture the multi-level features of protein interactions for linear epitope prediction, which leads to low performance in B-cell epitope prediction, and significantly improves the accuracy and reliability of epitope prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a schematic flowchart of the first embodiment of the method for predicting B-cell epitopes based on a geometric graph network of the present application;

[0041] Figure 2 It is a schematic flowchart of the second embodiment of the method for predicting B-cell epitopes based on a geometric graph network of the present application;

[0042] Figure 3 It is a schematic flowchart of the third embodiment of the method for predicting B-cell epitopes based on a geometric graph network of the present application;

[0043] Figure 4 It is a schematic flowchart of the sixth embodiment of the method for predicting B-cell epitopes based on a geometric graph network of the present application;

[0044] Figure 5 It is an architecture diagram of the method for predicting B-cell epitopes based on a geometric graph network of the present application;

[0045] Figure 6 This is a schematic structural diagram of the epitope prediction device for this application form.

[0046] The realization of the purpose, functional features and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Specific implementation manners

[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0048] In the related art, the random forest algorithm is mainly used for prediction, and multiple sequence features such as amino acid composition, secondary structure, and solvent accessibility are combined to represent the epitope region. However, B cell epitopes are composed of multiple discontinuous protein fragments, and the related art focuses on linear epitope prediction, making it difficult to comprehensively capture the multi-level features of protein interactions, thus resulting in low performance in B cell epitope prediction.

[0049] This application provides a solution: First, extract the sequence features corresponding to the antigen sequence and the structural features corresponding to the three-dimensional structure data of the protein; then, determine the geometric features of the antigen according to the antigen graph constructed based on the three-dimensional structure data; then, splice the geometric features, the sequence features, and the structural features into an antigen graph data set; finally, process the antigen graph data set based on a pre-trained target neural network to determine the predicted epitope.

[0050] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an epitope prediction device, etc. that can implement the above functions. Hereinafter, taking the epitope prediction device as an example, this embodiment and the following embodiments will be described.

[0051] To better understand the technical solutions of this application, the following will be described in detail with reference to the accompanying drawings of the specification and specific implementation manners.

[0052] The embodiment of this application provides a B cell epitope prediction method based on a geometric graph network, referring to Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the B cell epitope prediction method based on a geometric graph network of this application.

[0053] In this embodiment, the B cell epitope prediction method based on a geometric graph network includes steps S10 to S40:

[0054] Step S10, extract the sequence features corresponding to the antigen sequence and the structural features corresponding to the three-dimensional structure data of the protein.

[0055] In this embodiment, the antigen sequence features are extracted by analyzing the amino acid sequence and physicochemical properties, such as hydrophobicity and charge, converting them into numerical vectors, and using a language model, such as SaProt (Structure-aware Protein Language Model). The three-dimensional structure features are converted into multi-dimensional data by measuring the protein surface exposure, secondary structure types, such as α-helix, and spatial geometric information, such as curvature, and are extracted using a structure tool, such as DSSP (Define Secondary Structure of Proteins).

[0056] As an alternative implementation of extracting sequence features, the newly proposed pre-trained protein language model SaProt_650M_PDB (Structure-aware Protein Language Model (650M parameters, PDB-trained)) is used to extract the features of each antigen sequence, convert each amino acid into a vector representation, and integrate to obtain sequence features.

[0057] Two structural properties were extracted from the PDB file of the antigen structure using the DSSP program. One is the relative solvent accessibility (RSA) of amino acids, which is used to characterize the exposure degree of amino acids in the solvent; the other is the secondary structure type of the antigen, and the extracted secondary structure information is converted into an 8-dimensional one-hot encoded vector. Finally, these two types of features together constitute a 9-dimensional DSSP feature vector.

[0058] Step S20, determine the geometric features of the antigen according to the antigen graph constructed from the three-dimensional structure data.

[0059] In this embodiment, the antigen graph is a graph network constructed based on the three-dimensional structure data of the antigen, with amino acid residues as nodes and edges representing the spatial or covalent connection relationships between residues. The geometric features calculate the curvature and normal vector through the Cα atom coordinates to quantify the topological morphology in three-dimensional space.

[0060] As an alternative implementation, the three-dimensional coordinates of all Cα atoms are extracted from the antigen PDB file to construct point cloud data. An algorithm is used to search for the k nearest neighbor atoms in the space of the Cα atom of each residue to form a local surface. In each residue, first calculate the negative bisector of the angle formed by the N, Cα, and C atoms, then calculate the normal vector of the plane formed by the N, Cα, and C atoms, and then construct a local three-dimensional coordinate system. Based on this coordinate system, geometric node features and geometric edge features that satisfy rotation and translation invariance are defined, and the geometric node features and geometric edge features are integrated into geometric features.

[0061] Step S30: Concatenate the geometric features, the sequence features, and the structural features into an antigen map dataset.

[0062] In this embodiment, the antigen map dataset is a high-dimensional biomolecular dataset that integrates the amino acid sequence features, three-dimensional structural features, and geometric features of antigens, describes the sequence-structure-function associations of antigens in the form of graph data, encapsulates node attributes, edge connections, and edge attributes, and is used for graph neural network training, supporting epitope prediction, antibody design, and antigen-receptor interaction analysis.

[0063] As an alternative implementation, first ensure strict alignment of the antigen file with the sequence residue numbers through preprocessing, use the DSSP tool to parse the file to generate the solvent accessibility and secondary structure types of each residue, constituting the structural features; at the same time, call the SaProt_650M_PDB model to tokenize and encode the sequence, and extract the embedding vectors of each residue as the sequence features; then construct a three-dimensional point cloud based on the Cα atom coordinates, search for neighboring atoms of each residue, construct a local three-dimensional coordinate system, and based on this coordinate system, define the geometric features; concatenate the three types of features in the residue order into a node feature matrix, construct a bidirectional edge connection relationship based on covalent bonds and spatial proximity, define the edge type label and the normalized distance value as the edge attributes, and finally, after encapsulating the node features, edge indices, and edge attributes and checking the consistency of the total number of residues, save it as a standardized antigen map dataset.

[0064] Step S40: Process the antigen map dataset based on a pre-trained target neural network to determine the predicted epitopes.

[0065] In this embodiment, the target neural network is the I3NN (SE(3)-invariant geometric graph neural network) module. I3NN is an improved version based on E3NN, which can capture both evolutionary and geometric information simultaneously. This model achieves geometric invariance of the model by introducing spherical harmonic function convolution of dihedral angles, enhancing the model's perception ability of geometric structures. B cell epitopes are regions on the antigen surface that are specifically recognized by B cell receptors or antibodies.

[0066] As an alternative implementation, first divide the standardized antigen map data into a training set, and then input the features into the I3NN neural network. Inside the I3NN module, the node feature vectors and edge feature vectors of the graph are convolved on the invariance layer, and information is transmitted through spherical harmonic functions to aggregate the features of neighboring nodes and edge features, enabling the model to effectively integrate local geometric information, obtain updated node features and edge features through the local geometric information, and finally predict the epitope probabilities of each residue through the output layer of the multi-layer perceptron with activation functions.

[0067] For example, first, perform node feature normalization and training set division on the standardized antigen graph data (including 1280-dimensional sequence embeddings extracted from SaProt_650M_PDB, 9-dimensional structural features parsed by DSSP, and geometric features defined by constructing a local three-dimensional coordinate system based on Cα atom coordinates). Subsequently, input the features into the I3NN neural network. This neural network performs convolution on the node feature vectors and edge feature vectors of the graph on the invariance layer, and transmits information through spherical harmonic functions, aggregating the features of neighboring nodes and edge features, enabling the model to effectively integrate local geometric information. After obtaining the updated node feature representation, further update the edge features by using two nodes connected by an edge to achieve message passing. Then, concatenate the source node features, edge features, and target node features together, and process them through a linear transformation and an activation function to obtain the updated edge features. Finally, predict the epitope probability of each residue through the output layer of the multi-layer perceptron activated by the Sigmoid (S-shaped function, Sigmoid Function) activation function.

[0068] Exemplarily, the input data includes the sequence and three-dimensional structure of the antigen. By integrating the sequence features extracted by the protein language model and the structural features extracted from the three-dimensional geometric information of the antigen, the graph representation of the antigen is obtained. Then, an I3NN neural network based on protein surface features is built, and message passing is performed through graph convolution to update the node and edge feature representations of the graph. After the graph convolution layer, the epitope prediction result is obtained through a multi-layer perceptron. Then, the model is pre-trained using protein-protein interaction site prediction data, and the prediction accuracy is further improved through transfer learning. Generally speaking, the present invention makes full use of the sequence and spatial structure information of proteins, combined with the transfer learning method, and significantly improves the accuracy of antigen epitope prediction.

[0069] Since the key geometric features of the protein surface and its interactions are accurately captured, the problem of predicting complex epitope conformations is effectively solved, overcoming the technical problem that the related technology is difficult to comprehensively capture the multi-level features of protein interactions for linear epitope prediction, which leads to the low performance of B cell epitope prediction, and significantly improving the accuracy and reliability of epitope prediction.

[0070] Based on any of the above embodiments, in the second embodiment of the present application, refer to Figure 2 , Figure 2 is a schematic flowchart of the second embodiment of the B cell epitope prediction method based on a geometric graph network of the present application. The step S10 includes steps A11 to A13:

[0071] Step A11, convert the antigen sequence into a vector representation through a protein language model, and gradually extract the sequence features from the vector representation.

[0072] In this embodiment, the protein language model is the SaProt language model, which encodes the antigen sequence and uses the high-dimensional and information-rich representation as part of the node encoding. Second, a large amount of protein-protein complex data is used for pre-training, thereby increasing more training data, and then the antigen dataset is used for fine-tuning. We have demonstrated through a large number of experiments that the pre-training step helps the model parameters get closer to the optimal values. Vector representation means mapping each residue into a dense vector in a high-dimensional space through the model embedding layer.

[0073] As an alternative implementation, load the pre-trained protein language model and its dedicated tokenizer, input the preprocessed target sequence into the tokenizer to convert it into a token ID sequence, map each token into an initial vector through the model embedding layer, calculate the context correlation weights through the self-attention mechanism of the multi-layer Transformer encoder, extract the hidden state of the last layer as the sequence feature after layer-by-layer transmission, remove the vectors corresponding to the special tokens, only retain the 1280-dimensional context embedding of the original amino acid residues, use the sliding window block processing for long sequences to ensure context continuity, and at the same time fuse the block features through mean pooling or position weighting, and finally output the sequence feature.

[0074] Step A12, extract the relative solvent accessibility of amino acids and the secondary structure information of the antigen from the three-dimensional structure data, and convert the secondary structure information into a one-hot encoded vector.

[0075] In this embodiment, the relative solvent accessibility (RSA) is obtained by calculating the solvent accessible surface area of the amino acid residue and dividing it by the maximum reference value of the residue type in the fully extended conformation. The secondary structure information describes the local folding pattern of the protein. The one-hot encoded vector is a coding method that converts discrete categorical variables into numerical forms. Each category corresponds to an equal-length binary vector with only a single element being 1 (the rest being 0), which is used to solve the problem of non-ordered categorical feature input, but may cause computational burden or overfitting due to high-dimensional sparsity.

[0076] For example, the one-hot encoded vector uniquely identifies each type of secondary structure with an 8-dimensional binary vector, and the 9-dimensional structural feature vector is composed of a 1-dimensional RSA value and an 8-dimensional one-hot encoding, which is used to characterize the exposure degree and local conformation of the residue.

[0077] As an alternative implementation, based on the PDB file of the antigen, use the DSSP tool to parse the solvent accessible surface area (SASA) of each amino acid residue, calculate the normalized relative solvent accessibility through the formula RSA = SASA / Max_SASA using the amino acid type-specific maximum solvent accessibility, and at the same time extract the secondary structure type encoding defined by DSSP, and map the secondary structure type of each residue into an 8-dimensional one-hot encoded vector.

[0078] For example, the secondary structure type encoding can be H = α-helix, E = β-sheet, G = 310-helix, I = π-helix, B = β-bridge, T = hydrogen bond turn, S = bend, space = random coil.

[0079] Step A13: Concatenate the relative solvent accessibility and the one-hot encoded vector along the dimension to obtain the structural feature.

[0080] In this embodiment, dimension concatenation means connecting the RSA and the one-hot encoding along the feature axis to form a 9-dimensional vector. The structural feature is the final generated multi-dimensional numerical representation that fuses the residue exposure degree and the local conformation, and is used to input into the graph neural network to model the antigen functional region.

[0081] As an alternative implementation, the RSA value of each residue is concatenated with the secondary structure one-hot encoding along the feature dimension to generate a structural feature vector. After checking the consistency between the residue numbers in the file and the DSSP output, the feature vectors are arranged in the residue order and saved as the structural feature.

[0082] For example, it can be to concatenate the RSA value (1-dimensional) of each residue with the secondary structure one-hot encoding (8-dimensional) along the feature dimension to generate a 9-dimensional structural feature vector.

[0083] It should be noted that step A11, step A12, and step A13 are executed in parallel.

[0084] Exemplarily, taking antigen antigen_A (FASTA sequence: MALLHSARVLSGVASAFH PGLAAAAS[SEP]) as an example, using the SaProt tokenizer (supporting Unigram tokenization), it is converted into a token ID sequence [0, 12, 5, 8, 8,... 1], padded to a length of 1024 (padding with [PAD] at the end), and then input into the SaProt model. The last layer hidden state (shape = [1024, 1280]) is extracted and the vectors corresponding to the special tokens are removed, retaining the 1280-dimensional embedding of the original residues (shape = [24, 1280]). At the same time, the SASA of each residue (such as the ) and the secondary structure type (such as H = α-helix) of the first residue M are calculated by DSSP. The RSA is obtained as 85 / 113 ≈ 0.75 after normalizing with alanine, and the secondary structure H is converted into a one-hot encoding [1, 0, 0, 0, 0, 0, 0, 0]. Finally, the RSA and the one-hot encoding are concatenated to form a structural feature vector [0.75, 1, 0, 0, 0, 0, 0, 0, 0], and a structural feature matrix (shape = [24, 9]) is generated in the residue order. After aligning with the sequence feature matrix, it is used as the antigen graph node input.

[0085] By extracting structural features and using a protein language model for sequence feature extraction, the simultaneous fusion of sequence features and structural features improves the performance of protein representation in prediction and model training.

[0086] Based on any of the above embodiments, in the third embodiment of the present application, referring to Figure 3 , Figure 3 is a schematic flowchart of the third embodiment of the B-cell epitope prediction method based on a geometric graph network of the present application. The step S20 includes steps B11 to B14:

[0087] Step B11, based on the coordinate matrix of each antigen, and combining the spatial structure and shape information of the antigen in the three-dimensional structure data, model the antigen as the antigen graph.

[0088] In this embodiment, the coordinate matrix is a numerical matrix storing the three-dimensional coordinates of each atom in the antigen. The spatial structure describes the arrangement of antigen atoms in three-dimensional space. The shape information includes the geometric features on the surface of the antigen and the overall morphology. Modeling refers to abstracting the antigen into a graph structure based on atomic coordinates and shape features. The antigen graph is a topological model of the antigen represented by graph theory, where the node attributes include sequence embedding and structural features, and the edge attributes define the type of interaction or distance threshold between residues, and are used for graph neural network analysis of epitope distribution and functional site prediction.

[0089] As an alternative implementation, extract all Cα atom coordinates from the PDB file of the antigen to construct a three-dimensional coordinate matrix, add spatial adjacent edges based on the Cα atom distance threshold, the edge index matrix of the graph and attach edge attributes; combine the sequence features, structural features and geometric features obtained by preprocessing, splice them into a node feature matrix in residue order, encapsulate the node features, edge index and edge attributes, and after verifying the consistency between the number of nodes N and the number of residues in the PDB file, save it as a standardized antigen graph.

[0090] For example, it can be to combine the sequence features (1280 dimensions), structural features (9 dimensions) and geometric features (geometric node features and geometric edge features) obtained by preprocessing, splice them into a node feature matrix in residue order, encapsulate the node features, edge index and edge attributes (such as distance and type), and after verifying the consistency between the number of nodes N and the number of residues in the PDB file, save it as a standardized antigen graph.

[0091] Step B12, based on the local coordinate system, compare the distance, direction and angle characteristics between atoms inside the node, and obtain geometric node features by calling the node function.

[0092] In this embodiment, the local coordinate system is a 3D reference system (x = normal vector, y = principal curvature direction, z = tangential direction) established with the Cα atom of each residue as the center. The geometric node feature is to fuse the sequence feature and the structure feature in the local coordinate system and call the node function to generate high-dimensional geometric node features for characterizing the sequence-structure-space joint attributes of residues.

[0093] As an alternative implementation, a local coordinate system is constructed with the Cα atom coordinates of each residue of the antigen as the center (the x-axis is the normal vector direction, the y-axis is along the principal curvature direction, and the z-axis is perpendicular to the x-y plane). By calculating the distance, direction, and angle characteristics of the residues, the geometric node features are defined, and the features are spatially transformed in the local coordinate system, and geometric operations are superimposed to output a geometric node feature vector that fuses sequence semantics, structural attributes, and spatial topological relationships. After being standardized, it is input into the graph convolutional layer together with the edge features, and finally a unified node representation matrix of the antigen graph dataset is generated.

[0094] Step B13: Based on each of the antigen graphs, by comparing the Euclidean distance between each pair of nodes with a distance threshold and calling the edge function, geometric edge features are determined.

[0095] In this embodiment, the Euclidean distance is calculated by computing the straight-line distance of the three-dimensional coordinates (x, y, z) of the Cα atoms of two nodes. The distance threshold is a preset spatial proximity determination value. The geometric edge features include the position, distance, direction, and orientation of the edge, which are used for the neural network to learn the spatial interaction patterns between residues.

[0096] As an alternative implementation, the three-dimensional coordinates of the Cα atoms of each residue are extracted from the node coordinate matrix of the antigen graph. The k-nearest neighbors of each node are searched through an algorithm and their Euclidean distances are calculated. A distance threshold is preset, and covalent edges are automatically generated between adjacent residues. The edge function is called to convert the distance into a normalized edge weight, and at the same time, the curvature difference between the two endpoints is calculated as the geometric edge feature.

[0097] Step B14: Merge the geometric node features and the geometric edge features into the geometric features.

[0098] In this embodiment, the merging is to input the node features and the edge features as the node attributes and edge attributes of the graph data into the graph convolutional network to obtain the geometric features.

[0099] As an alternative implementation of obtaining geometric features, in order to more fully capture the spatial structure and shape information of the antigen, we performed in-depth geometric feature extraction based on the coordinate matrix of each antigen, thereby comprehensively characterizing the geometric information of the antigen nodes and edges. Considering that the spatial relationship of amino acids has rotational and translational invariance, we defined a local coordinate system centered on Cα. As shown in formulas (3) and (4), for each residue A iFirst, through three coordinates calculate the vector u i , v i . Then calculate the negative bisector b of the angle formed by the N, Cα, and C atoms i . Next, calculate the normal vector m of the plane formed by the N, Cα, and C atoms i . Then, the local three-dimensional coordinate system Q i = [b i , m i , b i × m i can be constructed. Based on this coordinate system, we define geometric node features and geometric edge features that satisfy rotation and translation invariance.

[0100]

[0101] b i = Normalize(u i - v i ) ; m i = Normalize(u i × v i )#(4)

[0102] Among them, Normalize is a normalization function, which is a method of scaling data to a unified range (such as [0, 1] or [-1, 1]) or adjusting its distribution (such as mean 0 and variance 1) according to specific rules to eliminate the dimension difference and improve the model stability and convergence speed.

[0103] Geometric node features: For residue A i , we calculate its distance, direction, and angle characteristics. The distance feature is embedded through the radial basis function RBF (||x ij - x iz ||), where j ≠ z, which defines the distance relationship between different atoms within the residue; the direction feature is defined as calculating the direction of atoms other than the Cα atom in the residue relative to the Cα atom in the local coordinate system; the angle feature is defined as the sine and cosine values of the dihedral angles (φ i , ψ i , ω i ) and bond angles (α i , β i , γ i ), so as to fully reflect the geometric information of the backbone. These features help to capture the local and global information of the graph structure. Combining these features with sequence features can further enhance the expression ability of the nodes.

[0104] Geometric edge features: For each antigen graph, calculate the Euclidean distance between all nodes based on the Cα atom coordinates, and then set a distance threshold. An edge exists between nodes whose Euclidean distance is within the threshold. This threshold is determined according to the performance of the model on the training dataset and is finally set to We calculated the distance, direction, and orientation characteristics of each edge. The distance feature is obtained by calculating the radial basis function RBF(||x i -x j ||) of adjacent nodes; the direction feature is defined as The orientation feature shows the relative spatial rotation information between two nodes and is defined as where q is the quaternion encoding function, which can represent a 3D rotation matrix as a quaternion vector. These edge features can help the model better learn the spatial relationship and structural information between nodes.

[0105] Exemplarily, the SaProt language model is trained on approximately 40 million protein sequence and structure data. By combining sequence phrases with structure phrases generated from the three-dimensional structures of proteins processed by the Foldseek (Structure-based Protein Alignment and Search Tool), a "structure-aware vocabulary" is introduced, which greatly improves the model's understanding of structural information. In addition, SaProt adopts a self-supervised learning method and is trained using masked prediction, enabling it to learn effective feature representations from a large amount of data under unsupervised conditions. Different from protein language models such as ProtTrans (Protein Transformer) and ESM-2 (Evolutionary Scale Modeling-2), SaProt further improves the performance of protein representation in downstream tasks by simultaneously integrating sequence and structure information.

[0106] By capturing the local and global information of the graph structure, the comprehensive description of node and edge features improves the accuracy of epitope prediction.

[0107] Based on any of the above embodiments, in the fourth embodiment of the present application, step S30 includes steps C11 to C12:

[0108] Step C11, splice the antigen graph with the corresponding sequence features, structure features, and geometric features of the antigen graph to obtain antigen graph data.

[0109] In this embodiment, splicing refers to stacking sequence features, structural features, and geometric node features along the feature dimension as node attributes, independently retaining edge features, and jointly encapsulating them into an antigen graph data set.

[0110] For example, splicing refers to stacking sequence features, structural features, and geometric node features along the feature dimension (1280 + 9 + 256 = 1545 dimensions) as node attributes.

[0111] As an alternative implementation, an antigen graph data object is constructed based on the geometric expansion library framework. The sequence feature matrix, structural feature matrix, and geometric node feature matrix are concatenated in residue order into a node attribute matrix. The geometric edge feature matrix serves as the edge attribute, and the edge index matrix is generated through dimension constraints and covalent connection rules. After verifying the consistency of the number of validation nodes with the number of residues and sequence length, the node features are standardized and it is ensured that there are no duplicate or isolated nodes in the edge index. Finally, by encapsulating the node features, edge index, edge attributes, and global attributes, it is saved as antigen graph data.

[0112] Step C12, integrating the antigen graph data to obtain the antigen graph data set.

[0113] In this embodiment, integrating the antigen graph data to obtain the antigen graph data set means storing the objects of multiple antigen graph data in a unified list or encapsulating them as a batch tensor.

[0114] As an alternative implementation, use the data collection class of PyTorch Geometric to construct an antigen graph data set. Traverse the files and preprocessed features of all antigens, load the objects of each antigen in turn, store them in a list after checking the consistency of the feature dimensions, convert the list data into a batch tensor through the function of the in-memory data set, add global attributes such as antigen ID and label, standardize the node features across antigens, divide the training set, validation set, and test set, and then save them as standardized files, supporting batch loading through a data loader, ensuring compatibility of mask processing and dynamic padding for heterogeneous antigen graphs, and finally generating a standardized antigen graph data set that can be directly input into graph neural networks such as I3NN.

[0115] Exemplarily, an antigen can be described in multiple complementary forms, including amino acid sequences, residue graphs, atomic point clouds, and molecular surfaces, etc. Each representation method can capture different functional features. By utilizing the three-dimensional structure of the antigen, an antigen is modeled as a graph, where amino acids represent the nodes in the graph, the Cα atom coordinates are defined as the positions of the nodes, and nodes within a set threshold distance are adjacent nodes and are connected by edges. To comprehensively describe the features of nodes and edges, we define multiple types of features from both the sequence and structure aspects. By constructing this graph structure, the model can effectively capture the complex relationships between different nodes in the antigen, transforming the epitope prediction problem into a graph node classification task.

[0116] Since a graph model is constructed based on the antigen structure and detailed geometric information of the graph structure is calculated, it helps to deeply analyze the overall structure of the graph and the internal connections between its various parts, thus effectively solving the problem of predicting complex epitope conformations and improving the reliability of the prediction method.

[0117] Based on any of the above embodiments, in the fifth embodiment of the present application, step S40 includes steps D11 to D14:

[0118] Step D11, based on the coordinate matrix of the antigen, determine the coordinates of each atom in the antigen and the central coordinates of each side-chain atom.

[0119] In this embodiment, the coordinate matrix is a numerical array (with shape [N, 3], where N is the number of atoms) storing the three-dimensional spatial coordinates (x, y, z) of all atoms of the antigen. Atomic coordinates refer to the position of a single atom in three-dimensional space (such as the coordinates of the Cα atom). Side-chain atoms are the atoms other than the main chain in an amino acid residue. The central coordinates are obtained by calculating the geometric mean of the side-chain atom coordinates and characterize the spatial distribution center of the side chain.

[0120] As an alternative implementation, extract the coordinates of all atoms from the PDB file of the antigen to construct a coordinate matrix group, screen the side-chain atoms for each residue, and if there are no atoms in the side chain, skip it; otherwise, calculate the mean of its coordinates to generate the side-chain central coordinates of each residue.

[0121] For example, it can be to extract the coordinates of all atoms from the PDB file of the antigen to construct a coordinate matrix (with shape [N, 3]), group it by residue ID (such as sequence position), screen the side-chain atoms for each residue, and if there are no atoms in the side chain (such as glycine), represent it with the Cα atom coordinates; otherwise, calculate the mean of its coordinates to generate the side-chain central coordinates of each residue.

[0122] Step D12, based on the coordinates of each atom and the central coordinates of each side-chain atom, call a vector function to calculate and obtain vector features.

[0123] In this embodiment, the vector feature is the dMaSIF feature, which is an approximate vector representation of surface curvature.

[0124] As an alternative implementation, define each antigen coordinate matrix as X, and define the coordinates of each residue A i as x i . x i is a 5x3-dimensional tensor, including the coordinates of N, Cα, C, O atoms and the central coordinates of side-chain atoms (R groups). This method uses the Cα atom coordinates of each amino acid to obtain the dMaSIF feature at the residue level.

[0125]

[0126] Among them, SDF represents the "soft distance" function between residues. For each point calculate its Euclidean distance from all other points and amplify the influence of long distances through an exponential function. n is the dMaSIF feature representing the normal vector of the antigen surface, which is an approximate vector representation of the surface curvature. Expand the partial derivative of SD F with respect to i and perform normalization processing.

[0127] Step D13: Obtain the target antigen map dataset according to the vector feature and the coordinate matrix.

[0128] As an alternative implementation, assign the vector features of each antigen to the antigen map node attributes in the residue order, extract the Cα coordinates from the coordinate matrix to construct the node spatial positions, generate the edge indices by using the spatially adjacent nodes, calculate the edge attributes and assign them to the edge attributes. After verifying the consistency between the node feature dimensions and the edge indices, merge all the antigen map data into a unified target antigen map dataset through the in-memory dataset.

[0129] For example, calculating the edge attributes can be the assignment of types, normalized distances, and curvature differences.

[0130] Step D14: Determine the predicted epitopes based on the target antigen map dataset.

[0131] In this embodiment, the predicted epitope is the probability that the antigen surface residue output by the model belongs to the B cell epitope, which needs to be binarized by a threshold, and its spatial distribution needs to be aligned and verified with the experimental data.

[0132] For example, the threshold can be binarized to 0 (non-epitope) or 1 (epitope).

[0133] As an alternative implementation, load the pre-trained I3NN model, batch load the target antigen map data through a data loader, obtain the epitope probabilities of each residue through the forward propagation of the input model, generate binary prediction labels after normalization by setting a threshold, and combine with the coordinate file of the antigen to filter out isolated residues and merge adjacent epitope regions through spatial clustering, and output the prediction results of the epitopes.

[0134] ​Exemplarily, we compared the performance of the method with GraphBepi on the GraphBepi dataset. The method proposed in this paper shows significant performance improvement in all metrics. Among them, on the test set, the AUPRC (Area Under the Precision-Recall Curve) value reaches 0.32, which is 6% higher than GraphBepi; the AUC (Area Under the Receiver Operating Characteristic Curve) reaches 0.81, which is 5% higher than GraphBepi; in terms of the F1 value and MCC (Matthews Correlation Coefficient), both exceed 4%. These results indicate that the method of the present invention has obvious advantages in capturing the key features of antigenic epitopes and prediction accuracy, and can more accurately identify epitope regions. On this basis, we pre-trained the model using the protein-protein interaction site data of PeSTo (Protein Binding Interface Predictor, Protein Structure Transformer) and fine-tuned the model through transfer learning. The experimental results show that the pre-training strategy further improves the performance of the model in the epitope prediction task. In addition, we directly applied the pre-trained model to the test set of protein-protein interaction sites, and the model also showed excellent performance. This not only verifies the effectiveness of the pre-training strategy but also reflects the scalability of the model of the present invention in cross-task scenarios.

[0135] Since the dihedral angle is calculated through vector features and the dihedral angle has rotational and translational invariance, it ensures that the model can stably handle different spatial transformations, improving the accuracy and stability of the prediction method.

[0136] Based on any of the above embodiments, in the sixth embodiment of the present application, referring to Figure 4 , Figure 4 is a schematic flowchart of the sixth embodiment of the B cell epitope prediction method based on the geometric graph network of the present application. Step D13 includes steps E11 to E14:

[0137] Step E11, based on the vector features, call the dihedral angle calculation formula to calculate the dihedral angle between the normal vectors of two residues and the direction vector connecting the residues, and input the dihedral angle into the mapping formula to obtain the input features.

[0138] In this embodiment, the dihedral angle calculation formula calculates the angle between two normal vectors through vector cross product and dot product:

[0139]

[0140] The input feature is a numerical vector that fuses dihedral angle encoding and the original vector feature.

[0141] For example, the input feature can be a 21-dimensional numerical vector that fuses dihedral angle encoding (4-dimensional) and the original vector feature (such as 17-dimensional).

[0142] As an alternative implementation, based on the residue pairs in the antigen map data, the normal vectors and Cα coordinates of two residues are extracted from the node attributes, the connecting vector is calculated, the included angle θ is obtained through cross product and dot product, and the azimuth angle φ is calculated using the arctangent function combined with the projection on the normal vector plane. Then, θ and φ are input into the mapping formula to generate a four-dimensional periodic encoding, which is concatenated with the original vector features to form the input features.

[0143] Step E12: Based on the input features and the coordinate matrix, perform convolution on the node features and edge features, and perform message passing through spherical harmonic functions to integrally obtain the target node features.

[0144] In this embodiment, the spherical harmonic functions are a set of orthogonal functions defined on the sphere, which can effectively capture the angular information on the sphere. Convolution refers to three-dimensional geometric convolution based on the graph structure, and the learnable spherical harmonic function coefficients are used as the convolution kernels. Message passing is the process of exchanging and aggregating features between nodes along the edges in the graph neural network. The target node features are the node embeddings generated after multiple layers of geometric convolution and message passing, which fuse local structure, global topology, and direction-sensitive features. The node features include sequence features, DSSP features, and geometric node features, and the edge features include geometric edge features.

[0145] As an alternative implementation for obtaining the target node features, we calculate the dihedral angle through dMaSIF features, as shown in formula (5). Specifically, we calculate the angle between the normal vectors of two residues and the direction vector connecting the residues. This angle has rotational and translational invariance, ensuring that the model can stably handle different spatial transformations. The spherical harmonic functions are a set of orthogonal functions defined on the sphere, which can effectively capture the angular information on the sphere. By mapping the dihedral angle to the Hilbert space, input features with invariance are obtained, as shown in formula (6). This mapping enables the convolution operation to be performed without destroying the geometric invariance.

[0146]

[0147] Y l (α,β) = S l (α)P l (cosβ)#(6)

[0148] where α and β are the calculated dihedral angles; S l is an n-dimensional sphere, P l is the Legendre polynomial, Y l (α,β) is the output of the spherical harmonic function for the angular coordinates α and β. The variation law of the function in the polar angle direction is defined by P l (cosβ), and then through S l(α) Control the periodicity of the function in the azimuth direction, multiply the modulation results of the two angles, and obtain the complete spatial dependence Y l (α,β).

[0149] For an antigen containing l amino acids, the input of the I3NN module includes the graph structure of the node feature matrix H and the edge feature matrix E, as well as the antigen coordinate matrix X. Inside the I3NN module, the node feature vectors and edge feature vectors of the graph are convolved on the invariance layer, and information is passed through the spherical harmonic function. The method in this paper uses the first-order spherical harmonic function in the four-layer network of I3NN. The convolutional layer operation is shown in Equation (7). By passing information in the spherical harmonic function space, the features of neighboring nodes and edge features are aggregated, enabling the model to effectively integrate local geometric information.

[0150]

[0151] where h i and h j represent the feature vectors of nodes i and j, e ij represents the edge feature vector between nodes i and j, CONCAT represents the concatenation operation; NB(i) represents the neighbors of node i, and z represents the degree of node i. After obtaining the local geometric information, to capture the global information, by calculating the mean value of the node features in each batch, the comprehensive feature of the edge (i,j) (concatenating the edge feature e ij with the features of the two end nodes h i , h j ), the edge feature e ij , the source node feature h i , and the target node feature h j are concatenated (CONCAT) to form the comprehensive edge representation E ij . Through the learnable function f, a non-linear transformation is performed on E ij to generate the weight vector f(E ij ). Using the first-order spherical harmonic function Y l to encode the geometric direction information Y l (α,β) between nodes i and j, sum the weighted features of all neighbors, and normalize through . Add the aggregation result to the original feature h i to obtain the updated feature h i ′. As shown in Equation (8), a global context vector is learned for each antigen, and then the node features are updated to obtain the target node features.

[0152]

[0153] where MLP2 represents the multi-layer perceptron, σ represents the Sigmoid activation function, h idenote the original feature vector of node i, and c is the context information vector, usually obtained by aggregating the features of neighbor nodes, such as MLP2 is a multi-layer perceptron used to generate gating weights. The σ Sigmoid activation function compresses the output to the interval [0, 1], h′ k is the feature of neighbor node k, and l is the number of neighbor nodes. The aggregated context vector c is input into MLP2, and an intermediate vector is generated through a non-linear transformation. The output is mapped to the interval [0, 1] using the Sigmoid function to obtain the gating weight σ(MLP2(c)).

[0154] Step E13: Concatenate the edge feature and the target node feature, and process them through a linear transformation and an activation function to obtain the target edge feature.

[0155] In this embodiment, the target edge feature is the edge attribute generated after the above processing, which is used to guide the message passing weight assignment in the graph attention mechanism or directly used as the input for edge classification. The linear transformation maps the high-dimensional features to the low-dimensional space through a fully connected layer. The activation function enhances the feature expression ability by introducing non-linearity.

[0156] As an alternative implementation, after obtaining the updated node feature representation, the edge feature is further updated by using the two nodes connected by the edge to achieve message passing. The specific process is as shown in formula (9). We concatenate the source node feature, the edge feature, and the target node feature, and then process them through a linear transformation and an activation function to obtain the updated edge feature.

[0157] e i ′ j = e ij + MLP1(h j ′ || e ij || h i ′) #(9)

[0158] where, e ij is the original feature vector of edge (i, j), || represents the concatenation operation, MLP1 represents the multi-layer perceptron, which concatenates the edge feature e ij with the features of its two end nodes h j ′ (source node) and h i ′ (target node) to form a comprehensive input vector (h j ′ || e ij || h i ′), adds the original edge feature e ij to the incremental feature to achieve feature update and obtain e i ′ j .

[0159] Step E14: In the antigen map dataset, update the node features and edge features to the target node features and target edge features to obtain the target antigen map dataset.

[0160] In this embodiment, the update is to replace the node feature edge features in the antigen map dataset with the target node features and target edge features.

[0161] As an alternative implementation, based on a graph neural network, assign the target node feature matrix, retain the original edge index and global attributes. After verifying the consistency of the number of nodes N and the number of edges M, perform cross-antigen standardization on the node features and normalization on the edge features. Traverse all antigen map data through the attributes of the in-memory dataset and replace the features, and finally encapsulate them into the updated target antigen map dataset.

[0162] For example, it can be a Data class based on a graph neural network. Assign the target node feature matrix to Data.x and the target edge feature matrix to Data.edge_attr. Retain the original edge index and global attributes. After verifying the consistency of the number of nodes N and the number of edges M, perform Z-score standardization (mean = 0, variance = 1) on the node features across antigens, and normalize the edge features (scale to 0-1). Traverse all antigen map data through the data_list attribute of the in-memory dataset and replace the features, and finally encapsulate them into the updated target antigen map dataset.

[0163] Exemplarily, taking antigen antigen_C (PDB file: antigen_C.pdb, FASTA sequence: MGARS...) as an example, parse its PDB file to extract the Cα coordinate matrix (shape [200, 3]) of each residue and the side chain atom coordinates, calculate the normal vector (based on local surface fitting) and connection direction vector between residue 25 (Lys) and residue 30 (Glu), call the dihedral angle formula (cross product and dot product) to calculate the included angle θ = 1.2 rad and azimuth angle φ = 0.8 rad, and map them to a 4D periodic encoding [sinθ, cosθ, sinφ, cosφ], which is concatenated with the original 17D geometric vector to generate a 21D input feature; input the coordinate matrix and input feature into the geometric convolution layer based on the e3nn library, expand the relative position direction through spherical harmonic functions, and calculate the neighborhood message (5D spherical harmonic basis × 8D radial basis) to generate 256D node features. After concatenating the node features (256D) with the geometric features at both ends of the edge (each 21D), generate the target edge features through a fully connected layer (512 → 128D + GELU activation); merge the updated node features (256D) with the edge features (128D + 3D original edge attributes) into the antigen map data, replace the geometric features in the original dataset, and save them as the standardized dataset after standardization for the I3NN model to load and train.

[0164] Since a graph model is constructed through the antigen structure and the detailed geometric information of the graph structure is calculated, it helps to deeply analyze the overall structure of the graph and the internal connections between its various parts, thus effectively solving the problem of predicting complex epitope conformations and improving the accuracy of the prediction method.

[0165] Based on any of the above embodiments, in the seventh embodiment of the present application, step D14 includes steps F11 to F12:

[0166] Step F11, the target neural network extracts antigen representations based on the target antigen graph dataset.

[0167] In this embodiment, the antigen representation is an embedded vector at the antigen residue level or global level extracted by the model.

[0168] As an alternative implementation, load the target antigen graph dataset into the data loader of the graph neural network, input the pre-trained target neural network, in the forward propagation, the node features are weighted and aggregated with neighborhood messages through the graph attention layer, and multi-scale features are retained through skip connections and layer normalization, and finally the antigen representation at the node level is output.

[0169] Step F12, based on the antigen representation, weight matrix and bias term, and calling a logical function, determine the predicted epitope.

[0170] In this embodiment, the logical function is used to calculate the probability score of epitope prediction. The weight matrix is a learnable parameter of the logistic regression layer, and the bias term is a constant offset of the classifier.

[0171] As an alternative implementation, the antigen representation learned by the multi-layer perceptron module, and finally the epitope probability of each amino acid is obtained through an MLP as follows:

[0172] Y = Sigmoid(H (L) W + b)

[0173] where H (L) is the antigen representation output by the last layer, W is the weight matrix, and b is the bias term; the Sigmoid function is used to normalize the network output to the probability score of the epitope.

[0174] Exemplarily, to verify the role of the modules used in this method, we conducted a systematic ablation experiment on the ScanNet dataset. The experimental results show that the introduction of different modules has a significant impact on the model performance. Specifically: from not using sequence information at all, to using one-hot encoding, then to introducing the ESM-2 language model, and finally to adopting the SaProt language model, the model performance gradually improves. This trend indicates that rich sequence information makes an important contribution to the epitope prediction task. In particular, the SaProt language model achieves a better feature representation effect than the ESM-2 language model by combining structural information, further improving the model performance. Without using antigen structure information, the model performance drops significantly and reaches the worst level, which fully proves the indispensability of structural information for accurate epitope prediction. The experimental results show that the model combining sequence information, structural information, and geometric information can achieve high performance. This multi-feature fusion strategy makes full use of the complementarity of different data modalities and significantly improves the model's prediction ability.

[0175] Since the geometric perception ability of the model is enhanced by the I3NN module, and by integrating local and global information, the accuracy of this prediction method is significantly improved.

[0176] Based on any of the above embodiments, in the eighth embodiment of the present application, before step S40, steps G11 to G14 are further included:

[0177] Step G11, dividing the antigen graph dataset by cross-validation, taking one fold of the antigen graph dataset as the validation set, and the remaining antigen graph datasets as the training datasets.

[0178] In this embodiment, cross-validation is an evaluation method that divides the dataset into k subsets, and takes one fold as the validation set to evaluate the model in turn, and the remaining k - 1 folds as the training set to train the model, so as to reduce the evaluation bias caused by data division.

[0179] As an optional implementation manner, use the random splitting function of the graph neural network to randomly divide the antigen graph dataset into k folds, set a random seed to ensure reproducibility, and specify the antigen graph corresponding to the current fold index as the validation set when traversing each fold, and the rest as the training set.

[0180] For example, use the KFold (k-fold cross-validation) or RandomSplit (random splitting) function of the graph neural network to randomly divide the antigen graph dataset into k folds (such as k = 5), set a random seed (seed = 42) to ensure reproducibility, and specify the antigen graph corresponding to the current fold index (such as fold = 0) as the validation set (accounting for 20%), and the remaining 4 folds (80%) as the training set when traversing each fold.

[0181] Step G12: Train the basic neural network using the training dataset to obtain an output result.

[0182] In this embodiment, the basic neural network refers to an untrained initial model architecture. Training is a process of optimizing model parameters by minimizing the loss function through the backpropagation algorithm. The output result is the predicted value of the model for the input antigen map.

[0183] As an alternative implementation, load the training dataset into a data loader, initialize the basic neural network, set the optimizer and loss function, input the antigen map data batch by batch, perform forward propagation to calculate the prediction probability, and obtain the output result.

[0184] Step G13: Adjust the parameters of the basic neural network by backpropagating according to the deviation between the output result and the true epitope label.

[0185] In this embodiment, the true epitope label is a binary label verified by experiments. The deviation quantifies the prediction error through the loss function. Backpropagation calculates the gradient of the loss with respect to the network parameters based on the chain rule. Adjusting the parameters updates the weights in the gradient direction by the optimizer to minimize the loss.

[0186] As an alternative implementation, call the automatic differentiation module of the neural network to calculate the gradient of the loss function with respect to the model parameters, use the optimizer to perform gradient descent to update the parameters, enable mixed precision acceleration calculation during training, clip the gradients to prevent explosion, record the loss and metrics, and adjust the parameters of the basic neural network according to the records.

[0187] Step G14: Evaluate the performance of each basic neural network according to the validation set, and set the basic neural network with the highest performance as the target neural network.

[0188] In this embodiment, the performance evaluation quantifies the matching degree between the model prediction and the true epitope label through preset metrics. The highest performance means the best metrics in the validation set. The target neural network is the selected optimal model for final inference or deployment.

[0189] As an alternative implementation of model training, the acquisition of the training antigen map dataset is to first obtain the amino acid sequence and three-dimensional structure data of the antigen from the open-source dataset. The subsequent processing process is similar to that of Embodiment 1 and will not be elaborated here. After obtaining the training antigen map dataset, divide the training antigen map dataset into k subsets. Use one fold as the validation set to evaluate the model, and the remaining k - 1 folds as the training set to train the model. Train the model by inputting it into the I3NN model through the training set, then verify the output result with the validation set, and finally select the one with the best prediction result as the target neural network.

[0190] As an optional architectural approach for the B-cell epitope prediction method based on the geometric graph network, refer to Figure 5 , Figure 5 is the architecture diagram of the B-cell epitope prediction method based on the geometric graph network in this application. The model mainly consists of three core modules: the Featurizer module, the I3NN module, and the Multi-Layer Perceptron (MLP) module. The Featurizer module is responsible for converting the antigen sequence and structure into a feature matrix. For each amino acid in the antigen sequence with length l, it generates 1280-dimensional SaProt features, 9-dimensional DSSP features, and 184-dimensional geometric features, and finally obtains an lx1473-dimensional node feature matrix; at the same time, it generates 450-dimensional geometric features for each edge. The I3NN module is used to process the input graph representation and dynamically update the feature representations of nodes and edges by performing convolutional operations on the rotation and translation invariant layers. After sufficient message passing, the node features containing local and global information of the antigen are input into the MLP module, and finally the epitope probability score of each residue is predicted.

[0191] Exemplarily, the cross-validation method is adopted to train and evaluate the model to ensure the generalization ability and stability of the model. Specifically, we select one fold as the validation set to evaluate the model performance each time, and the remaining folds as the training set for model training. During the training process, by monitoring the performance metrics on the validation set, the early stopping strategy is adopted to prevent the model from overfitting. In addition, based on the performance of the model, we further optimize the feature combination and the selection of hyperparameters to achieve the best prediction performance. After multiple experiments, the following hyperparameter settings are finally determined: the number of convolutional layers of the I3NN model is set to 4; the batch size is set to 8; the dimension of the hidden layer is set to 64; the number of training iterations of the model is set to 100 rounds. The model training uses the Adam optimizer, and its learning rate is set to 5e-4; the binary cross-entropy loss (BCE) is selected as the loss function. In terms of technical implementation, this method uses Python version 3.8.16 and Pytorch version 1.13.1 to build the model.

[0192] Since the cross-validation method is used to train and evaluate the model, the generalization ability and stability of the model are ensured, and the accuracy and reliability of the prediction method are improved.

[0193] This application provides an epitope prediction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the B-cell epitope prediction method based on the geometric graph network in the first embodiment above.

[0194] Next, refer to Figure 6, which shows a schematic structural diagram of an epitope prediction device suitable for implementing the embodiments of the present application. The epitope prediction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown epitope prediction device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0195] As Figure 6 shown, the epitope prediction device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the epitope prediction device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the epitope prediction device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an epitope prediction device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0196] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0197] The epitope prediction device provided in the present application adopts the B cell epitope prediction method based on the geometric graph network in the above embodiments, and can solve the technical problem that in the related art, for linear epitope prediction, it is difficult to comprehensively capture the multi-level characteristics of protein interactions, thereby resulting in low B cell epitope prediction performance. Compared with the prior art, the beneficial effects of the epitope prediction device provided in the present application are the same as those of the B cell epitope prediction method based on the geometric graph network provided in the above embodiments, and other technical features in the epitope prediction device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.

[0198] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0199] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0200] The present application provides a computer-readable storage medium, which has computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the B cell epitope prediction method based on the geometric graph network in the above embodiments.

[0201] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0202] The above computer-readable storage medium may be included in the epitope prediction device; or it may exist separately and not be assembled into the epitope prediction device.

[0203] The above computer-readable storage medium carries one or more programs, which, when executed by the epitope prediction device, cause the epitope prediction device to: extract the sequence features corresponding to the antigen sequence and the structural features corresponding to the three-dimensional structure data of the protein; determine the geometric features of the antigen based on the antigen map constructed from the three-dimensional structure data; splice the geometric features, the sequence features, and the structural features into an antigen map dataset; and process the antigen map dataset based on a pre-trained target neural network to determine the predicted epitope.

[0204] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0206] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0207] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned B-cell epitope prediction method based on the geometric graph network, which can solve the technical problem that in the related art, it is difficult to comprehensively capture the multi-level features of protein interactions for linear epitope prediction, thereby resulting in low performance of B-cell epitope prediction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the B-cell epitope prediction method based on the geometric graph network provided in the above embodiments, and will not be elaborated here.

[0208] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.

Claims

1. A B-cell epitope prediction method based on a geometric graph network, characterized in that, The method includes: extracting sequence features corresponding to the antigen sequence and extracting structural features corresponding to the three-dimensional structure data of the protein; determining geometric features of the antigen according to the antigen map constructed based on the three-dimensional structure data; concatenating the geometric features, the sequence features, and the structural features into an antigen map dataset; processing the antigen map dataset based on a pre-trained target neural network to determine predicted epitopes.

2. The B cell epitope prediction method based on a geometric graph network according to claim 1, wherein The steps of extracting sequence features corresponding to the antigen sequence and extracting structural features corresponding to the three-dimensional structure data of the protein include: converting the antigen sequence into a vector representation through a protein language model and gradually extracting the sequence features based on the vector representation; and, extracting the relative solvent accessibility of amino acids and the secondary structure information of the antigen from the three-dimensional structure data and converting the secondary structure information into a one-hot encoded vector; concatenating the relative solvent accessibility and the one-hot encoded vector according to dimensions to obtain the structural features.

3. The B cell epitope prediction method based on a geometric graph network according to claim 1, wherein The steps of determining geometric features of the antigen according to the antigen map constructed based on the three-dimensional structure data include: modeling the antigen as the antigen map based on the coordinate matrix of each antigen and combining the spatial structure and shape information of the antigen in the three-dimensional structure data; comparing the distance, direction, and angle characteristics between atoms inside the nodes based on a local coordinate system and obtaining geometric node features by calling a node function; determining geometric edge features based on each antigen map by comparing the Euclidean distance between each node and a distance threshold and calling an edge function; combining the geometric node features and the geometric edge features into the geometric features.

4. The B cell epitope prediction method based on a geometric graph network according to claim 1, wherein The steps of concatenating the geometric features, the sequence features, and the structural features into an antigen map dataset include: concatenating the antigen map with the corresponding sequence features, structural features, and geometric features of the antigen map to obtain antigen map data; integrating each antigen map data to obtain the antigen map dataset.

5. The B cell epitope prediction method based on a geometric graph network according to claim 1, wherein The steps of processing the antigen map dataset based on a pre-trained target neural network to determine predicted epitopes include: determining the coordinates of each atom and the central coordinates of each side-chain atom in the antigen based on the coordinate matrix of the antigen; calling a vector function based on the coordinates of each atom and the central coordinates of each side-chain atom to calculate vector features; obtaining a target antigen map dataset according to the vector features and the coordinate matrix; determining predicted epitopes based on the target antigen map dataset.

6. The B cell epitope prediction method based on a geometric graph network according to claim 5, characterized in that The steps of obtaining a target antigen map dataset according to the vector features and the coordinate matrix include: calling a dihedral angle calculation formula based on the vector features to calculate the dihedral angle between the normal vector of two residues and the direction vector connecting the residues, and inputting the dihedral angle into a mapping formula to obtain input features; performing convolution on node features and edge features based on the input features and the coordinate matrix and performing message passing through a spherical harmonic function to integrate and obtain target node features; Concatenate the edge feature and the target node feature, and process them through linear transformation and activation function to obtain the target edge feature; In the antigen graph dataset, update the node feature and the edge feature to the target node feature and the target edge feature to obtain the target antigen graph dataset.

7. The B cell epitope prediction method based on a geometric graph network according to claim 5, characterized in that The step of determining the predicted epitope based on the target antigen graph dataset includes: The target neural network extracts the antigen representation based on the target antigen graph dataset; Based on the antigen representation, the weight matrix, and the bias term, and invoking the logical function, determine the predicted epitope.

8. The B cell epitope prediction method based on a geometric graph network according to claim 1, wherein Before the step of processing the antigen graph dataset by the pre-trained target neural network to determine the predicted epitope, it further includes: Use cross-validation to divide the antigen graph dataset, use one-fold of the antigen graph dataset as the validation set, and use the remaining antigen graph datasets as the training datasets; Train the basic neural network by using the training datasets to obtain the output result; Adjust the parameters of the basic neural network through backpropagation according to the deviation between the output result and the true epitope label; Evaluate the performance of each basic neural network according to the validation set, and set the basic neural network with the highest performance as the target neural network.

9. An epitope prediction device, characterized in that, The epitope prediction device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the B cell epitope prediction method based on the geometric graph network according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the steps of the B cell epitope prediction method based on the geometric graph network according to any one of claims 1 to 8.