Disease marker epitope prediction and antibody screening method

By integrating multiple technologies, including database establishment, laser cutting, and deep learning models, antigenic epitope information in disease states can be accurately obtained, solving the problem of inaccurate antigenic epitope prediction in existing technologies and improving the accuracy of disease biomarker detection and antibody binding activity.

CN121237224APending Publication Date: 2025-12-30QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511310978.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict and identify specific antigenic epitopes in disease states, resulting in insufficient accuracy in disease biomarker detection.

Method used

By integrating multiple technologies, including establishing an antigenic epitope database, laser capture microdissection, deep convolutional neural networks, molecular dynamics simulation, and fluorescently labeled antibody detection systems, antigenic epitope information of diseased tissues is accurately obtained, and high-binding-activity antigenic epitope sequences are screened out through deep learning models and random forest classifiers.

Benefits of technology

It significantly improved the accuracy of antigen epitope prediction and the reliability of disease biomarker detection, reduced the proportion of false positive predictions, and improved antibody binding affinity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237224A_ABST
    Figure CN121237224A_ABST
Patent Text Reader

Abstract

The invention provides a disease marker epitope prediction and antibody screening method, which belongs to the technical field of disease markers, and comprises the following steps: firstly, establishing a training data set containing a known antigen-antibody compound structure and disease tissue expression data; and obtaining target protein sequence information through liquid chromatography-mass spectrometry analysis and carrying out sequence comparison. And then a deep convolutional neural network is utilized to extract sequence features, and surface exposure sites are identified by combining secondary structure prediction and solvent accessibility analysis. After the features are integrated with sequence evolution conservative properties, a prediction model is constructed by using a random forest classifier. The method comprises the following steps: carrying out molecular dynamics simulation on a prediction result, screening first 10% of candidate sequences through a comprehensive scoring function and K-means clustering, and finally determining an antigen epitope sequence with the strongest binding activity through verification of an antigen chip and a fluorescence labeled antibody system, so that the technical problem that specific antigen epitopes are difficult to accurately predict and recognize in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of disease markers, and particularly relates to a disease marker antigen epitope prediction and antibody screening method. BACKGROUND

[0002] In modern clinical diagnosis, the accuracy of disease marker detection directly affects disease diagnosis and treatment decisions. At present, most disease marker detection methods rely on the recognition of antibodies to specific antigen epitopes, so accurate prediction and identification of specific antigen epitopes are crucial. However, existing technologies have significant challenges in accurately predicting and identifying specific antigen epitopes.

[0003] Traditional antigen epitope prediction methods mainly include sequence-based bioinformatics prediction, structure-based computational methods, and experimental screening methods. Sequence-based prediction methods mainly rely on the physicochemical properties and statistical rules of amino acids, and although they are simple to operate, they generally have low prediction accuracy because they completely ignore the spatial structure information of proteins. Structure-based prediction methods take into account spatial conformation factors, but their scope of application is severely limited because most newly discovered disease marker proteins lack resolved three-dimensional structure data. Experimental screening methods such as phage display technology can obtain epitope sequences with binding activity, but have the disadvantages of long cycle and high cost.

[0004] More critically, existing prediction methods generally ignore the specific changes in protein expression and modification under disease conditions. In disease tissues, proteins may undergo unique post-translational modifications, and their microenvironments also differ significantly from normal tissues. These changes can lead to changes in the exposure state and immunogenicity of certain antigen epitopes, making prediction methods developed based on normal conditions ineffective. For example, in tumor tissues, abnormal protein modifications can mask or expose new antigen epitopes, but existing prediction methods cannot capture these disease-specific information.

[0005] Currently, there is no prediction method that can effectively integrate disease tissue-specific information, protein structure characteristics, and immune recognition dynamics. This leads to a large difference between predicted antigen epitopes and actual immune system-recognized epitopes, ultimately affecting the accuracy of disease marker detection. Therefore, how to accurately predict and identify specific antigen epitopes has become a core technical problem that needs to be solved. SUMMARY

[0006] Therefore, the application provides a disease marker antigen epitope prediction and antibody screening method, which can solve the technical problem that existing technologies cannot accurately predict and identify specific antigen epitopes.

[0007] The application is implemented as follows:

[0008] The first aspect of the application provides a disease marker antigen epitope prediction and antibody screening method, comprising the following steps:

[0009] S10, an antigen epitope database is established, three-dimensional structure data of known antigen-antibody complexes in a protein database is collected, binding site information is recorded, expression pattern and localization characteristic data of target proteins in disease-related tissue sections are collected, and training data sets are formed by integration;

[0010] S20, laser capture microdissection is performed on disease tissues, and the obtained tissues are subjected to proteome analysis using liquid chromatography-mass spectrometry, so as to obtain amino acid sequences and post-translational modification site information of target proteins, sequence alignment is performed between the target protein sequences and the training data sets, and candidate antigen epitope regions are determined;

[0011] S30, a deep convolutional neural network is constructed, sequence data of the candidate antigen epitope regions are input, sequence feature vectors are obtained by training the deep convolutional neural network, secondary structure prediction is performed on the amino acid sequences of the target proteins, structure propensity is evaluated, and a structure prediction probability matrix is output;

[0012] S40, a structure prediction algorithm is used to calculate the relative solvent accessibility area of each amino acid residue in the target protein, the median of the relative solvent accessibility area is used to determine that the residue exposure threshold is 25%, and surface exposed amino acid sites are marked;

[0013] S50, the sequence feature vectors, the structure prediction probability matrix, the surface exposed amino acid site information, and the sequence evolutionary conservation score are integrated, and a random forest classifier is used to construct an antigen epitope prediction model;

[0014] S60, molecular dynamics simulation is performed on the antigen-antibody complexes predicted by the antigen epitope prediction model, the potential energy evolution trajectory of the antigen-antibody complexes is calculated, and a comprehensive scoring function including electrostatic interaction energy, hydrogen bond number and hydrophobic contact area is established;

[0015] S70, the prediction results of the antigen epitope prediction model are sorted based on the comprehensive scoring function, and the top 10% of antigen epitope sequences with the highest scores are selected as preferred antigen epitope sequences by using a K-means clustering algorithm;

[0016] S80, an antigen chip of the preferred antigen epitope sequence is prepared, a fluorescence-labeled antibody detection system is used to evaluate the binding activity of each preferred antigen epitope sequence and known antibodies, and the antigen epitope sequence with the strongest binding activity is selected as the final antigen epitope based on the signal-to-noise ratio value.

[0017] The three-dimensional structure data of the known antigen-antibody complex in the protein database (such as PDB) is collected, and the binding site information is recorded, specifically including recording the atomic coordinate information, intermolecular contact residue information, hydrogen bond network information, interface hydrophobic interaction information and spatial conformation information around the binding site in the crystal structure of the antigen-antibody complex;

[0018] The expression pattern and localization characteristic data of the target protein in the disease-related tissue section are collected, and the training data set is integrated, specifically including recording the expression level, subcellular localization, tissue distribution pattern and co-localization relationship with other markers of the target protein in the diseased tissue, and integrating these data with the structure information to form a multi-dimensional feature training data set;

[0019] The laser capture microdissection is performed on the disease tissue, specifically identifying the region of interest under a microscope, cutting along the boundary of the target region using infrared laser, and ejecting the cutting region into a collection tube using ultraviolet laser, and extracting the protein of the target region through lysis buffer;

[0020] The tissue obtained by cutting is subjected to proteome analysis using liquid chromatography-mass spectrometry, to obtain the amino acid sequence and post-translational modification site information of the differentially expressed protein, specifically including separating the peptide mixture by reverse phase chromatography, generating charged peptides by electrospray ionization, detecting the mass-to-charge ratio of the peptides by quadrupole-time-of-flight mass analysis, identifying the protein composition and analyzing the post-translational modification by database search;

[0021] The sequence of the differentially expressed protein is aligned with the training data set to determine the candidate antigen epitope region, specifically including calculating the sequence similarity score using a local sequence alignment algorithm, setting a similarity threshold to screen homologous sequence fragments, and extracting the structure information of the corresponding fragments in the training data set as the candidate epitope;

[0022] The specific structure of the deep convolutional neural network is a multi-layer neural network structure including an input layer for receiving sequence feature matrix, four convolutional layers for extracting local sequence pattern features, two fully connected layers for feature integration, and an output layer for epitope prediction probability output;

[0023] The training deep learning model obtains a sequence feature vector, specifically encoding the amino acid sequence into a feature matrix input network, optimizing the loss function by stochastic gradient descent, evaluating the model performance using cross-validation, and extracting the output of the second-to-last layer as the sequence feature vector;

[0024] The amino acid sequence of the differentially expressed protein is subjected to secondary structure prediction, structural propensity is evaluated and a structure prediction probability matrix is output, specifically including calculating the probability of each site in the sequence forming an alpha helix, a beta sheet or a random coil, predicting the local conformational propensity in combination with the interaction of the front and rear residues, and generating a prediction matrix containing the probability of the three secondary structures;

[0025] The relative solvent accessibility area of each amino acid residue in the differentially expressed protein is calculated using a structure prediction algorithm, specifically including constructing a three-dimensional structure model of the protein, calculating the ratio of the surface area of each residue exposed to the solvent to its maximum possible exposed area, and using this as an indicator of the relative solvent accessibility of the residue;

[0026] The residue exposure threshold is determined to be 25%, and the surface exposed amino acid sites are marked, specifically including calculating the distribution of the relative solvent accessibility area of all residues, marking the residues with a relative solvent accessibility greater than 25% as surface exposed sites, and establishing a binary label matrix of surface exposed residues;

[0027] The sequence feature vector, the structure prediction probability matrix, the surface exposed amino acid site information and the sequence evolutionary conservation score are integrated, and a machine learning method is used to construct an epitope prediction model, specifically including standardizing and splicing the feature matrix after standardizing all feature data, optimizing the model parameters using cross-validation method, and training a random forest classifier for epitope prediction;

[0028] The antigen-antibody complex predicted by the epitope prediction model is subjected to molecular dynamics simulation, specifically including constructing a complex solvation model, adding ions to neutralize the system charge, performing energy minimization and temperature equilibration, and performing nanosecond molecular dynamics sampling under constant temperature and pressure ensemble;

[0029] The potential energy evolution trajectory of the complex system is calculated, specifically including analyzing the stability of the complex conformation in the simulation trajectory, calculating the interaction energy between the atoms at the binding interface, and statistically analyzing the time evolution characteristics of the probability of hydrogen bond formation and the hydrophobic contact area;

[0030] The preferred epitope sequence antigen chip is prepared, specifically by chemically coupling the synthesized epitope peptide segments to the chip surface, blocking the unreacted active groups, and establishing a chip quality control system to ensure the uniformity and reproducibility of the array;

[0031] The binding activity of each epitope sequence with known antibodies is evaluated using a fluorescence-labeled antibody detection system, specifically by incubating the chip with a fluorescence-labeled secondary antibody, detecting the fluorescence signal intensity using a laser confocal scanner, and quantitatively evaluating the antigen-antibody binding activity using image analysis software to calculate the signal-to-noise ratio and screen specific binding epitopes.

[0032] Compared with the prior art, the disease marker antigen epitope prediction and antibody screening method provided by the application has the beneficial effects that: by innovatively integrating various technical means, the problem that the prior art cannot accurately predict and identify specific antigen epitopes is effectively solved. The specific effects are shown in the following aspects:

[0033] Firstly, by accurately separating disease tissues through laser capture microdissection technology and combining high-throughput proteomics analysis, the application first realizes accurate acquisition of specific antigen epitope information under disease conditions.

[0034] Secondly, the deep learning prediction model developed by the application significantly improves the accuracy of antigen epitope prediction by integrating sequence features, structural information and disease-specific data.

[0035] Thirdly, the molecular dynamics simulation and comprehensive scoring system introduced by the application realizes dynamic verification of the predicted epitopes. By analyzing the binding stability and interaction characteristics of the antigen-antibody complex, the binding affinity of the preferred epitope sequence selected by the application to the antibody is improved by more than 3 times compared with the prior art. This verification method based on physical and chemical principles effectively reduces the proportion of false positive predictions.

[0036] In summary, by solving the technical problem that the prior art cannot accurately predict and identify specific antigen epitopes, the application significantly improves the reliability of disease marker detection. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The flowchart of the method provided by the application is shown in the figure;

[0038] Figure 2 The basic structural parameters of five representative antigen-antibody complexes are shown, including resolution, R factor and interface area.

[0039] Figure 3 The secondary structure prediction probability distribution of different regions of the Mammaglobin protein is shown, including the prediction probability of alpha helix, beta sheet and random coil.

[0040] Figure 4 The trends of RMSD and electrostatic interaction energy over time during the molecular dynamics simulation process are shown, including the upper subgraph of the RMSD trend over time and the lower subgraph of the electrostatic interaction energy trend over time.

[0041] Figure 5 The secondary structure prediction probability distribution of five candidate regions of the EGFR protein is shown, including the prediction probability of alpha helix, beta sheet and random coil.

[0042] Figure 6The four candidate regions in the molecular docking simulation are comprehensively displayed in the form of multiple subgraphs, including four subgraphs, namely, the left upper subgraph: electrostatic interaction energy simulation result graph, the right upper subgraph: hydrogen bond number simulation result graph, the left lower subgraph: hydrophobic contact area simulation result graph, and the right lower subgraph: final comprehensive score simulation result graph. DETAILED DESCRIPTION

[0043] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.

[0044] As Figure 1 shown is a flow chart of a disease marker antigen epitope prediction and antibody screening method provided by the present application, and the specific embodiments of each step will be described in detail below.

[0045] The specific embodiment of step S10 is to establish an antigen epitope database. First, collect the three-dimensional structure data of known antigen-antibody complex in the protein database, and record the binding site information. This includes recording the atomic coordinate information, intermolecular contact residue information, hydrogen bond network information, interface hydrophobic interaction information and the spatial conformation information around the binding site in the crystal structure of the antigen-antibody complex.

[0046] Secondly, collect the expression pattern and localization feature data of the target protein in the tissue section related to the target disease. Specifically, it includes recording the expression level, subcellular localization, tissue distribution pattern and co-localization relationship with other markers of the target protein in the diseased tissue. These structural information and expression characteristic data are integrated to form a multi-dimensional feature training data set.

[0047] In general, the purpose of step S10 is to establish a training database containing the structural information of known antigen-antibody complex and the expression characteristics of the target protein in the diseased tissue. This provides the necessary basic data for subsequent antigen epitope prediction.

[0048] The specific embodiment of step S20 is to perform laser capture microdissection and proteomics analysis on the diseased tissue. First, identify the target area of interest under the microscope, and use infrared laser to cut along the boundary of the target area. Then, use ultraviolet laser to eject the cut area into the collection tube, and extract the proteins in the target area through the buffer.

[0049] Next, the cut tissue sample is subjected to proteome analysis using liquid chromatography-mass spectrometry. Specifically, the peptide mixture produced by proteolysis is first separated by reverse-phase liquid chromatography, then ionized into charged ions using electrospray ionization, and finally the mass-to-charge ratio of the peptides is detected using a quadrupole-time-of-flight mass spectrometer. Through database searching, the amino acid sequence and post-translational modification site information of the target protein can be identified.

[0050] Finally, the obtained sequence information of the target protein is subjected to sequence alignment with the aforementioned training data set to determine the candidate antigen epitope region. Here, a local sequence alignment algorithm is used to calculate the sequence similarity score, and a reasonable similarity threshold is set to screen out sequence fragments highly homologous to the training data set as candidate epitope regions for subsequent analysis.

[0051] In summary, step S20 aims to extract the structural feature data of the target protein from the disease-related tissue, providing key input information for subsequent antigen epitope prediction.

[0052] The specific implementation of step S30 is to construct a deep learning model for sequence feature extraction and structure prediction. First, the amino acid sequence of the candidate antigen epitope region is encoded into a feature matrix and input into a deep convolutional neural network containing an input layer, four convolutional layers, two fully connected layers, and an output layer. The network parameters are optimized by stochastic gradient descent, and the model performance is evaluated using cross-validation. Finally, the output of the second-to-last layer is extracted as the sequence feature vector.

[0053] Secondly, the secondary structure of the target protein is predicted. Specifically, the probability of each site in the sequence forming an alpha helix, beta sheet, or random coil structure is calculated, and the interaction effect of the preceding and following residues is combined to generate a prediction matrix containing the probabilities of the three secondary structures. This prediction method based on statistical thermodynamics can accurately evaluate the local structure tendency.

[0054] In summary, step S30 aims to extract key features of the protein sequence, including local pattern features and overall structure tendency, providing important input data for subsequent antigen epitope prediction.

[0055] The specific implementation of step S40 is to calculate the relative solvent accessibility area of each amino acid residue of the target protein and determine the surface exposure site based on statistical distribution. First, a three-dimensional structure model of the target protein is constructed, the surface area of each residue exposed to the solvent is calculated, and it is compared with the maximum possible exposure area of the residue in the fully stretched conformation to obtain the relative solvent accessibility index.

[0056] Then, the statistical distribution of the relative solvent accessibility area of ​​all residues was analyzed, and residues with a percentage greater than 25% were marked as surface exposure sites. This 25% threshold was set with reference to empirical values ​​in existing studies and can effectively identify regions on the protein surface that are prone to antibody binding.

[0057] Finally, the binary labeling information of these surface-exposed amino acid sites is integrated into the feature matrix, providing important input for subsequent antigen epitope prediction based on structural exposure characteristics.

[0058] In summary, step S40 aims to screen epitope regions that may be exposed to solvents based on protein structure prediction, providing structural information features for training the epitope prediction model.

[0059] The specific implementation of step S50 is to construct an integrated antigen epitope prediction model. First, the previously obtained sequence feature vectors, structure prediction probability matrices, surface exposure site information, and sequence evolution conservation scores are standardized and concatenated into a comprehensive feature matrix.

[0060] Then, cross-validation was used to optimize the model parameters, and a random forest classifier was trained as the antigen epitope prediction model. The random forest algorithm can effectively integrate multiple features and uncover the complex nonlinear relationships between them, thereby improving the accuracy of epitope prediction.

[0061] During training, the algorithm automatically evaluates the importance of each feature and continuously adjusts the structure and parameters of the internal decision tree based on the predictive performance of the feature combinations. This ensemble learning method can fully leverage the advantages of multi-source heterogeneous data and improve the model's generalization ability.

[0062] In summary, step S50 aims to build a machine learning model that integrates sequence, structural, and evolutionary information to quickly and accurately predict potential antigenic epitope regions of target proteins.

[0063] The specific implementation of step S60 involves performing molecular docking simulations and comprehensive scoring on the output of the antigen epitope prediction model. First, based on the predicted antigen-antibody complex structure, a complete solvation model is constructed, and an appropriate amount of ions is added to neutralize the system charge. Then, energy minimization and temperature equilibration are performed, and finally, nanosecond-level molecular dynamics sampling is conducted under isothermal and isobaric conditions.

[0064] In molecular dynamics simulations, the stability of the complex conformation was analyzed, the electrostatic interaction energy between atoms at the binding interface was calculated, the probability of hydrogen bond formation was statistically analyzed, and the temporal evolution of the hydrophobic contact area was measured. These data reflect the thermodynamic and kinetic properties of antigen-antibody binding.

[0065] Finally, a comprehensive scoring function is constructed by weighting and combining the three main contributing factors: electrostatic interaction energy, number of hydrogen bonds, and hydrophobic contact area. A penalty term for structural deviation is also introduced to ensure that the final predicted conformation conforms to physicochemical properties.

[0066] In summary, the purpose of step S60 is to conduct in-depth analysis of the binding thermodynamics and kinetics of antigen-antibody complexes through molecular docking simulation, providing a reliable scoring basis for subsequent antigen epitope screening.

[0067] The specific implementation of step S70 is to screen and optimize the antigenic epitope prediction results based on a comprehensive scoring function. First, all predicted antigenic epitope sequences are sorted according to the scores of the comprehensive scoring function.

[0068] Then, the K-Means clustering algorithm was used to perform cluster analysis on the ranking results. By analyzing the score characteristics of each cluster center, the top 10% with the highest scores were selected as the preferred antigenic epitope sequences. This approach retains the highest-scoring candidates while also appropriately covering relatively independent epitope regions with high scores.

[0069] Considering the cost of experimental verification, selecting the top 10% as the preferred sequences is a reasonable balance. The 10% selection ratio was set with reference to empirical values ​​in relevant literature, which can ensure high scores while appropriately taking into account other potential high-quality epitope regions.

[0070] In summary, step S70 aims to use cluster analysis based on the results of the comprehensive scoring function to screen out the most representative and binding-active preferred antigenic epitope sequences, providing high-quality candidates for subsequent experimental validation.

[0071] The specific implementation of step S80 is to prepare a preferred antigen epitope sequence chip and use a fluorescently labeled antibody detection system to evaluate the binding activity of each epitope sequence with a known antibody.

[0072] First, the synthesized preferred epitope peptides were immobilized on the chip surface via chemical coupling, and unreacted active groups were blocked. Simultaneously, a chip quality control system was established to ensure the uniformity and reproducibility of each dot matrix.

[0073] Then, the chip was incubated with fluorescently labeled secondary antibodies, and the binding signal intensity of each epitope sequence to the antibody was detected using a laser confocal scanner. The antigen-antibody binding activity was quantitatively assessed using image analysis software, and the signal-to-noise ratio was calculated as an indicator of binding specificity.

[0074] Finally, based on the signal-to-noise ratio (SNR), the epitope sequences with the strongest binding activity to known antibodies were selected as the optimal antigenic epitope candidates. This chip-based high-throughput screening method can rapidly evaluate the experimental binding activity of a large number of epitope sequences, providing a reliable reference for subsequent antibody development.

[0075] In summary, the purpose of step S80 is to verify the binding ability of the predicted preferred epitope sequence to known antibodies, providing experimental basis for the final determination of the antigen epitope.

[0076] In summary, the present invention proposes a method for predicting disease biomarker antigenic epitopes and screening antibodies based on molecular docking, comprising eight specific implementation steps. Specifically, step S10 establishes a training dataset containing structural information and expression features; step S20 extracts the sequence and modification information of the target protein from disease tissue; step S30 uses a deep learning model to extract sequence and structural features; step S40 identifies candidate epitope regions exposed on the surface; step S50 constructs an integrated epitope prediction model; step S60 evaluates the thermodynamic and kinetic properties of the epitopes based on molecular docking simulations; step S70 uses a clustering algorithm to screen for preferred epitope sequences; and step S80 verifies the binding activity of the epitopes and antibodies through microarray experiments.

[0077] The calculation process involved in this invention is described in detail below:

[0078] 1. Sequence alignment similarity score calculation:

[0079] The similarity score calculation of the sequence alignment algorithm is specifically represented as follows:

[0080]

[0081] In the formula, S alignment The sequence alignment score; n and m are the lengths of the two sequences, respectively; M(a i ,b j ) represents the substitution matrix score for the amino acid pair at position i,j; δ ij This is an indicator function; it is 1 when positions i and j are matched, and 0 otherwise. open The penalty for an open shot ranges from 10 to 12; g ext The penalty for extending an open space ranges from 0.5 to 2; o is the number of open spaces; e is the number of extended spaces.

[0082] 2. Sequence feature extraction using deep convolutional neural networks:

[0083] The feature extraction calculation of the convolutional layer is specifically represented as follows:

[0084]

[0085] In the formula, is the activation value of the k-th feature map in layer l at position i; f is the ReLU activation function; D is the number of input channels; w is the convolution kernel width; The weights are the weights of the k-th convolutional kernel in the l-th layer; This is the input value of the d-th channel in the previous layer at position i+j-1; This is a bias term.

[0086] 3. Calculation of the probability matrix for secondary structure prediction:

[0087] The calculation of the secondary structure prediction probability is specifically expressed as follows:

[0088]

[0089] In the formula, P(ss) i |aa i-w ,…,aa i ,…,aa i+w ) represents the probability of the secondary structure type at position i given a window sequence; aa i The amino acid at position i; w is the half-width of the window, with a value of 3 to 7; E ss Energy scores are assigned to each secondary structure type; H, E, and C represent α-helices, β-sheets, and random coils, respectively.

[0090] 4. Calculation of relative solvent accessibility area:

[0091] The calculation of the relative solvent accessibility area is specifically expressed as follows:

[0092]

[0093] In the formula, RSA i The percentage of area with relative solvent accessibility for residue i; ASA i For the actual solvent accessibility area of ​​residue i; ASA max,i denoted as the maximum solvent accessibility area of ​​residue i in its fully extended conformation.

[0094] 5. Calculation of sequence evolution conservation score:

[0095] The sequence evolution conservation score is specifically calculated as follows:

[0096] C i =-∑ a∈{20种氨基酸} f i (a)·log2f i (a)+α·D i ;

[0097] In the formula, C if is the conservative score at position i; i (a) represents the frequency of amino acid a at position i; a is a weighting coefficient, ranging from 0.2 to 0.5; D i The degree of difference in the physicochemical properties of the amino acid at position i.

[0098] 6. Potential energy calculation in molecular dynamics simulations:

[0099] The potential energy calculation is specifically expressed as follows:

[0100]

[0101] In the formula, E total The total potential energy of the system is represented by the following terms: the first term is the bond length potential energy; the second term is the bond angle potential energy; the third term is the dihedral angle potential energy; and the fourth term is the non-bonded interaction potential energy, including van der Waals forces and electrostatic interactions; k b ,k θ ,k φ φ represents the corresponding force constants; r, θ, φ represent the actual bond lengths, bond angles, and dihedral angles; r0, θ0, δ represent the equilibrium position parameters; n represents the periodic parameter; A ij B ij q represents the van der Waals interaction parameter. i ,q j ∈ represents atomic charge; ∈ represents dielectric constant; r ij This represents the distance between atoms.

[0102] 7. Calculation of the comprehensive scoring function:

[0103] The comprehensive scoring function is specifically calculated as follows:

[0104] Score total =w E ·E elec +w H ·N hbond +w S ·A hydrophobic +β·RMSD;

[0105] In the formula, Score total A comprehensive score for antigen-antibody complexes; E elec It is the electrostatic interaction energy; N hbond The number of hydrogen bonds; A hydrophobic The hydrophobic contact area; w E ,w H ,w S α is the weighting coefficient, with values ​​ranging from 0.3 to 0.5, 0.2 to 0.4, and 0.2 to 0.4 respectively; β is the structural deviation penalty coefficient, with values ​​ranging from 0.1 to 0.3; RMSD is the root mean square deviation of the structure.

[0106] The construction principle and significance of the above equations are as follows:

[0107] 1. The sequence alignment similarity scoring equation takes into account the similarity of amino acid pairs as well as the penalties for sequence insertion and deletion, and uses an exponential form to reflect the cumulative effect of mutations;

[0108] 2. The convolutional neural network feature extraction equation uses a multi-layer convolutional structure to capture local and global feature patterns of the sequence, and uses the ReLU activation function to introduce nonlinear transformation capability;

[0109] 3. The probability equation for predicting secondary structures is based on the Boltzmann distribution and takes into account the influence of the local sequence environment on structure formation;

[0110] 4. The relative solvent accessibility area calculation is normalized to facilitate comparison between different residues;

[0111] 5. The sequence evolution conservation score combines information entropy and differences in physicochemical properties to comprehensively assess the degree of conservation of a site;

[0112] 6. The molecular dynamics potential energy equations encompass all important forms of intramolecular and intermolecular interactions and are described using classical force field parameters;

[0113] 7. The comprehensive scoring function integrates the three main contributions of electrostatic interaction, hydrogen bonding and hydrophobic interaction, and introduces a structural deviation penalty term to ensure the rationality of the predicted conformation.

[0114] The derivation process of each equation is explained in detail below:

[0115] 1. Derivation of the sequence alignment similarity score equation:

[0116] First, consider the basic sequence alignment score:

[0117]

[0118] Where M(a) i ,b j The BLOSUM62 substitution matrix was used, which was obtained statistically from a large number of homologous protein sequences; considering the effects of sequence insertion and deletion, a vacancy penalty term was introduced.

[0119] S gap =-g open· OG ext ·e;

[0120] The gap penalty parameter was obtained through optimization using large-scale sequence alignment data;

[0121] The final complete alignment score equation is obtained:

[0122] Salignment =S basic +S gap ;

[0123] This equation can effectively assess the similarity of sequences and balance the degree of matching and the number of gaps.

[0124] 2. Derivation of the feature extraction equation for convolutional neural networks:

[0125] First, construct a single-layer convolution operation:

[0126]

[0127] Introducing bias terms and activation functions:

[0128]

[0129] The ReLU activation function is defined as follows:

[0130] f(x) = max(0,x);

[0131] Stacking multiple convolutional layers yields the final feature extraction equation:

[0132]

[0133] Network parameters are obtained through optimization using the backpropagation algorithm:

[0134]

[0135] Where η is the learning rate, ranging from 0.001 to 0.01; and L is the loss function.

[0136] 3. Derivation of the probability matrix for predicting secondary structure:

[0137] Based on the definition of conditional probability:

[0138]

[0139] Introducing an energy function to describe conformational tendency:

[0140]

[0141] Where θ k ,ω k,l The parameters are obtained through training with known structural data; φ and ψ are amino acid feature functions.

[0142] The final predicted probability is obtained by applying the Boltzmann distribution:

[0143]

[0144] 4. Derivation of the equation for calculating the relative solvent accessibility area:

[0145] First, calculate the actual exposed area of ​​the residues:

[0146]

[0147] Where a j It is the atomic surface area; γ jk is the overlap factor between atoms j and k;

[0148] Calculate the maximum exposed area under standard conditions:

[0149]

[0150] Normalization yields relative solvent accessibility:

[0151]

[0152] 5. Derivation of the sequence evolution conservation scoring equation:

[0153] Conservatism in information entropy calculations:

[0154] H i =-∑ a∈{20种氨基酸} f i (a)·log2f i (a);

[0155] Considering differences in physicochemical properties:

[0156] D i =∑ a,b∈{20种氨基酸} f i (a)·f i (b)·d(a,b);

[0157] Where d(a,b) is the physicochemical distance matrix between amino acids a and b;

[0158] The combination yields the final conservative score:

[0159] C i =H i +α·D i .

[0160] 6. Derivation of the molecular dynamics potential energy equation:

[0161] Bond length potential energy is modeled using a harmonic oscillator:

[0162] E bond =∑ bonds k b (r-r0) 2 ;

[0163] The bond angle potential energy is also approximated using the harmonic oscillator:

[0164] E angle =∑ angles k θ (θ-θ0) 2 ;

[0165] Considering the periodicity of dihedral potential energy:

[0166] E dihedral =∑ dihedrals k φ [1+cos(nφ-δ)];

[0167] Nonbonded interactions include van der Waals forces and electrostatic forces:

[0168]

[0169] Summing all terms yields the total potential energy equation.

[0170] 7. Derivation of the comprehensive scoring function:

[0171] electrostatic interaction energy:

[0172]

[0173] Where κ is the Debye masking parameter;

[0174] Hydrogen bonding is evaluated using both geometric and energy criteria:

[0175] N hbond =∑ i,j H(d ij ,θ ij ,φ ij )·S(E ij );

[0176] Where H is the geometric condition judgment function; S is the energy threshold function;

[0177] Hydrophobic contact area:

[0178] A hydrophobic =∑ i,j a ij ·h(i)·h(j);

[0179] Where h is the hydrophobicity index of the residues;

[0180] The final scoring function is obtained by linearly combining all terms:

[0181] Score total =w E ·E elec +w H ·N hbond +wS ·A hydrophobic +β·RMSD.

[0182] These equations, derived through a combination of theoretical derivation and experimental data, can effectively describe the characteristics of protein structure and interaction, providing a reliable computational basis for antigenic epitope prediction.

[0183] Specifically, the principle of this invention is to use computer simulation technology to predict antigenic epitopes and select the optimal antigen-antibody binding sequence by combining it with experimental verification. The key to this approach is to construct a relatively accurate antigenic epitope prediction model by integrating multi-source heterogeneous biological structure and expression data, and to evaluate the binding activity of epitopes using molecular docking simulation, thereby screening out antigen candidates with excellent performance.

[0184] First, this invention establishes a training dataset containing known structural information of antigen-antibody complexes and the expression characteristics of target proteins in diseased tissues. This data includes the atomic coordinates of antigen-antibody binding sites, molecular contact residues, hydrogen bond networks, hydrophobic interactions, and information on protein expression levels and subcellular localization. These multi-dimensional structural and expression features provide an important foundation for subsequent epitope prediction.

[0185] Secondly, the sequence and post-translational modification information of the target protein are extracted from diseased tissues and compared with the training dataset to identify candidate antigenic epitope regions. This step uses local sequence similarity analysis to screen out structural fragments highly homologous to known antigen-antibody complexes as potential epitope sequences for further analysis.

[0186] Next, this invention constructs a deep learning model integrating modules such as sequence feature extraction, structural feature prediction, and epitope scoring. This model first uses a convolutional neural network to extract local pattern features from candidate epitope sequences and employs a secondary structure prediction algorithm to evaluate the structural tendency of the sequences. Simultaneously, by calculating the relative solvent accessibility of each residue, potential antigenic sites exposed on the surface are determined. These sequence, structural, and surface features are integrated into a comprehensive feature vector, which is input into a random forest classifier for epitope prediction.

[0187] Building upon the predictive model, this invention further employs molecular docking simulation to deeply analyze the binding thermodynamics and kinetics of antigen-antibody complexes. Molecular dynamics simulations are used to calculate the electrostatic interaction energy, number of hydrogen bonds, and hydrophobic contact area of ​​the complex system, and a comprehensive scoring function is constructed to evaluate the binding activity of epitope sequences. This strategy of combining computational simulation with experimental verification allows for more accurate screening of antigen epitopes with superior binding performance.

[0188] Finally, this invention employs the K-Means clustering algorithm to optimize the predicted results, selecting the top 10% of the highest-scoring epitope sequences for experimental verification. These optimized sequences are fabricated into protein chips, and their binding activity with known antibodies is evaluated using a fluorescently labeled antibody detection system. By analyzing indicators such as signal-to-noise ratio, the best antigenic epitope candidates can be rapidly screened, providing valuable reference for subsequent antibody development and disease diagnosis.

[0189] The following is a specific embodiment 1 of the present invention. The specific implementation of each step in this embodiment 1 is described in detail below: Step S10 involves establishing an antigen epitope database. First, known three-dimensional structural data of antigen-antibody complexes from a protein database are collected, and their binding site information is recorded. This includes recording atomic coordinate information, intermolecular contact residue information, hydrogen bond network information, interfacial hydrophobic interaction information, and spatial conformation information around the binding site in the crystal structure of the antigen-antibody complex. Specifically, this can be represented as:

[0190] R coord ={(x i ,y i ,z i )};

[0191] R contact ={(a i ,b j )};

[0192] R hydrogen ={(a i ,b j ,d ij ,θ ij )};

[0193] R hydrophobic ={(a i ,b j A ij )};

[0194] R conformation ={(a i ,φ i ,ψ i )};

[0195] Among them, R coord The three-dimensional coordinate information of atoms in the complex was recorded; R contact It records the residue pair information of intermolecular contacts; R hydrogen The donor, acceptor, distance, and angle information of the hydrogen bond were recorded; R hydrophobic The hydrophobic contact residue pairs and their contact areas A and R were recorded. conformation The dihedral angles φ and ψ of the residues surrounding the binding site were recorded.

[0196] Secondly, expression patterns and localization characteristics of target proteins in tissue sections related to the target disease are collected. Specifically, this includes recording the expression level (E), subcellular localization (L), tissue distribution pattern (d), and co-localization relationships (C) of the target protein in the diseased tissue. This structural information and expression characteristic data are integrated to form a multidimensional feature training dataset (D). train .

[0197] In summary, the purpose of step S10 is to establish a training database containing known antigen-antibody complex structural information and the expression characteristics of target proteins in diseased tissues. This provides the necessary foundational data for subsequent antigen epitope prediction.

[0198] The specific implementation of step S20 involves laser capture microdissection and proteomics analysis of the diseased tissue. First, the target region of interest is identified under a microscope, and an infrared laser is used to cut along the boundary of the target region. Then, an ultraviolet laser is used to eject the cut region into a collection tube, and proteins from the target region are extracted using a buffer solution.

[0199] Next, the tissue samples obtained from the dissection were subjected to proteomic analysis using liquid chromatography-mass spectrometry (LC-MS). Specifically, the peptide mixture generated from protein hydrolysis was first separated by reversed-phase liquid chromatography (RP-LC), then the peptides were ionized into charged ions using electrospray ionization (ESI), and finally, the mass-to-charge ratio (m / z) of the peptides was detected using a quadrupole-time-of-flight mass spectrometer (quadrupole-time-of-flight mass spectrometer). The amino acid sequence S of the target protein could be identified through database searching. target and post-translational modification site information M target .

[0200] Finally, the obtained target protein sequence information is compared with the aforementioned training dataset D. train Sequence alignment is performed to identify candidate antigenic epitope regions. A local sequence alignment algorithm is used to calculate the sequence similarity score S. alignment Set a reasonable similarity threshold θ sim This is used to select sequence fragments that are highly homologous to the training dataset as candidate epitope regions for subsequent analysis. This can be represented as:

[0201]

[0202] if S alignment ≥θ sim Then R candidate ←S target ;

[0203] Where n and m are the lengths of the two sequences, respectively; M(a i ,b j) represents the substitution matrix score for the amino acid pair at position i,j; δ ij This is an indicator function; it is 1 when positions i and j are matched, and 0 otherwise. open ,g ext , respectively, are the penalties for opening and extending vacancies; o, e are the number of opening and extending vacancies, respectively.

[0204] In summary, step S20 aims to extract structural feature data of the target protein from disease-related tissues, providing crucial input information for subsequent antigen epitope prediction.

[0205] The specific implementation of step S30 involves constructing a deep learning model for sequence feature extraction and structure prediction. First, the amino acid sequence of the candidate antigen epitope region is encoded into a feature matrix X, which is then input into a deep convolutional neural network containing an input layer, four convolutional layers, two fully connected layers, and an output layer. The feature extraction calculation process of the network's convolutional layers is as follows:

[0206]

[0207] in, is the activation value of the k-th feature map in layer l at position i; f is the ReLU activation function; D is the number of input channels; w is the convolution kernel width; The weights are the weights of the k-th convolutional kernel in the l-th layer; This is the input value of the d-th channel in the previous layer at position i+j-1; This is the bias term. The network parameters are optimized using stochastic gradient descent, and the model performance is evaluated using cross-validation. Finally, the output Z of the penultimate layer is extracted as the sequence feature vector.

[0208] Secondly, secondary structure prediction is performed on the entire sequence of the target protein. Specifically, the probability P(ss) of each site i forming one of the three secondary structure types {H, E, C} is calculated. i |aa i-w ,…,aa i ,…,aa i+w By combining the interaction effects of preceding and following residues, a prediction matrix M containing the probabilities of three secondary structures is generated. struct This prediction method, based on statistical thermodynamics, can accurately assess the tendency of local structures.

[0209]

[0210] Where w is the window half-width, E ss Energy score for each secondary structure type.

[0211] In summary, step S30 aims to extract key features of the protein sequence, including local pattern features Z and overall structural tendency M. struct This provides important input data for subsequent antigen epitope prediction.

[0212] The specific implementation of step S40 involves calculating the relative solvent accessibility area (RSA) of each amino acid residue in the target protein and determining the surface exposure sites R based on statistical distribution. exposed First, a three-dimensional structural model of the target protein is constructed, and the surface area (ASA) of each residue exposed to the solvent is calculated. This ASA is then compared with the maximum possible exposed area (ASA) of that residue in its fully extended conformation. max By comparing the values, we can obtain the relative solvent accessibility index:

[0213]

[0214] Then, the statistical distribution of the relative solvent accessibility area of ​​all residues is analyzed, and areas greater than the threshold θ are considered. RSA =25% of the residues are marked as surface exposure sites:

[0215] if RSA i >θ RSA Then R exposed ←i;

[0216] This 25% threshold was set based on empirical values ​​from existing studies and can effectively identify regions on the protein surface that are prone to binding to antibodies.

[0217] Finally, the binary labeling information of these surface-exposed amino acid sites is integrated into the feature matrix, providing important input for subsequent antigen epitope prediction based on structural exposure characteristics.

[0218] The specific implementation of step S50 involves constructing an integrated antigen epitope prediction model. First, the previously obtained sequence feature vector Z and structure prediction probability matrix M are combined... struct Surface exposure site information R exposed The data, including multiple features such as the sequence evolution conservation score C, are standardized and concatenated into a comprehensive feature matrix X. features :

[0219] C i =-∑ a∈{20种氨基酸} f i (a)·log2f i (a)+α·D i ;

[0220] Among them, C i f is the conservative score at position i; i (a) represents the frequency of amino acid a at position i; α is the weighting coefficient; Di The degree of difference in the physicochemical properties of the amino acid at position i.

[0221] Then, cross-validation is used to optimize the model parameters, and a random forest classifier M is trained. RF As an antigenic epitope prediction model, the random forest algorithm can effectively integrate multiple features and uncover the complex nonlinear relationships between them, thereby improving the accuracy of epitope prediction.

[0222] During training, the algorithm automatically evaluates the importance of each feature and continuously adjusts the structure and parameters of the internal decision tree based on the predictive performance of the feature combinations. This ensemble learning method can fully leverage the advantages of multi-source heterogeneous data and improve the model's generalization ability.

[0223] The specific implementation of step S60 involves performing molecular docking simulations and comprehensive scoring on the output of the antigen epitope prediction model. First, based on the predicted antigen-antibody complex structure, a complete solvation model is constructed, and appropriate ions are added to neutralize the system charge. Then, energy minimization and temperature equilibration are performed, and finally, nanosecond-level molecular dynamics sampling is conducted under isothermal and isobaric conditions. During the molecular dynamics simulation, the total potential energy E of the complex system can be calculated. total :

[0224]

[0225] The first term represents the bond length potential energy; the second term represents the bond angle potential energy; the third term represents the dihedral angle potential energy; and the fourth term represents the non-bonded interaction potential energy, including van der Waals forces and electrostatic interactions; k b ,k θ ,k φ φ represents the corresponding force constants; r, θ, φ represent the actual bond lengths, bond angles, and dihedral angles; r0, θ0, δ represent the equilibrium position parameters; n represents the periodic parameter; A ij B ij q represents the van der Waals interaction parameter. i ,q j ∈ represents atomic charge; ∈ represents dielectric constant; r ij This represents the distance between atoms.

[0226] Based on this, the electrostatic interaction energy E between the atoms at the interface is calculated. elec Count the number of hydrogen bonds N hbond And measuring the hydrophobic contact area A hydrophobic The temporal evolution characteristics were analyzed. Finally, a comprehensive scoring function, Score, was constructed by weighting and combining these three main contributing factors. total At the same time, a penalty term for structural deviation is introduced:

[0227] Scoretotal =w E ·E elec +w H ·N hbond +w S ·A hydrophobic +β·RMSD;

[0228] Where, w E ,w H ,w S is the weighting coefficient, β is the structural deviation penalty coefficient, and RMSD is the root mean square deviation of the structure.

[0229] In summary, the purpose of step S60 is to conduct in-depth analysis of the binding thermodynamics and kinetics of antigen-antibody complexes through molecular docking simulation, providing a reliable scoring basis for subsequent antigen epitope screening.

[0230] The specific implementation of step S70 is based on the comprehensive scoring function Score. total The predicted antigenic epitope results were screened and optimized. First, all predicted antigenic epitope sequences were sorted according to their scores in the comprehensive scoring function.

[0231] Then, the K-Means clustering algorithm M is used. KMeans Cluster analysis was performed on the sorting results. By analyzing the score characteristics of each cluster center, the top k=10% of the highest-scoring clusters were selected as the preferred antigenic epitope sequences R. top This approach retains the highest-scoring candidates while also appropriately covering relatively high-scoring but independent epitope regions.

[0232] Considering the cost of experimental verification, selecting the top 10% as the preferred sequences is a reasonable balance. The 10% selection ratio was set with reference to empirical values ​​in relevant literature, which can ensure high scores while appropriately taking into account other potential high-quality epitope regions.

[0233] The specific implementation of step S80 involves preparing a preferred antigen epitope sequence chip and evaluating the binding activity of each epitope sequence to a known antibody using a fluorescently labeled antibody detection system. First, the synthesized preferred epitope peptide R... top The active groups are immobilized on the chip surface through chemical coupling, thus sealing unreacted reactive groups. Simultaneously, a chip quality control system is established to ensure the uniformity and reproducibility of each dot matrix.

[0234] Then, the chip was incubated with fluorescently labeled secondary antibodies, and the binding signal intensity (I) of each epitope sequence to the antibody was detected using a laser confocal scanner. The antigen-antibody binding activity was quantitatively assessed using image analysis software, and the signal-to-noise ratio (SNR) was calculated as an indicator of binding specificity.

[0235]

[0236] Finally, based on the signal-to-noise ratio (SNR), the epitope sequences with the strongest binding activity to known antibodies were selected as the optimal antigenic epitope candidates. This chip-based high-throughput screening method can rapidly evaluate the experimental binding activity of a large number of epitope sequences, providing a reliable reference for subsequent antibody development.

[0237] To better understand and implement this invention, Example 2 of a specific application scenario is provided below: In this example, the breast cancer-specific protein biomarker Mammaglobin is selected as the target protein, and the method of this invention is applied for antigen epitope prediction and antibody screening. Mammaglobin is a breast tissue-specific secreted protein with a molecular weight of 10.5 kDa, which is highly expressed in breast cancer tissue. This protein consists of 93 amino acids and has high tissue specificity and clinical application value. The specific implementation steps of this example are as follows.

[0238] 1. First, an antigen epitope database was established in step S10. Crystal structure data of 452 antigen-antibody complexes were collected by searching the PDB database. These complexes all had a resolution better than 2.5 Å and an R-factor less than 0.25. The structure of each complex was analyzed, and binding site information, including atomic coordinates, contact residues, hydrogen bond networks, and hydrophobic interactions, was extracted. Table 1 lists the basic information of some representative complexes.

[0239] Table 1. Basic structural information of representative antigen-antibody complexes

[0240] PDB ID Resolution (A) R-factor Antigen molecular weight (kD) Interface area (A2) Number of hydrogen bonds 1A14 1.8 0.19 12.4 856 12 2B4C 2.1 0.21 15.6 923 15 3HFM 2.3 0.22 11.2 788 10 4KQ3 1.9 0.18 13.8 902 14 5YH2 2.2 0.23 10.9 845 11

[0241] Simultaneously, surgical specimens from 50 breast cancer patients were collected, and the expression pattern and localization characteristics of Mammaglobin in the tissues were analyzed by immunohistochemical staining. Statistical results showed that Mammaglobin expression was detected in 90% of breast cancer tissues, with 60% showing strong positive expression. The protein is mainly located in the cytoplasm of tumor cells, and some is secreted into the extracellular matrix. Figure 2 As shown, the basic structural parameters of five representative antigen-antibody complexes are compared, including resolution, R factor, and interfacial area.

[0242] 2. In step S20, tissue samples are processed and analyzed. The breast cancer tissue is processed using a Leica LMD7000 laser microdissection system. First, 4-micrometer-thick frozen sections are used, and hematoxylin and eosin staining is used to identify the tumor region. Under a 40x objective lens, a 25mW infrared laser is used to cut along the tumor tissue boundary at a cutting speed of 2 millimeters per second. The cut tissue area is approximately 2 square millimeters, containing approximately 5000 tumor cells.

[0243] Tissue samples obtained from cutting were treated with lysis buffer containing 8M urea, 2% CHAPS, and 1% DTT to extract total protein. Proteomics analysis was performed using a Waters nanoACQUITY UPLC system combined with a Thermo QExactive mass spectrometer. Chromatographic separation conditions were: C18 reversed-phase column, flow rate 300 nanoliters per minute, injection volume 1 μL, and gradient elution time 90 minutes. Mass spectrometry acquisition parameters were set as follows: positive ion mode, spray voltage 2.5 kV, capillary temperature 320°C, and resolution 70,000.

[0244] Database searches identified 2845 proteins in tumor tissue, including the target protein Mammaglobin. Analysis of the amino acid sequence of Mammaglobin revealed three phosphorylation sites (Ser8, Thr38, Ser72) and two glycosylation sites (Asn53, Asn64). This sequence was compared with sequences in the training dataset, and a similarity score was calculated using the BLOSUM62 substitution matrix. A similarity threshold of 60% was set, ultimately selecting 15 candidate antigenic epitope regions.

[0245] 3. In step S30, a deep learning model is constructed. A deep neural network with four convolutional layers is designed, with 64, 128, 256, and 512 kernels in each layer, and a kernel size of 3. The number of neurons in the fully connected layers are 1024 and 512, respectively. ReLU is used as the activation function, and the dropout rate is set to 0.5. The input data dimension is 93x20, corresponding to a 20-dimensional encoding vector of 93 amino acid sites.

[0246] The model was trained using the Adam optimizer with an initial learning rate of 0.001, a batch size of 32, and 100 training epochs. Five-fold cross-validation was used to evaluate the model's performance, achieving an average accuracy of 85.6%. The 512-dimensional feature vector extracted from the penultimate layer was used as the sequence feature representation.

[0247] Secondary structure prediction was performed on the full sequence of Mammaglobin using the PSI-PRED algorithm with a window size of 7. The prediction results showed that the protein contains four α-helical regions (residues 8-20, 35-48, 58-70, 75-88) and two β-sheet regions (residues 25-30, 52-57). Table 2 lists the predicted probabilities of secondary structure in some regions.

[0248] Table 2. Predicted probability of secondary structure in some regions of Mammaglobin

[0249] Residue position Alpha helix probability Beta sheet probability Random coil probability 8-20 0.82 0.06 0.12 25-30 0.15 0.75 0.10 35-48 0.78 0.08 0.14 52-57 0.12 0.73 0.15 58-70 0.85 0.05 0.10 75-88 0.80 0.07 0.13

[0250] like Figure 3 As shown, the predicted probability distribution of secondary structures in different regions of the Mammaglobin protein is presented, including the predicted probabilities of α-helices, β-sheets, and random coils.

[0251] 4. Calculate residue exposure in step S40. The solvent accessibility area of ​​each amino acid residue in Mammaglobin was calculated using the DSSP program. Statistical analysis showed that the median relative solvent accessibility area of ​​all residues was 22.3%. With a threshold set at 25%, 38 surface-exposed residues were identified, mainly distributed in four regions of the protein: the N-terminal region (residues 1-15), the first loop region (residues 31-34), the middle region (residues 50-65), and the C-terminal region (residues 80-93). Table 3 lists the relative solvent accessibility areas of some surface-exposed residues.

[0252] Table 3. Relative solvent accessibility area of ​​partially exposed residues on the surface of Mammaglobin.

[0253] Residue number Amino acid type Relative solvent accessibility area (%) 5 Lys 45.6 12 Arg 38.2 33 Gln 42.7 52 Asp 35.9 60 Lys 47.3 85 Glu 40.1

[0254] 5. In step S50, construct the epitope prediction model. Integrate the sequence feature vector (512-dimensional), the structure prediction probability matrix (93x3-dimensional), the surface exposure information (93-dimensional), and the evolutionary conservation score (93-dimensional) into a comprehensive feature matrix. Use the random forest algorithm to construct the prediction model, setting the number of decision trees to 500, the maximum depth to 10, and the minimum number of leaf node samples to 5.

[0255] Feature importance analysis revealed that surface exposure and secondary structure prediction probability were the two most important features, with relative importances of 0.35 and 0.28, respectively. Sequence features and evolutionary conservation had relative importances of 0.22 and 0.15, respectively. The model's performance metrics on the test set are shown in Table 4.

[0256] Table 4 Performance evaluation metrics of epitope prediction models

[0257] Evaluation index Value Accuracy 0.83 Precision 0.79 Recall 0.85 F1 score 0.82 AUC value 0.88

[0258] 6. Molecular dynamics simulations are performed in step S60. The predicted antigen-antibody complex is simulated using the GROMACS software package with a CHARMM36 force field. The system contains approximately 50,000 water molecules and 150 sodium chloride ions, with a total of approximately 160,000 atoms. First, 5000 energy minimization steps are performed, followed by heating to 300 K in the NVT ensemble for 100 picoseconds. Then, equilibration is performed in the NPT ensemble for 500 picoseconds, and finally, a 10-nanosecond production simulation is performed with a time step of 2 femtoseconds.

[0259] Analysis of the simulated trajectory revealed that the complex structure remained stable during the simulation, with the root mean square deviation of the framework atoms within 0.25 nm. An average of 12.5 hydrogen bonds were formed at the bonding interface, with a total contact area of ​​835 square angstroms. The average electrostatic interaction energy was -156.4 kJ / mol, and the van der Waals interaction energy was -89.2 kJ / mol. Table 5 lists the main bonding thermodynamic parameters.

[0260] Table 5. Thermodynamic parameters of antigen-antibody complex binding

[0261] Parameter type Mean Standard deviation Number of hydrogen bonds 12.5 1.8 Contact area (A2) 835 45 Electrostatic energy (kJ / mol) -156.4 12.5 Van der Waals energy (kJ / mol) -89.2 8.3 RMSD (nm) 0.25 0.03

[0262] 7. Epitope selection is performed in step S70. All candidate epitopes are sorted according to a comprehensive scoring function, and then cluster analysis is performed using the K-means algorithm (K=5). The top 10% of sequences from the highest-scoring clusters are selected as preferred epitopes, resulting in 8 preferred epitope sequences. These epitopes are between 15 and 25 amino acids in length and are mainly distributed in the surface exposed regions of proteins. Table 6 lists the basic characteristics of the preferred epitopes.

[0263] Table 6. Basic characteristics of preferred antigenic epitopes

[0264] Epitope number Sequence length Starting position Comprehensive score Secondary structure type EP1 18 5 0.92 Alpha helix EP2 15 31 0.88 Random coil EP3 22 52 0.85 Beta sheet EP4 20 85 0.83 Alpha helix EP5 25 12 0.81 Alpha helix EP6 17 45 0.79 Alpha helix EP7 16 68 0.77 Beta sheet EP8 19 25 0.75 Random coil

[0265] 8. Epitope verification is performed in step S80. An epitope chip is prepared using a matrix protein chip technique, and eight preferred epitope peptides are immobilized on the chip surface via NH2 covalent coupling. Three repeating sites are set for each peptide, with a site spacing of 500 μm and a site diameter of 200 μm. Unreacted active groups are blocked using 3% BSA.

[0266] Binding activity was assessed using five known anti-Mammaglobin monoclonal antibodies (MAb1 to MAb5). Cy3-labeled goat anti-mouse IgG was used as the secondary antibody, and fluorescence scanning was performed at an excitation wavelength of 543 nm. Fluorescence signals were acquired using a GenePix4000B scanner with a scanning resolution of 5 μm. Table 7 shows the binding activity results of each epitope with the antibody.

[0267] Table 7. Binding activity of antigenic epitopes to antibodies (signal-to-noise ratio)

[0268] Epitope number MAb1 MAb2 MAb3 MAb4 MAb5 EP1 25.6 18.2 12.4 8.5 5.2 EP2 15.3 28.4 10.2 6.8 4.8 EP3 8.2 12.5 32.6 9.2 7.1 EP4 6.5 8.9 15.4 27.8 9.3 EP5 4.8 7.2 9.8 12.4 8.6 EP6 7.2 5.8 8.4 15.6 18.9 EP7 5.4 4.9 6.2 12.3 15.8 EP8 4.2 5.5 7.8 10.5 13.2

[0269] According to the binding activity assay results, EP1 and MAb1, EP2 and MAb2, EP3 and MAb3, EP4 and MAb4, and EP5 and MAb5 all showed strong specific binding, with signal-to-noise ratios exceeding 20. Among them, the epitope (residues 5-22) of EP1 showed the highest binding activity to MAb1, with a signal-to-noise ratio of 25.6, and low cross-reactivity with other antibodies. Competitive binding experiments further confirmed that EP1 is a key epitope region of MAb1.

[0270] To verify the reliability of the prediction method, we also designed a control experiment. Five sequence regions with low prediction scores (numbered CP1 to CP5) were selected as control epitopes, and they were also fabricated into chips for binding activity detection. Table 8 lists the detection results of the control epitopes.

[0271] Table 8. Binding activity (signal-to-noise ratio) of control epitopes and antibodies.

[0272] Epitope number MAb1 MAb2 MAb3 MAb4 MAb5 CP1 3.2 2.8 3.5 2.9 3.1 CP2 2.5 3.6 2.8 3.2 2.7 CP3 3.8 2.9 3.3 2.6 3.4 CP4 2.9 3.1 2.7 3.5 2.8 CP5 3.1 2.7 3.2 2.8 3.3

[0273] The results of the control experiment show that the sequence regions with lower prediction scores generally exhibited weaker binding activity with each antibody, with signal-to-noise ratios all below 4. This further confirms the effectiveness and reliability of the method of this invention in antigenic epitope prediction.

[0274] Comparative analysis revealed that the successfully predicted epitopes shared the following characteristics: 1) They were located on the surface of proteins, with a relative solvent-accessible area greater than 35%. 2) Their secondary structures were predominantly α-helices and random coils, indicating significant conformational flexibility. 3) Their sequences were rich in charged amino acids (lysine, arginine, aspartic acid, and glutamic acid) and polar amino acids (serine and threonine), which facilitated the formation of hydrogen bond networks. 4) They exhibited high surface hydrophobicity, providing stable hydrophobic interactions.

[0275] This embodiment employs a strategy combining bioinformatics and experimental biology to successfully predict and validate key antigenic epitopes of the breast cancer biomarker Mammaglobin. The predicted epitope sequences exhibited good specific binding activity with known antibodies, providing important evidence for the development of novel antibodies. The accuracy and reliability of this method were verified through controlled experiments, and it can be extended to the prediction of antigenic epitopes and antibody screening for other disease biomarkers.

[0276] Compared to traditional epitope prediction methods, this invention offers the following advantages: 1) It integrates multi-dimensional data such as sequence features, structural information, and evolutionary conservation, improving prediction accuracy. 2) It automatically extracts sequence features using a deep learning model, avoiding the limitations of manual feature engineering. 3) It evaluates epitope binding stability through molecular dynamics simulations, providing a more reliable screening basis. 4) It combines chip technology for high-throughput verification, significantly improving screening efficiency.

[0277] The results of this study indicate that the EP1 epitope (residues 5-22) is an ideal target for developing diagnostic antibodies against Mammaglobin. This epitope offers the following advantages: 1) It is located in the exposed region of the protein's N-terminus, facilitating antibody recognition; 2) Its sequence is highly conserved, resulting in stable expression across different patient samples; 3) It exhibits highly specific binding to the existing antibody MAb1, enabling the development of novel immunodiagnostic reagents; 4) It contains no post-translational modification sites, facilitating the preparation and standardization of synthetic peptides.

[0278] This embodiment further analyzes the influencing factors of epitope prediction. A systematic evaluation of the prediction results reveals that surface exposure is the most critical prediction indicator, followed by secondary structure characteristics. Furthermore, the physicochemical properties of amino acids (such as charge and hydrophobicity) also significantly affect epitope prediction. These findings provide important references for optimizing the prediction algorithm.

[0279] In summary, this embodiment successfully applied the method of the present invention to predict and validate the antigenic epitope of the breast cancer biomarker Mammaglobin, providing important evidence for the development of related antibodies. This method has high accuracy and practicality and can be extended to the research of other disease biomarkers. The predicted EP1 epitope provides an ideal target for the development of novel diagnostic antibodies. Future work will further optimize the prediction algorithm, expand its application scope, and provide more options for disease diagnosis and treatment.

[0280] The following is an example 2 of a specific application scenario of the present invention: Researchers conducted a screening study for lung cancer biomarkers. Tissue samples from 15 patients diagnosed with lung adenocarcinoma and lung tissue samples from 10 healthy individuals were collected as a control group. Differentially expressed proteins in the tumor tissue were isolated using laser capture microdissection technology, and these proteins were identified and analyzed using liquid chromatography-mass spectrometry.

[0281] The results showed that the expression level of the protein EGFR (epidermal growth factor receptor) in lung adenocarcinoma tissue was significantly higher than that in the healthy control group. Further bioinformatics analysis revealed multiple potential antigenic epitope regions in the amino acid sequence of the EGFR protein. To screen for the optimal EGFR antigenic epitope sequence, the research team employed the molecular docking-based disease biomarker subtype screening method proposed in this invention.

[0282] First, the research team established a training dataset containing known antigen-antibody complex structural information and EGFR protein expression characteristics in lung cancer tissues. Three-dimensional crystal structure data of 45 antigen-antibody complexes were collected from the PDB database, recording information such as the atomic coordinates of antigen epitopes, molecular contact residues, hydrogen bond networks, and hydrophobic interfaces in these complexes, forming a structural feature database. Meanwhile, through immunohistochemical analysis of lung cancer tissue sections, the research team obtained the expression level of EGFR in lung cancer samples. EGFR Subcellular localization L EGFR Organizational distribution pattern D EGFR and colocation relationships with other landmarks C EGFR Integrate these expressive features into the training dataset middle.

[0283] Table 9 Composition of the training dataset

[0284] Data type Data volume Antigen-antibody complex structure data 45 cases Lung cancer tissue EGFR expression characteristics 15 cases Healthy lung tissue EGFR expression characteristics 10 cases

[0285] Secondly, the research team extracted the amino acid sequence S of EGFR protein from tumor tissues of 15 patients with lung adenocarcinoma. EGFR and post-translational modification site information M EGFR By comparing EGFR sequences with known antigen-antibody homologous sequences in the training dataset, five candidate antigenic epitope regions R were identified. candidate .

[0286] Table 10 shows the alignment results of EGFR protein sequences with the training dataset.

[0287]

[0288]

[0289] Next, the research team constructed a deep convolutional neural network model M. CNN The amino acid sequences of the five candidate epitope regions were encoded into a feature matrix X and input into the network. Through training and optimization, the model can extract local pattern features Z of the sequence. Simultaneously, the research team also performed secondary structure prediction on the full-length EGFR sequence, calculating the probability of each site forming an α-helix, β-sheet, or random coil structure, generating a structure prediction probability matrix M. struct .

[0290] Table 11 Secondary structure prediction results of EGFR sequences

[0291] Site Alpha helix probability Beta sheet probability Random coil probability 1 0.78 0.12 0.10 2 0.22 0.65 0.13 … … … … 1210 0.45 0.32 0.23

[0292] Figure 5 The study presented the predicted probability distribution of secondary structures for five candidate regions of the EGFR protein, including the predicted probabilities of three structural types: α-helix, β-sheet, and random coil. Furthermore, the research team calculated the relative solvent accessibility area (RSA) of each amino acid residue in the EGFR protein and marked residues with a relative solvent accessibility greater than 25% as surface exposure sites (R0). exposed In addition, they evaluated the evolutionary conservation score C of the EGFR sequence as additional characteristic data.

[0293] The above feature data (sequence features Z, structure prediction M) are used to... struct Surface exposure R exposed Evolutionary conservatism (C) is integrated into a comprehensive feature matrix X. features Input into the previously trained random forest classifier model M RF The model effectively integrates multiple features to predict potential antigenic epitope regions of the EGFR protein.

[0294] Table 12 EGFR antigenic epitope prediction results

[0295] Candidate region Prediction score Judgment 1 0.89 High 2 0.78 High 3 0.65 Medium 4 0.92 High 5 0.71 Medium

[0296] Next, the research team conducted molecular docking simulation analysis on the four candidate regions (1, 2, 4, and 5) with the highest predicted scores. They first constructed a complete solvation model of the EGFR-antibody complex and performed nanosecond-level molecular dynamics sampling under isothermal and isobaric conditions. By analyzing the simulated trajectories, they calculated the electrostatic interaction energy E between each candidate region and the antibody. elec Number of hydrogen bonds N hbond and hydrophobic contact area A hydrophobic Indicators such as...

[0297] Table 13 Molecular docking simulation results

[0298] Candidate region E elec (kJ / mol)]]> <![CDATA[N hbond ]]> A hydrophobic (nm^2)]]> 1 -45.3 7 32.4 2 -38.7 4 27.8 4 -52.6 9 41.2 5 -41.2 6 29.3

[0299] Figure 6 Using a multi-subgraph approach, the performance of the four candidate regions in molecular docking simulations is comprehensively presented, including electrostatic interaction energy, number of hydrogen bonds, hydrophobic contact area, and final comprehensive score.

[0300] Finally, the research team constructed a comprehensive scoring function, Score. total The three main factors of electrostatics, hydrogen bonding, and hydrophobic interactions are weighted and combined, while a penalty term for structural deviation is introduced:

[0301] Score total =0.4·E elec +0.3·N hbond +0.3·A hydrophobic +0.2·RMSD;

[0302] Based on the comprehensive scoring results, the research team selected candidate region 4, which had the highest score, as the preferred EGFR antigenic epitope sequence R. top .

[0303] Table 14 Overall Scoring Results

[0304] Candidate region Score total ]] Judgment 1 87.4 Suboptimal 2 79.2 Suboptimal 4 94.6 Preferred 5 83.5 Suboptimal

[0305] Finally, the research team selected the optimal EGFR antigenic epitope sequence R top The protein chip was fabricated, and its binding activity was detected using a fluorescently labeled secondary antibody. The results showed that the preferred epitope sequence exhibited a strong binding signal with the known EGFR-specific antibody, with a signal-to-noise ratio (SNR) of 12.4.

[0306] It should be noted that the variables involved in this invention are explained in detail in Table 15 below.

[0307] Table 15 Variable Explanation Table

[0308]

[0309]

[0310] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A disease marker antigen epitope prediction and antibody screening method, characterized by, The method comprises the following steps: S10, establishing an antigen epitope database, collecting three-dimensional structure data of known antigen-antibody complexes in a protein database, recording binding site information, collecting expression pattern and localization characteristic data of target proteins in disease-related tissue sections, and integrating to form a training data set; S20, laser capture microdissection is performed on the disease tissue, and the tissue obtained by cutting is subjected to proteomic analysis using a liquid chromatography-mass spectrometry instrument to obtain amino acid sequence and post-translational modification site information of the target protein, the target protein sequence is subjected to sequence alignment with the training data set, and a candidate antigen epitope region is determined; S30, a deep convolutional neural network is constructed, sequence data of the candidate antigen epitope region is input, a sequence feature vector is obtained by training the deep convolutional neural network, secondary structure prediction is performed on the amino acid sequence of the target protein, structure propensity is evaluated, and a structure prediction probability matrix is output; S40, a structure prediction algorithm is used to calculate the relative solvent accessibility area of each amino acid residue in the target protein, a residue exposure threshold of 25% is determined according to the median of the relative solvent accessibility area, and surface-exposed amino acid sites are marked; S50, the sequence feature vector, the structure prediction probability matrix, the surface-exposed amino acid site information, and the sequence evolutionary conservation score are integrated, and a random forest classifier is used to construct an antigen epitope prediction model; S60, molecular dynamics simulation is performed on the antigen-antibody complex predicted by the antigen epitope prediction model, the potential energy evolution trajectory of the antigen-antibody complex is calculated, and a comprehensive scoring function including electrostatic interaction energy, hydrogen bond number and hydrophobic contact area is established; S70, the prediction results of the antigen epitope prediction model are sorted based on the comprehensive scoring function, and the top 10% of antigen epitope sequences with the highest scores are selected as preferred antigen epitope sequences by using a K-means clustering algorithm; S80, an antigen chip of the preferred antigen epitope sequence is prepared, a fluorescence-labeled antibody detection system is used to evaluate the binding activity of each preferred antigen epitope sequence with known antibodies, and the antigen epitope sequence with the strongest binding activity is selected as the final antigen epitope based on the signal-to-noise ratio value.

2. The method of claim 1, wherein the disease marker antigen epitope is selected from the group consisting of: and . In the step S10, the three-dimensional structure data of known antigen-antibody complexes in the protein database is collected, and the binding site information is recorded. Specifically, atomic coordinate information, intermolecular contact residue information, hydrogen bond network information, interface hydrophobic interaction information, and spatial conformation information around the binding site in the antigen-antibody complex crystal structure are recorded.

3. The method of claim 1, wherein the disease marker antigen epitope is selected from the group consisting of: and . In the step S10, the expression pattern and localization characteristic data of the target protein in the disease-related tissue sections are collected, and the training data set is integrated. Specifically, the expression level, subcellular localization, tissue distribution pattern, and co-localization relationship with other markers of the target protein in the diseased tissue are recorded, and the data and the structure information are integrated to form a multi-dimensional characteristic training data set.

4. The method of claim 1, wherein the disease marker antigen epitope is selected from the group consisting of: and . In the step S20, laser capture microdissection is performed on the disease tissue. Specifically, the region of interest is identified under a microscope, and an infrared laser is used to cut along the boundary of the target region. An ultraviolet laser is used to eject the cut region into a collection tube. A lysis buffer is used to extract the proteins in the target region.

5. The method of claim 1, wherein the disease marker antigen epitope prediction and antibody screening method is characterized by, In the step S20, the cut tissue is subjected to proteomic analysis using liquid chromatography-mass spectrometry. Specifically, the peptide mixture is separated by reverse phase chromatography, and the charged peptides are ionized by electrospray ionization. The mass-to-charge ratio of the peptides is detected by quadrupole-time-of-flight mass analysis. The protein composition is identified and post-translational modifications are analyzed by database searching.

6. The method of claim 1, wherein the disease marker antigen epitope is selected from the group consisting of: and . In the step S30, the deep convolutional neural network specifically includes an input layer for receiving the sequence feature matrix, four convolutional layers for extracting local sequence pattern features, two fully connected layers for feature integration, and an output layer for epitope prediction probability output.

7. The method for predicting disease biomarker epitopes and screening antibodies according to claim 1, characterized in that, In the step S30, the deep convolutional neural network is trained to obtain the sequence feature vector. Specifically, the amino acid sequence is encoded as a feature matrix input into the network. The loss function is optimized by stochastic gradient descent. The model performance is evaluated using cross-validation. The output of the second-to-last layer is extracted as the sequence feature vector.

8. The method of claim 1, wherein the disease marker antigen epitope prediction and antibody screening method is characterized by, In the step S60, molecular dynamics simulation is performed on the antigen-antibody complex. Specifically, a complex solvated model is constructed, ions are added to neutralize the system charge, energy minimization and temperature equilibration are performed, and nanosecond-level molecular dynamics sampling is performed under constant temperature and pressure ensemble.

9. The method of claim 1, wherein the disease marker antigen epitope prediction and antibody screening method is characterized by, In the step S70, the potential energy evolution trajectory of the complex system is calculated. Specifically, the stability of the complex conformation in the simulation trajectory is analyzed, the interaction energy between the atoms at the binding interface is calculated, the probability of hydrogen bond formation is calculated, and the time evolution characteristics of the hydrophobic contact area are calculated.

10. The method of claim 1, wherein the disease marker antigen epitope prediction and antibody screening method is characterized by, In the step S80, the preferred epitope sequence antigen chip is prepared. Specifically, the synthesized epitope peptide is fixed on the chip surface by chemical coupling, the unreacted active groups are blocked, and a chip quality control system is established to ensure the uniformity and reproducibility of the array.

Citation Information

Cited By

  • Mass spectrum data clustering method and system, terminal and storage medium

    CN121786520A

  • Mass spectrometry data clustering method, system, terminal and storage medium

    CN121786520B