Method and device for analyzing antigenic dynamic evolution of h3n2 influenza virus

By constructing an antigenicity analysis method based on a two-layer prediction model, the accuracy and systematic problems of H3N2 influenza virus antigenicity evolution analysis in existing technologies have been solved, achieving high-precision antigenicity relationship prediction and antigen cluster identification, supporting influenza virus surveillance and vaccine development.

CN119832993BActive Publication Date: 2025-10-21SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510086854.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-10-21
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Existing technologies are insufficient for quickly and accurately analyzing the antigenic evolution of the H3N2 influenza virus, resulting in inadequate accuracy and comprehensiveness in vaccine strain recommendations, as well as a lack of systematic and model-based predictive capabilities.

Method used

By acquiring multiple neuraminidase virus pairs and multiple effective antigenic features of H3N2 influenza virus, an antigen relationship prediction dataset was constructed. A two-layer prediction model was used for model training to establish an antigen similarity prediction model. Based on the antigen relationship prediction results, an antigen correlation network was constructed for network clustering to achieve high-precision antigen evolution analysis.

Benefits of technology

It has achieved high-precision prediction of the antigenic relationship of H3N2 influenza virus, and can quickly identify new circulating strains and reveal the distribution and evolution of antigen clusters, providing a scientific basis for influenza surveillance and vaccine selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832993B_ABST
    Figure CN119832993B_ABST
Patent Text Reader

Abstract

The application discloses an H3N2 influenza virus antigen dynamic evolution analysis method and device, and aims to solve the problems of poor prediction accuracy, great limitation and difficulty in accurately identifying antigen clusters and evolution rules in the current H3N2 influenza virus antigen analysis method. The method comprises the following steps: obtaining a plurality of neuraminidase virus pairs and a plurality of effective antigen characteristics of the H3N2 influenza virus; performing feature quantization on the plurality of neuraminidase virus pairs based on the plurality of effective antigen characteristics, and constructing an antigen relationship prediction dataset; performing model training on a pre-constructed double-layer prediction model according to the antigen relationship prediction dataset, and obtaining an antigen similarity prediction model; inputting the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigen relationship prediction; constructing an antigen correlation network based on the antigen relationship prediction result, and performing network clustering on the antigen correlation network to obtain an antigen evolution analysis result of the H3N2 influenza virus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of influenza virus analysis, and in particular to a method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, a device for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, an electronic device, and a storage medium. Background Art

[0002] Hemagglutinin (HA) and neuraminidase (NA), the primary surface proteins of the H3N2 influenza virus, play a crucial role in antigenic drift. Currently, most influenza vaccines are designed primarily based on HA antigenicity. While effective in suppressing influenza transmission, HA vaccines offer limited protection against A / H3N2. Studies have shown that NA antibodies can provide a degree of cross-protection, playing a crucial role in preventing influenza infection, inhibiting viral spread, alleviating clinical symptoms, and enhancing vaccine efficacy. The antigenic evolution of the H3N2 influenza virus has a significant impact on vaccine strain selection. Analyzing the antigenic evolution of H3N2 can provide insights into influenza virus antigenic evolution and future vaccine development. Therefore, incorporating NA antigenicity into vaccine development and exploring its evolutionary patterns has become a crucial research direction for improving vaccine effectiveness.

[0003] To better understand the antigenicity of H3N2 NA, current research typically uses experimental methods (such as neuraminidase inhibition assays) to detect antigenic differences between different NA strains and to classify antigenic clusters based on experimental data. However, these traditional methods are complex and time-consuming, and they struggle to obtain rapid results. Furthermore, the fragmented nature of sample sources and their varying quality further limit the efficiency and applicability of these experiments. Furthermore, current experimental methods are typically only capable of analyzing antigenicity between a limited number of viral strains, making them inadequate for large-scale antigenic analysis. Consequently, antigenic analysis results are limited, unable to fully reflect the antigenic evolution of NA, and severely impacting the accuracy and comprehensiveness of vaccine strain recommendations.

[0004] Current analytical methods also lack systematic and model-based predictive capabilities, limiting their ability to accurately predict viral antigenic relationships. Furthermore, due to insufficient training data, existing models are prone to overfitting or prediction bias, failing to effectively capture the complex patterns of NA antigenic evolution. Furthermore, there is currently a lack of systematic analysis of antigenic relationships between viral strains. This makes it difficult to accurately identify antigenic clusters and their evolutionary patterns, making it difficult to provide a scientific basis for vaccine strain recommendations. Summary of the Invention

[0005] The present invention provides a method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, a device for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, an electronic device, and a storage medium, which are used to solve or partially solve the technical problems existing in current H3N2 influenza virus antigen analysis methods, such as poor prediction accuracy, large limitations, and difficulty in accurately identifying antigen clusters and their evolution patterns.

[0006] The present invention provides a method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus, comprising:

[0007] Obtain multiple neuraminidase-virus pairs and multiple effective antigenic features of H3N2 influenza virus;

[0008] quantifying the features of the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features to construct an antigen relationship prediction dataset;

[0009] Performing model training on a pre-built two-layer prediction model according to the antigen relationship prediction dataset to obtain an antigen similarity prediction model;

[0010] Inputting the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigen relationship prediction;

[0011] An antigen correlation network was constructed based on the antigen relationship prediction results, and network clustering was performed on the antigen correlation network to obtain the antigen evolution analysis results of the H3N2 influenza virus.

[0012] The present invention also provides a device for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus, comprising:

[0013] a data acquisition unit for acquiring multiple neuraminidase-virus pairs and multiple effective antigen features of the H3N2 influenza virus;

[0014] a feature quantification unit, configured to perform feature quantification on the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features, and construct an antigen relationship prediction data set;

[0015] A model training unit, configured to perform model training on a pre-built two-layer prediction model based on the antigen relationship prediction data set to obtain an antigen similarity prediction model;

[0016] an antigenic relationship prediction unit, configured to input the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigenic relationship prediction;

[0017] The antigen evolution analysis unit is used to construct an antigen correlation network based on the antigen relationship prediction results, and perform network clustering on the antigen correlation network to obtain the antigen evolution analysis results of the H3N2 influenza virus.

[0018] The present invention further provides an electronic device, comprising a processor and a memory:

[0019] The memory is used to store program code and transmit the program code to the processor;

[0020] The processor is configured to execute the method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus as described above according to the instructions in the program code.

[0021] The present invention also provides a computer-readable storage medium for storing program code, wherein the program code is used to execute the method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus as described in any one of the above items.

[0022] It can be seen from the above technical solutions that the present invention has the following advantages:

[0023] A method for analyzing the dynamic evolution of antigenicity of the H3N2 influenza virus is provided. For model construction, multiple neuraminidase-virus pairs and multiple effective antigenic features of the H3N2 influenza virus are first obtained. Next, based on the multiple effective antigenic features, the multiple neuraminidase-virus pairs are characterized and a dataset for antigenic relationship prediction is constructed. A pre-constructed two-layer prediction model is then trained using the dataset to obtain an antigenic similarity prediction model. By fully integrating the multiple effective antigenic features of the neuraminidase and using an ensemble learning approach based on the two-layer model, a high-precision machine-learning-based antigenic relationship prediction model is developed. This allows for high-precision prediction of antigenic relationships between virus strains in subsequent processes. For the actual analysis of H3N2 influenza virus antigenicity, the sequence data of the neuraminidase to be tested is input into the antigenic similarity prediction model for antigenic relationship prediction. Based on the antigenic relationship prediction results, an antigenic correlation network is constructed, and network clustering is performed on the antigenic correlation network to obtain the results of the antigenic evolution analysis of the H3N2 influenza virus. Based on the prediction results, an antigen correlation network covering global H3N2 NA data is further constructed. Through dynamic network and cluster analysis, it is helpful to quickly identify possible new epidemic strains when viral antigens mutate, reveal the distribution and evolution of antigen clusters, and provide new scientific basis for influenza monitoring and vaccine selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A flowchart of a method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus;

[0026] Figure 2 A schematic diagram of an antigen epitope region and enzyme catalytic site;

[0027] Figure 3 This is a schematic diagram comparing the test set results when using different antigen similarity prediction models;

[0028] Figure 4 This is a schematic diagram comparing cross-validation results when using different antigen similarity prediction models;

[0029] Figure 5 A schematic diagram of an antigen correlation network and phylogenetic tree;

[0030] Figure 6 It is a heat map of intra-cluster and inter-cluster similarity of an antigen cluster;

[0031] Figure 7 A graph showing the average cluster size and modularity changing with the expansion parameter for a network clustering;

[0032] Figure 8 A schematic diagram showing the temporal and spatial prevalence of H3N2 viruses based on neuraminidase and hemagglutinin;

[0033] Figure 9 The figure is a structural block diagram of a device for analyzing the dynamic evolution of H3N2 influenza virus antigenicity. DETAILED DESCRIPTION

[0034] The embodiments of the present invention provide a method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, a device for analyzing the dynamic evolution of H3N2 influenza virus antigenicity, an electronic device, and a storage medium, which are used to solve or partially solve the technical problems existing in current H3N2 influenza virus antigen analysis methods, such as poor prediction accuracy, large limitations, and difficulty in accurately identifying antigen clusters and their evolution patterns.

[0035] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0036] As an example, to better understand the antigenicity of H3N2 NA, current research typically uses experimental methods (such as neuraminidase inhibition assays) to detect antigenic differences between different NA strains and to classify antigenic clusters based on experimental data. However, these traditional methods are complex and time-consuming, making it difficult to obtain rapid results. Furthermore, the dispersed and variable sample sources further limit the efficiency and applicability of the experiments. Furthermore, current experimental methods are typically only capable of analyzing antigenicity between a limited number of viral strains, making it difficult to meet the needs of large-scale antigenic analysis. This results in limited antigenic analysis results that fail to fully reflect the antigenic evolution of NA, seriously impacting the accuracy and comprehensiveness of vaccine strain recommendations.

[0037] Current analytical methods also lack systematic and model-based predictive capabilities, limiting their ability to accurately predict viral antigenic relationships. Furthermore, due to insufficient training data, existing models are prone to overfitting or prediction bias, failing to effectively capture the complex patterns of NA antigenic evolution. Furthermore, there is currently a lack of systematic analysis of antigenic relationships between viral strains. This makes it difficult to accurately identify antigenic clusters and their evolutionary patterns, making it difficult to provide a scientific basis for vaccine strain recommendations.

[0038] Therefore, one of the core inventive aspects of the present invention is to provide a neuraminidase-based method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity. By fully integrating multiple key features of neuraminidase (such as epitope characteristics, physicochemical properties, glycosylation sites, and catalytic sites), and using an ensemble learning approach with multiple machine learning models, a high-precision machine learning-based antigenic relationship prediction model is developed, enabling high-precision prediction of antigenic relationships between virus strains. Based on the prediction results, an antigenic correlation network covering global H3N2 NA data is further constructed. Through dynamic network and cluster analysis, this helps to rapidly identify potential new epidemic strains when viral antigens mutate, revealing the distribution and evolution of antigenic clusters, and providing new scientific evidence for influenza surveillance and vaccine selection.

[0039] Reference Figure 1 , shows a flowchart of a method for analyzing the dynamic evolution of antigenicity of an H3N2 influenza virus provided by an embodiment of the present invention, which may specifically include the following steps:

[0040] Step 101, obtaining multiple neuraminidase-virus pairs and multiple effective antigen features of H3N2 influenza virus;

[0041] In some embodiments, the process of obtaining multiple neuraminidase-virus pairs and multiple effective antigenic features of H3N2 influenza virus can be achieved by performing the following sub-steps S01 to S03:

[0042] Step S01: collecting neuraminidase sequence data and neuraminidase inhibition test data of H3N2 influenza virus;

[0043] First, the present invention collected all H3N2 subtype influenza virus neuraminidase sequence data from the Global Initiative on Sharing All Influenza Data (GISAID) database, which covers virus samples from different months and continents between 1968 and 2024.

[0044] In addition, the present invention also conducted extensive literature collection and screened for test data pairs with both homologous and heterologous titers. A total of 377 pairs of neuraminidase inhibition (NAI) test data were compiled. These data were derived from neuraminidase inhibition experiments between different virus strains and documented the antigenic differences between virus pairs. The neuraminidase inhibition test data provide experimental support for the antigenic relationships between viral neuraminidases and are important for calculating antigenic distances and training antigenic similarity prediction models. This data contains rich information on antigenic relationships, which, combined with subsequent rigorous screening and optimization, can improve the reliability and usability of subsequent analyses.

[0045] Step S02: performing sequence screening and double optimization on the neuraminidase sequence data to obtain optimized sequence data, and performing feature selection based on the optimized sequence data to extract multiple effective antigen features;

[0046] In order to improve the data quality and the effect of model training, it is necessary to perform multi-step data processing on the collected neuraminidase sequence data and neuraminidase inhibition test data.

[0047] In some embodiments, the process of performing sequence screening and double optimization on neuraminidase sequence data to obtain optimized sequence data can be achieved by executing the following sub-steps S11 to S17:

[0048] Step S11: performing uniformly distributed random sampling on the neuraminidase sequence data to obtain sampled neuraminidase sequence data;

[0049] The processing of neuraminidase sequence data mainly involves eliminating low-quality or incomplete sequences and adopting a unified naming format (such as virus strain name, date, and place of origin) for standardized management.

[0050] The sequence sampling process begins with a random sampling of all neuraminidase sequence data, selecting influenza samples from different continents (North America, Europe, Asia, South America, Africa, and Oceania) between 1968 and 2024. This process evenly distributes sampling by month to ensure global representativeness. Specifically, for each sampling point in each month for each continent, seven sequences are randomly sampled from the corresponding sequence data for that month. If fewer than seven sequences are available for a particular month, all sequences for that month are sampled.

[0051] This equal random sampling strategy ensures data balance across geographic and temporal dimensions. Uniformly distributed random sampling ensures that the amount of data from each continent is roughly equal, effectively mitigating bias in model training due to high-data-volume regions (such as North America and Europe), allowing the model to fully learn the antigenic characteristics of regions with lower data volumes. Furthermore, by drawing balanced samples monthly, the data maintains temporal continuity and representativeness, helping to reflect trends and patterns in the long-term evolution of H3N2 antigenicity. Uniformly distributed random sampling reduces sample duplication and redundancy, making the training data more representative and diverse. By selecting sample data from different time periods and geographic regions, the generalizability and broad applicability of the subsequent prediction model can be significantly improved, enabling it to accurately reflect antigenic changes in neuraminidase levels in the H3N2 virus.

[0052] Although the likelihood of this occurring is low, when sequencing is conducted globally, there may be cases where the number of sequences in certain months or regions is low or even zero. When the number of sequences is zero, methods such as duplicating data to supplement the data may introduce unrealistic noise or bias, affecting the accuracy of antigen analysis. Therefore, in this case, there is no need to duplicate or supplement the data. It is understandable that although sequences may be insufficient in certain months, this localized deficiency has little impact on the overall analysis. Global data representativeness can be supplemented and balanced with data from other months and continents, thereby reducing the impact of local deficiencies.

[0053] To further validate the reliability of the sampling strategy, the present invention also tested a percentage-based proportional sampling method. Specifically, 5% of the sequences (at least one sequence) were selected for each month and continent. Subsequent analysis showed that the two sequence sampling methods achieved highly consistent results in antigen cluster delineation and vaccine strain matching. This demonstrates that the uniform sampling method can better avoid bias caused by variations in sample size.

[0054] Step S12: performing a multiple sequence alignment on the sampled neuraminidase sequence data, performing sequence alignment based on the multiple sequence alignment results, and screening out low-quality neuraminidase sequences with gaps exceeding a preset gap threshold during the sequence alignment process;

[0055] The sampled neuraminidase sequence data can be aligned using the MAFFT (Multiple Alignment using Fast Fourier Transform) multiple sequence alignment tool, and sequence alignment can be performed based on the multiple sequence alignment results. Sequence alignment and comparison ensures relative consistency between sequences, eliminating potential sequence errors and insertions and deletions. Furthermore, the present invention incorporates a strict quality filter during the sequence alignment process. Low-quality neuraminidase sequences with gaps exceeding 10% are deleted.

[0056] Step S13: selecting at least one target neuraminidase sequence with a vacancy from all retained neuraminidase sequences;

[0057] Subsequently, the Bayesian interpolation method can be used to dynamically fill the small number of gaps in the retained neuraminidase sequences. Therefore, at least one target neuraminidase sequence with a gap can be selected from all the retained neuraminidase sequences.

[0058] Step S14: for each target neuraminidase sequence, determine the amino acid vacancy site of the target neuraminidase sequence and record the sequence context amino acids upstream and downstream of the amino acid vacancy site;

[0059] For each target neuraminidase sequence, determine the vacant amino acid site and record the upstream and downstream sequence context amino acids (e.g., 3 amino acids before and after).

[0060] Step S15: screening out, from all retained complete neuraminidase sequences, at least one target complete neuraminidase sequence having an amino acid similarity match with the sequence context higher than a preset match threshold;

[0061] Next, you can search for related sequences—sequences with similar contextual amino acids (three amino acids before and after) to the gap in the target neuraminidase sequence. Specifically, you can use complete neuraminidase sequences that don't require gap filling (which can be interpreted as other high-quality sequences in the alignment results) as references. Then, you can find the top 10 complete neuraminidase sequences that are most similar to the target neuraminidase sequence context. These sequences will serve as reference sequences for dynamically filling gaps in the target neuraminidase sequence.

[0062] Step S16: Based on at least one target complete neuraminidase sequence, predicting gap filling values ​​for amino acid gaps by Bayesian interpolation, and dynamically filling the target neuraminidase sequence based on the gap filling values ​​to obtain a filled neuraminidase sequence;

[0063] For example, suppose there is a gap at a position in the target neuraminidase sequence, and its context fragments are ALK and VGT. From all retained complete neuraminidase sequences (high-quality sequences), the top 10 sequences with the best context matches, such as ALKQVGT and ALKSVGT, are selected. The most likely value of the gap's amino acid is predicted using Bayesian interpolation, ultimately filling the gap with "Q." For the target gap's position in the collection of similar sequences, the distribution probability of the amino acid at that position is calculated.

[0064] According to Bayes' theorem, the posterior probability of the amino acid at the target position is calculated as follows:

[0065]

[0066] in, ) represents the prior probability of the amino acid in the database; represents the conditional probability of the context given the amino acid X; represents the overall probability of the context in all sequences.

[0067] After calculating the posterior probabilities of the amino acid positions that could be filled in each target neuraminidase sequence, the amino acid with the highest probability is used to fill the target gap, completing dynamic correction. Assuming a relatively uniform probability distribution (i.e., multiple amino acids with similar probabilities), the mode of the highest-probability amino acids is selected as the final imputation result. Dynamic correction improves data integrity and reduces data loss caused by removing low-quality sequences. This dynamic consideration of context and sequence similarity enhances imputation accuracy.

[0068] Step S17: Integrate all filled-in neuraminidase sequences and all complete neuraminidase sequences, and add the H3N2 recommended vaccine strains from previous years to obtain optimized sequence data.

[0069] All padded and complete neuraminidase sequences were integrated, and the H3N2 vaccine strains recommended by the World Health Organization over the years were added to this sequence data. Through the aforementioned processing steps, 8,992 neuraminidase sequences were obtained as optimized sequence data. Through alignment and filtering, the remaining data was high-quality neuraminidase sequence data, providing a reliable foundation for subsequent feature extraction and quantification.

[0070] Extracting effective features from viral sequences is key to predicting whether an antigen is similar to a known strain. In the present embodiment, antigenic features are divided into four categories based on the characteristics of neuraminidase: epitope features, physicochemical properties, N-glycosylation site features, and enzyme catalytic site features.

[0071] In some embodiments, the process of performing feature selection based on the optimized sequence data to extract multiple effective antigen features can be achieved by executing the following sub-steps S21 to S27:

[0072] Step S21: using a variety of different prediction tools to perform structural analysis on each neuraminidase sequence in the optimized sequence data, and recording all residues predicted to be antigenic epitopes;

[0073] The first step is to extract antigen epitope features. In this example, based on the 7U4E template from the PDB (Protein Data Bank) database, the structure of H3N2 neuraminidase was analyzed and its potential structural epitopes were predicted by combining four tools: Discotope-2.0, Discotope-3.0, SEMA2.0, and ScanNet, to improve the accuracy and reliability of the prediction.

[0074] These four tools each have their own unique prediction principles and algorithmic approaches, providing different perspectives on epitope recognition. This paper combines the prediction results of these tools to propose an optimized multi-tool comprehensive prediction strategy, overcoming the limitations of a single prediction method.

[0075] Among them, the main prediction principle of Discotope-2.0 is based on the solvent accessibility (Solvent Accessible Surface Area, SASA) of protein residues and the statistical tendency of amino acids. It assumes that antibodies mainly bind to areas exposed on the protein surface, and further combines the distribution tendency of amino acids in known epitope data to calculate the epitope score of each residue.

[0076] DiscoTope-3.0 is a structure-based epitope prediction tool. It uses a high-dimensional structural representation generated by an inverse protein folding model to predict B cell conformational epitopes. Unlike traditional methods that rely solely on solvent accessibility (SASA) and geometric properties, DiscoTope-3.0 converts the three-dimensional protein structure into a feature vector that can be processed by deep learning models. This includes residue surface accessibility, the AlphaFold structure quality score (predicted Local Distance Difference Test, pLDDT), and a sequence length correction score.

[0077] SEMA2.0 is a multimodal deep learning tool that combines protein sequence and three-dimensional structure information. By introducing a large-scale pre-trained language model (Evolutionary Scale Modeling 2, ESM2) and a structure-aware protein module (SaProt) that incorporates geometric patterns, SEMA2.0 is able to capture epitope characteristics from both sequence and structure dimensions.

[0078] ScanNet is a geometric deep learning model that extracts the geometric and physicochemical features of binding sites from protein 3D structures. Using geometric relationships among residue neighborhoods, combined with chemical information such as solvent accessibility and charge distribution, ScanNet employs a convolutional neural network (CNN) to perform multi-scale modeling of residue spatial properties and predict conformational epitopes.

[0079] Step S22: Calculating the number of times each residue is predicted to be an antigenic epitope, and taking the residues whose antigenic epitope prediction number is greater than or equal to a preset number threshold as potential structural epitopes;

[0080] To integrate the prediction results from the four tools above, embodiments of the present invention devise a weighted synthesis strategy. During data processing, the prediction results from each tool are collated and statistically analyzed, and the residue numbers predicted as epitopes by each tool are recorded. Next, the number of times each residue is predicted as an epitope by the four tools is calculated. For example, if a residue is predicted as an epitope by two tools, the number is recorded as 2. A predefined threshold (e.g., two) can be set, indicating that residues predicted as epitopes by at least two tools are considered highly reliable and can be selected as weighted epitope predictions.

[0081] Specifically, for each residue, if it is predicted twice or more by all four tools, it is considered a possible antigenic epitope, i.e., a potential structural epitope. This weighted comprehensive strategy combines the strengths of all four tools, retaining the robustness of DiscoTope-3.0 in predicting structural models while leveraging ScanNet's ability to model geometric details and the multimodal adaptability of SEMA2.0. This approach effectively reduces the risk of false positives or false negatives that can arise from a single tool, significantly improving the accuracy and reliability of predictions.

[0082] In the present embodiment, selecting twice as the threshold not only achieves a good balance between sensitivity and specificity, but also covers most potential true epitopes while avoiding omissions caused by overly high standards. Overall, selecting twice as the threshold can meet practical needs.

[0083] Step S23: performing cluster analysis based on all potential structural epitopes, determining the number of clusters according to the Euclidean distance relationship and silhouette coefficient between each potential structural epitope, and obtaining five antigen epitope features;

[0084] The predicted potential structural epitopes were clustered using the K-means clustering algorithm. The number of clusters was determined based on the Euclidean distance relationship and the silhouette score between the epitopes. The present invention finally identified five antigenic epitopes through clustering. They were named N2_A, N2_B, N2_C, N2_D, and N2_E. The antigenic epitope region is shown in Figure 2. Figure 2 These epitopes were derived by analyzing the frequency distribution of viral mutation positions, ensuring the accuracy and applicability of epitope definition. Furthermore, the extraction of epitope features provides important input features for subsequent antigenicity prediction.

[0085] The site information contained in each antigen epitope is shown in Table 1 below:

[0086]

[0087] Table 1: Site information contained in each antigen epitope

[0088] Step S24: screening hydrophobic residues, charge distribution residues, charge polarity residues, molecular volume residues, and solvent accessible surface area residues from each neuraminidase sequence in the optimized sequence data;

[0089] The present invention screened based on biological significance and selected five physicochemical properties that play an important role in antigen-antibody interaction, including hydrophobicity, charge, polarity, volume, and (solvent) accessible surface area.

[0090] Hydrophobicity is one of the key drivers of antigen-antibody binding. Antibodies tend to bind to regions of low hydrophobicity. These regions are often located on the surface of the antigen, facilitating antibody recognition. Furthermore, aromatic hydrophobic residues (such as tyrosine, phenylalanine, and tryptophan) play a crucial role at the antigen-antibody interface. These residues not only enhance binding affinity through π-π interactions but also provide stable binding sites for antigenic epitopes through their unique geometry.

[0091] Charge distribution directly influences the binding specificity between antibodies and antigens. The charge distribution in the epitope region typically contains a judicious mix of positively and negatively charged residues. This distribution enhances binding stability through electrostatic interactions.

[0092] The binding site of an antibody is typically enriched with negatively charged residues (such as aspartic acid and glutamic acid). In contrast, the epitope is enriched with positively charged residues (such as arginine and lysine). This charge complementarity not only enhances binding strength but also optimizes the geometric fit of the binding interface. Epitope regions are typically exposed to the solvent and have a high accessible surface area. The high exposure of these regions allows antibodies to more easily access and bind to the epitope.

[0093] Furthermore, solvent-accessible surface area is closely related to the conformational flexibility of the antigen epitope, increasing the adaptability of the antibody to the epitope. Polar residues (such as glutamine, asparagine, serine, and threonine) significantly enhance the affinity and stability of antigen-antibody binding by forming weak interactions such as hydrogen bonds. These residues tend to be distributed on the antigen surface, further enhancing binding specificity.

[0094] Molecular volume characteristics reflect the geometric contributions of residues. Particularly in complex epitopes, larger residues (such as tryptophan and phenylalanine) directly facilitate the antigen-antibody interaction by providing a larger binding surface area. Meanwhile, smaller residues (such as glycine and alanine) indirectly optimize the binding interface by providing conformational flexibility. These geometric properties enable the epitope to better fit within the antibody binding site.

[0095] Step S25: Based on all hydrophobic residues, charge distribution residues, charge polarity residues, molecular volume residues, and solvent accessible surface area residues, feature selection is performed using a random forest model, and five physicochemical property features representing hydrophobicity, charge, polarity, volume, and accessible surface area are screened out based on feature importance scores;

[0096] Each physicochemical property often has multiple quantitative representations in bioinformatics databases. To identify the specific parameter values ​​that best represent each physicochemical property, quantitative values ​​for five physicochemical properties of H3N2 neuraminidase were obtained from the AAindex (Amino Acid Index Database). In this embodiment, a random forest model was used for feature selection to construct a feature matrix.

[0097] Each row of the feature matrix corresponds to a residue in the H3N2 neuraminidase, and each column represents the value of a parameter. Each residue is labeled as an epitope (positive sample, 1) or a non-epitope (negative sample, 0) based on existing experimental data. During the training of the random forest model, multiple decision trees are used to classify whether a residue is an epitope. The contribution of each parameter to the prediction task is evaluated, and a feature importance score is generated based on the information gain. Through feature selection and screening using the random forest method, the representative features ultimately selected include BLAS910101 (hydrophobicity), FAUJ880111 (charge), RADA880108 (polarity), GRAR740103 (volume), and ROSG850101 (accessible surface area).

[0098] It's understandable that the purpose of using the random forest model in this step is to determine whether each residue is an epitope and identify which physicochemical property values ​​are most important for this determination. In other words, each physicochemical property is assigned a score, called a feature importance score, based on which properties are most conducive to correctly classifying the epitope. The predicted epitope serves as the classification label, while the physicochemical property values ​​serve as the input.

[0099] Step S26: predicting the N-glycosylation site of each neuraminidase sequence in the optimized sequence data to obtain N-glycosylation site features;

[0100] The N-glycosylation site characteristics can be obtained by predicting the N-glycosylation sites of the neuraminidase sequence using the NetNGlyc website prediction tool.

[0101] Step S27: Based on the data collection and analysis, the enzyme catalytic site characteristics of the H3N2 influenza virus neuraminidase are determined.

[0102] Based on the literature, the present invention collected 13 H3N2 neuraminidase catalytic sites as enzyme catalytic site characteristics, the site schematic is as follows Figure 2 shown.

[0103] It should be noted that the 12 features selected in the present invention have been validated and can significantly improve the accuracy and robustness of predictions based on H3N2 neuraminidase-antigen relationships. In addition, other feature selection methods can be used instead of the random forest model to select physicochemical property features. For example, automated feature selection algorithms (SelectKBest (a feature selection method based on univariate statistical testing) and recursive feature elimination) can quickly identify important features through statistical analysis or recursive screening. For another example, model-based feature selection (such as LASSO regression (Least Absolute Shrinkage and Selection Operator, a linear regression analysis method)) can effectively handle high-dimensional data and avoid overfitting. Alternatively, feature extraction methods based on deep learning (such as convolutional neural networks) can enhance the automation level of feature selection, further improving the accuracy and versatility of the model. It is understood that the present invention is not limited to this.

[0104] Step S03: Deduplication is first performed on the neuraminidase inhibition test data, and then antigen distance quantification is performed to obtain multiple neuraminidase-virus pairs.

[0105] For neuraminidase inhibition test data, we used the median of multiple neuraminidase inhibition test results for the same H3N2 virus pair as the final neuraminidase inhibition test titer value, removing redundant data to ensure data consistency. This step not only reduced redundancy in the experimental data but also eliminated the interference of duplicate experimental data on model training results.

[0106] The present invention obtains 377 pairs of neuraminidase virus data through the above data processing of the neuraminidase inhibition test data.

[0107] In order to effectively train the model later, it is necessary to quantify the antigenic differences between the 377 virus pairs collected by calculating the antigenic distance to obtain the final 377 neuraminidase virus pairs for model training. The calculation formula for the antigenic distance is as follows:

[0108]

[0109] Specifically, represents the NA antigenic distance between strains a and b; To use antiserum against virus b Detect the NAI titer (homologous titer) of virus b; To use antiserum against virus b The NAI titer (heterologous titer) of virus a was tested. If the absolute value of the NA antigenic distance between a pair of strains was greater than or equal to 2, they were classified as an antigenically dissimilar pair; otherwise, the pair of strains was considered an antigenically similar pair.

[0110] Step 102, quantifying the features of the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features, and constructing an antigen relationship prediction dataset;

[0111] After obtaining 12 effective antigen features in four categories through the previous steps, it is necessary to quantify the neuraminidase virus pairs based on these effective antigen features so that the subsequent model can perform machine learning training and antigen similarity prediction.

[0112] In some embodiments, the process of quantifying features of multiple neuraminidase-virus pairs based on multiple effective antigen features and constructing an antigen relationship prediction dataset can be achieved by performing the following sub-steps S31 to S35:

[0113] Step S31: for each antigenic epitope feature, comparing whether the amino acids of each site in the corresponding epitope region of the neuraminidase virus pair are different, and taking the number of differences as the antigenic epitope feature value of the neuraminidase virus pair in the antigenic epitope feature;

[0114] For each antigenic epitope, the amino acids of each site in the corresponding epitope region of each pair of viruses are compared to see if there are any differences, and the number of differences is used as a characteristic value to reflect the potential for antigenic changes of the virus at that epitope.

[0115] Step S32: For each physicochemical property characteristic, calculate the difference value of the neuraminidase virus pair at each site, and select the three largest difference values ​​to calculate the average value as the physicochemical property characteristic value of the neuraminidase virus pair in the physicochemical property characteristic;

[0116] For each physicochemical property, the difference in each physicochemical property was calculated for each pair of viral sequences. The specific calculation method was to use the 20 amino acid physicochemical property data from the AAindex database to calculate the difference value at each site between the virus pairs and take the average of the three largest differences.

[0117] Step S33: calculating the number of N-glycosylation site differences of the neuraminidase virus pair as the N-glycosylation characteristic value of the N-glycosylation site characteristics of the neuraminidase virus pair;

[0118] The N-glycosylation characteristic value was obtained by calculating the number of different N-glycosylation sites between virus pairs.

[0119] Step S34: Calculate the Euclidean distance value from each differential site to the enzyme catalytic site in the neuraminidase virus pair, and select the three largest Euclidean distance values ​​to calculate the average value as the enzyme catalytic site characteristic value of the neuraminidase virus pair at the enzyme catalytic site;

[0120] The Euclidean distance from each differential site in the virus pair to the catalytic site was calculated, and the average of the three shortest Euclidean distances was selected as the characteristic value of the enzyme catalytic site.

[0121] Step S35: Integrate the antigenic epitope characteristic values, physicochemical property characteristic values, N-glycosylation characteristic values, and enzyme catalytic site characteristic values ​​of all neuraminidase-virus pairs to construct an antigenic relationship prediction data set.

[0122] Step 103, performing model training on a pre-built two-layer prediction model based on the antigen relationship prediction dataset to obtain an antigen similarity prediction model;

[0123] Traditional related methods often select a single machine learning algorithm as the final prediction model. This invention uses a stacking ensemble learning method to combine multiple different machine learning algorithms to construct an antigenicity prediction model. Specifically, based on the classification capabilities of the models, five models were selected: Logistic Regression (Logistic Regression), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Random Forest (RF), and XGBoost (Extreme Gradient Boosting, a gradient boosting algorithm). This fully leverages the characteristics and advantages of multiple machine learning algorithms to construct a more efficient antigenicity prediction model.

[0124] More specifically, the two-layer prediction model proposed in this invention includes a first-layer basic prediction model and a second-layer comprehensive prediction model. The first-layer basic prediction model includes multiple different prediction sub-models. Logistic regression, support vector machine, K-nearest neighbor, random forest, and XGBoost can be selected as prediction sub-models. Logistic regression is selected as the second-layer comprehensive prediction model.

[0125] In some embodiments, the process of training a pre-built two-layer prediction model based on the antigen relationship prediction dataset to obtain an antigen similarity prediction model can be achieved by executing the following sub-steps S41 to S45:

[0126] Step S41: dividing the antigen relationship prediction data set into an antigen relationship prediction training set and an antigen relationship prediction test set according to a preset division ratio;

[0127] In practice, the number of antigenically similar and dissimilar virus pairs in an antigenic relationship prediction dataset is unbalanced. To avoid bias in the model's prediction results, the dataset can be balanced using the Synthetic Minority Over-sampling Technique (SMOTE). SMOTE generates an equal number of antigenically similar virus pairs as there are antigenically dissimilar virus pairs, thereby enhancing data balance. The balanced dataset is then divided into an antigenic relationship training set and an antigenic relationship test set in a 7:3 ratio. The antigenic relationship training set is used for model training, while the antigenic relationship test set is used to verify the model's generalization performance.

[0128] Step S42: performing model training on each prediction sub-model using the antigen relationship prediction training set;

[0129] In conjunction with the previous introduction, the present invention introduces a stacking ensemble learning method. A basic model is constructed in the first layer, and logistic regression, support vector machine, K-nearest neighbor, random forest, and XGBoost are used as prediction sub-models for model training. Each prediction sub-model learns from the antigen relationship prediction training set and generates corresponding prediction results as feature input to the second layer, i.e., a two-layer comprehensive prediction model.

[0130] Step S43: inputting the prediction results of each prediction sub-model into the second-layer comprehensive prediction model for model training based on comprehensive prediction;

[0131] The two-layer comprehensive prediction model uses a meta-learner and selects a logistic regression model with relatively stable performance. During the training process, the final prediction output is generated by integrating the prediction results of each prediction sub-model.

[0132] Step S44: During the model training process, cross-validation is used to perform layered training on the prediction results of each prediction sub-model, and model parameters of the two-layer comprehensive prediction model and each prediction sub-model are optimized through grid parameter adjustment and random parameter adjustment;

[0133] To avoid overfitting, cross-validation is used during model training to perform layered training on the prediction results of each prediction sub-model, ensuring better generalization of the model at each layer. Furthermore, grid and random parameter tuning methods can be used to optimize the hyperparameters of each algorithm to further improve model performance.

[0134] Step S45: using the antigen relationship prediction test set to perform model evaluation on the trained two-layer comprehensive prediction model and each trained prediction sub-model, and output the final antigen similarity prediction model.

[0135] Model evaluation mainly uses accuracy, precision, F1 score, recall rate and AUC (Area Under the Curve) as evaluation indicators, and verifies the actual effect of the model by drawing the ROC (Receiver Operating Characteristic Curve) characteristic curve.

[0136] In order to verify the feasibility of the technical solution provided by the present invention, illustratively, Figure 3 A schematic diagram showing the comparison of test set results when using different antigen similarity prediction models is shown. Figure 4 A schematic diagram showing the comparison of cross-validation results when using different antigen similarity prediction models is shown.

[0137] The stacking ensemble model proposed in this study, as a model for predicting antigenic similarity, outperformed single models in most comprehensive performance metrics. For example, in an independent test set, the stacking ensemble model achieved the best accuracy (0.852), F1 score (0.833), recall (0.833), and precision (0.833). In terms of AUC, the stacking ensemble model (AUC=0.918) performed slightly worse than the random forest (AUC=0.922). In cross-validation, the stacking ensemble model achieved the best accuracy (0.875), F1 score (0.882), recall (0.892), and precision (0.875). In terms of AUC, the stacking ensemble model (AUC=0.948) performed slightly worse than the random forest (AUC=0.949). This indicates that while the random forest has a slight advantage in discriminatory ability, the stacking ensemble model, by combining the predictions of multiple models, achieves overall performance improvements across multiple metrics. Compared with the traditional single model method, the Stacking ensemble model combines the advantages of multiple algorithms and effectively improves the prediction accuracy of antigen-similar and variant virus pairs through SMOTE data balancing processing.

[0138] It should be noted that while the present invention constructs an antigen relationship prediction model based on a stacking ensemble learning model, those skilled in the art can also employ other machine learning methods, such as LightGBM (Light Gradient Boosting Machine, a machine learning algorithm based on a gradient boosting framework), CatBoost (a machine learning algorithm based on a gradient boosting framework), and neural network models. These alternatives can achieve similar antigen relationship prediction results, further enriching the technical implementation and application scenarios of the present invention.

[0139] Among them, the models listed above can be used as alternatives to the Stacking integrated model and used alone for antigen relationship prediction. It can also be used as a prediction sub-model of the basic model in the Stacking integrated model, and used in combination with models such as Logistic Regression, Support Vector Machine, K Nearest Neighbor, Random Forest and XGBoost. It should be noted that the more sub-models there are in the Stacking integrated model, the more training time and computing resource requirements will increase significantly. In particular, the neural network model has a lower training efficiency than models such as LightGBM, and increased complexity will lead to increased difficulty in parameter adjustment and optimization. Therefore, those skilled in the art can select a suitable model as a first-layer basic model or a second-layer comprehensive prediction model based on the actual situation and comprehensive training requirements. It will be understood that the present invention is not limited to this.

[0140] Step 104, inputting the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigen relationship prediction;

[0141] This step mainly uses the trained antigenic similarity prediction model to predict the antigenic relationship of the virus sequence data actually required for detection. The antigenic relationship prediction result indicates that the H3N2 neuraminidase virus pair in the neuraminidase inhibition test data is antigenically similar or antigenically dissimilar.

[0142] In practice, once a prediction model is trained, antigenic relationships can be predicted by pairing collected neuraminidase sequence data, yielding the probability of antigenic similarity and dissimilarity between virus pairs. The core principle of antigenic similarity prediction is the probability ratio of antigenic variation. This involves first calculating the ratio of the probability of similarity to the probability of dissimilarity (oddsratio), then taking the logarithm of this ratio as the criterion for determining antigenic relationships.

[0143] When the logarithm of the ratio is greater than 0, the viruses are considered antigenically similar. When the logarithm of the ratio is less than or equal to 0, the viruses are considered antigenically dissimilar. Compared to traditional methods for calculating antigenic relationships, the present invention uses logarithmic transformation of probability ratios to more intuitively reflect the significance of antigenic relationships between virus strains, avoiding the inaccuracy that may be caused by a single probability threshold.

[0144] This eliminates the need to rely on complex experimental facilities or consumables. By building a machine learning model, the prediction of antigen similarity can be quickly completed, reducing experimental time and saving experimental condition costs. It can also be expanded to the analysis of large-scale viral data, significantly improving the efficiency of antigen monitoring.

[0145] Step 105: construct an antigen correlation network based on the antigen relationship prediction result, and perform network clustering on the antigen correlation network to obtain an antigen evolution analysis result of the H3N2 influenza virus.

[0146] In some embodiments, the process of constructing an antigen correlation network based on the antigen relationship prediction results and performing network clustering on the antigen correlation network to obtain the antigen evolution analysis results of the H3N2 influenza virus can be achieved by executing the following sub-steps S51 to S55:

[0147] Step S51: network visualization of the antigen relationship prediction results;

[0148] The antigen relationship prediction results were visualized using Cytoscape software (an open source biological network visualization and analysis software).

[0149] Step S52: Each virus strain in the neuraminidase sequence data to be tested is treated as a node, and the antigenic similarity between virus strains is used as an edge connecting the nodes. Based on each node and each node-connecting edge, an antigenic correlation network of H3N2 influenza virus is constructed;

[0150] Each H3N2 virus strain is treated as a node in the network, and pairs of antigenically similar viruses are connected by edges, forming an Antigen Correlation Network (ACNet). Antigenically similar viruses are clustered together, making it easier to observe and analyze the antigenic relationships between different virus strains.

[0151] Step S53: constructing a phylogenetic tree of the H3N2 influenza virus based on the antigen correlation network, wherein the phylogenetic tree is used to reflect the genetic continuity and evolutionary process between virus strains;

[0152] Through the antigen correlation network, five major antigen clusters can be preliminarily identified and a phylogenetic tree can be constructed. Figure 5 Vaccine strains are distributed in clusters and named after the earliest vaccine strain. Phylogenetic trees are based on genetic sequence variation, while antigenic clusters more closely reflect phenotypic differences due to epitope variation.

[0153] Figure 5 The center left panel focuses on visualizing the distribution of antigenic clusters and highlighting the vaccine strain name for each cluster. These clusters reflect important trends in viral antigenicity and simplify the complex information in the phylogenetic tree. The right panel displays the phylogenetic relationships of all viral strains, demonstrating the genetic continuity and evolutionary process among strains. In practice, antigenic and genetic changes are not necessarily completely linear; small genetic changes may occur while significant antigenic changes occur.

[0154] Figure 6A heatmap of intra-cluster and inter-cluster similarity for an antigenic cluster is shown. The heatmap shows that virus strains within a cluster have been verified to have higher antigenic similarity (darker colors indicate higher similarity values) than those between clusters, demonstrating the effectiveness of the cluster.

[0155] Step S54: performing network clustering on the antigen correlation network using Markov clustering to identify virus clusters with similar antigens in the antigen correlation network, and obtaining multiple antigen clusters;

[0156] The Markov Clustering Algorithm (MCL) was used for network clustering to identify antigenically similar virus clusters within the antigen-association network. During the clustering process, edge weights were set to the logarithm of the probability ratio of antigenic change to ensure the rationality of virus clustering. Clustering parameters included average cluster size and modularity. By adjusting these parameters, clustering results can be optimized and reasonable antigen clustering results can be obtained.

[0157] Step S55: performing antigenic evolution analysis of the H3N2 influenza virus based on multiple antigenic clusters to obtain corresponding antigenic evolution analysis results.

[0158] Figure 7 Figure 2 shows a graph showing the average cluster size and modularity of a network cluster as a function of the expansion parameter. Figure 7 As shown, blue represents the average cluster size, corresponding to the left y-axis. Green represents the modularity, corresponding to the right y-axis. Figure 7 The parameter corresponding to the first point of the first plateau period is 1.4. The higher the modularity, the better the clustering effect.

[0159] Through cluster analysis, multiple antigenic clusters were obtained, and each cluster contains a group of viruses with similar antigenicity. The division of these antigenic clusters provides data support for analyzing the antigenic evolution and transmission path of the H3N2 virus. On the basis of the antigenic clusters, by tracking the spatiotemporal distribution of each antigenic cluster, observing the replacement process of old and new antigenic clusters, analyzing the emergence and spread trends of different antigenic clusters around the world or in each continent, and comparing the similarities and differences in the evolution of hemagglutinin (HA) and neuraminidase (NA). The schematic diagram of the spatiotemporal epidemic map of H3N2 virus based on neuraminidase and the comparison based on hemagglutinin is shown in the figure below. Figure 8 As shown. A comparative analysis of the hemagglutinin and neuraminidase of the viruses in the antigenic cluster revealed a phenomenon of co-evolution between the two during antigenic evolution. Compared with the rapid mutation of hemagglutinin, the antigenic changes of neuraminidase are more stable. Therefore, through this multi-level antigenic evolution analysis, the present invention not only reveals the antigenic evolution pattern of H3N2 viruses, but also provides important scientific basis for vaccine selection and antigen prediction.

[0160] It should be noted that the present invention utilizes Cytoscape software for the construction and visualization analysis of antigen-association networks. Alternatively, network construction tools (Gephi or NetworkX) can be used in place of Cytoscape for visualization and topological analysis of antigen networks. Gephi is more suitable for interactive analysis, while NetworkX is more suitable for scripted processing. Other network analysis algorithms, such as the graph-embedding-based Node2Vec or GraphSAGE algorithms, can also be used to further explore hidden antigenic relationships between viral strains. Dynamic network models, such as time series network models, can also be used to dynamically demonstrate the formation and spread trends of antigen clusters. It is understood that the present invention is not limited to this.

[0161] In an embodiment of the present invention, a method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity based on neuraminidase is provided. By fully integrating multiple key features of neuraminidase (such as epitope characteristics, physicochemical properties, glycosylation sites, and catalytic sites), and using an ensemble learning approach with multiple machine learning models, a high-precision machine learning-based antigenic relationship prediction model is developed, enabling high-precision prediction of antigenic relationships between virus strains. Based on the prediction results, an antigenic correlation network covering global H3N2 NA data is further constructed. Through dynamic network and cluster analysis, this helps to quickly identify potential new epidemic strains when viral antigens mutate, revealing the distribution and evolution of antigenic clusters, and providing new scientific basis for influenza surveillance and vaccine selection. Thus, the technical solution provided by the present invention enables dynamic evolution analysis of antigenic clusters based on H3N2 antigenicity prediction, allowing tracking the formation of new antigenic clusters and predicting their spread trends. Furthermore, this antigenic similarity prediction model and network analysis method can be extended to analyze the antigenicity of whole genome data or other viral proteins, better adapting to the complexity and diversity of viral antigenic changes.

[0162] Reference Figure 9 , shows a structural block diagram of a device for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus provided by an embodiment of the present invention, which may specifically include:

[0163] A data acquisition unit 901 is used to acquire multiple neuraminidase-virus pairs and multiple effective antigen features of the H3N2 influenza virus;

[0164] A feature quantification unit 902 is configured to perform feature quantification on the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features to construct an antigen relationship prediction dataset;

[0165] A model training unit 903 is configured to perform model training on a pre-built two-layer prediction model based on the antigen relationship prediction dataset to obtain an antigen similarity prediction model;

[0166] The antigen relationship prediction unit 904 is used to input the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigen relationship prediction;

[0167] The antigen evolution analysis unit 905 is used to construct an antigen correlation network based on the antigen relationship prediction result, and perform network clustering on the antigen correlation network to obtain the antigen evolution analysis result of the H3N2 influenza virus.

[0168] In an optional embodiment, the data acquisition unit 901 includes:

[0169] Influenza virus data collection unit, used to collect neuraminidase sequence data and neuraminidase inhibition test data of H3N2 influenza virus;

[0170] an effective antigen feature extraction unit, configured to perform sequence screening and dual optimization on the neuraminidase sequence data to obtain optimized sequence data, and perform feature selection based on the optimized sequence data to extract a plurality of effective antigen features;

[0171] The neuraminidase-virus pair generating unit is used to first deduplicate the neuraminidase inhibition test data and then perform antigen distance quantification processing to obtain multiple neuraminidase-virus pairs.

[0172] In an optional embodiment, the effective antigen feature extraction unit includes:

[0173] a uniformly distributed random sampling unit, configured to perform uniformly distributed random sampling on the neuraminidase sequence data to obtain sampled neuraminidase sequence data;

[0174] a multiple sequence alignment unit, configured to perform multiple sequence alignment on the sampled neuraminidase sequence data, perform sequence alignment based on the multiple sequence alignment results, and screen out low-quality neuraminidase sequences whose gaps exceed a preset gap threshold during the sequence alignment process;

[0175] a target neuraminidase sequence screening unit, for selecting at least one target neuraminidase sequence with a vacancy from all retained neuraminidase sequences;

[0176] an amino acid vacancy determination unit, configured to determine, for each target neuraminidase sequence, an amino acid vacancy of the target neuraminidase sequence and record the sequence context amino acids upstream and downstream of the amino acid vacancy;

[0177] a target complete neuraminidase sequence screening unit, configured to screen out, from all retained complete neuraminidase sequences, at least one target complete neuraminidase sequence having an amino acid similarity match with the sequence context higher than a preset match threshold;

[0178] a site dynamic filling unit for predicting gap filling values ​​of the amino acid sites by Bayesian interpolation based on the at least one target complete neuraminidase sequence, and dynamically filling the target neuraminidase sequence based on the gap filling values ​​to obtain a filled neuraminidase sequence;

[0179] The sequence data integration unit is used to integrate all the filled neuraminidase sequences and all the complete neuraminidase sequences, and add the H3N2 recommended vaccine strains of previous years to obtain optimized sequence data.

[0180] In an optional embodiment, the effective antigen feature extraction unit includes:

[0181] a sequence structure analysis unit, configured to perform structural analysis on each neuraminidase sequence in the optimized sequence data using a plurality of different prediction tools, and record all residues predicted to be antigenic epitopes;

[0182] a potential structural epitope screening unit, for respectively calculating the number of times each of the residues is predicted to be an antigenic epitope, and selecting residues whose antigenic epitope prediction number is greater than or equal to a preset number threshold as potential structural epitopes;

[0183] an antigen epitope feature determination unit, configured to perform cluster analysis based on all the potential structural epitopes, determine the number of clusters according to the Euclidean distance relationship and silhouette coefficient between each of the potential structural epitopes, and obtain five antigen epitope features;

[0184] a sequence residue screening unit, for screening hydrophobic residues, charge distribution residues, charge polarity residues, molecular volume residues and solvent accessible surface area residues from each neuraminidase sequence in the optimized sequence data;

[0185] a physicochemical property feature determination unit, configured to perform feature selection using a random forest model based on all the hydrophobic residues, the charge distribution residues, the charge polarity residues, the molecular volume residues, and the solvent accessible surface area residues, and screen out five physicochemical property features representing hydrophobicity, charge, polarity, volume, and accessible surface area, respectively, based on feature importance scores;

[0186] An N-glycosylation site feature determination unit is used to predict the N-glycosylation site of each neuraminidase sequence in the optimized sequence data to obtain the N-glycosylation site feature;

[0187] The enzyme catalytic site feature determination unit is used to determine the enzyme catalytic site features of the H3N2 influenza virus neuraminidase based on data collection and analysis.

[0188] In an optional embodiment, the feature quantization unit 902 includes:

[0189] an antigenic epitope feature quantification unit, for comparing, for each antigenic epitope feature, whether the amino acids of each site in the corresponding epitope region of the neuraminidase virus pair are different, and taking the number of differences as the antigenic epitope feature value of the neuraminidase virus pair in the antigenic epitope feature;

[0190] a physicochemical property characteristic quantification unit, configured to calculate, for each physicochemical property characteristic, a difference value of the neuraminidase virus pair at each site, and select the three largest difference values ​​to calculate an average value as the physicochemical property characteristic value of the neuraminidase virus pair in the physicochemical property characteristic;

[0191] An N-glycosylation site characteristic quantification unit is used to calculate the number of differences in the N-glycosylation sites of the neuraminidase virus pair as the N-glycosylation characteristic value of the neuraminidase virus pair at the N-glycosylation site characteristics;

[0192] an enzyme catalytic site characteristic quantification unit, configured to calculate the Euclidean distance value from each differential site in the neuraminidase virus pair to the enzyme catalytic site, and select the three largest Euclidean distance values ​​to calculate the average value as the enzyme catalytic site characteristic value of the neuraminidase virus pair at the enzyme catalytic site;

[0193] The antigenic relationship prediction data set construction unit is used to integrate the antigenic epitope characteristic values, physicochemical property characteristic values, N-glycosylation characteristic values ​​and enzyme catalytic site characteristic values ​​of all the neuraminidase virus pairs to construct an antigenic relationship prediction data set.

[0194] In an optional embodiment, the two-layer prediction model includes a first-layer basic prediction model and a second-layer comprehensive prediction model; the first-layer basic prediction model includes multiple different prediction sub-models; the model training unit 903 includes:

[0195] A data set division unit, configured to divide the antigen relationship prediction data set into an antigen relationship prediction training set and an antigen relationship prediction test set according to a preset division ratio;

[0196] A prediction sub-model training unit, configured to perform model training on each of the prediction sub-models using the antigen relationship prediction training set;

[0197] A comprehensive prediction model training unit, configured to input the prediction results of each prediction sub-model into the second-layer comprehensive prediction model for model training based on comprehensive prediction;

[0198] A model parameter optimization unit is used to perform layered training on the prediction results of each prediction sub-model using cross-validation during the model training process, and to optimize the model parameters of the two-layer comprehensive prediction model and each prediction sub-model respectively through grid parameter adjustment and random parameter adjustment;

[0199] The model evaluation unit is used to use the antigen relationship prediction test set to perform model evaluation on the trained two-layer comprehensive prediction model and each trained prediction sub-model, and output the final antigen similarity prediction model.

[0200] In an optional embodiment, the antigenic relationship prediction result is that the H3N2 neuraminidase virus pair in the neuraminidase inhibition test data is antigenically similar, or antigenically dissimilar; the antigenic evolution analysis unit 905 includes:

[0201] A network visualization unit, configured to perform network visualization on the antigen relationship prediction result;

[0202] An antigen correlation network construction unit is used to treat each virus strain in the neuraminidase sequence data to be tested as a node, and treat the antigenic similarity relationship between virus strains as node edges, and construct an antigen correlation network of H3N2 influenza virus based on each of the nodes and each of the node edges;

[0203] A phylogenetic tree construction unit, configured to construct a phylogenetic tree of the H3N2 influenza virus based on the antigen correlation network, wherein the phylogenetic tree is configured to reflect the genetic continuity and evolutionary process between virus strains;

[0204] a network clustering unit, configured to perform network clustering on the antigen correlation network using Markov clustering to identify virus clusters with similar antigens in the antigen correlation network and obtain a plurality of antigen clusters;

[0205] The antigen evolution analysis subunit is used to perform antigen evolution analysis of the H3N2 influenza virus based on the multiple antigen clusters to obtain corresponding antigen evolution analysis results.

[0206] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the aforementioned method embodiment.

[0207] An embodiment of the present invention further provides an electronic device, the device including a processor and a memory:

[0208] The memory is used to store program codes and transmit the program codes to the processor;

[0209] The processor is configured to execute the method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to any embodiment of the present invention according to the instructions in the program code.

[0210] An embodiment of the present invention further provides a computer-readable storage medium for storing program code, and the program code is used to execute the method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to any embodiment of the present invention.

[0211] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0212] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0213] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0214] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0215] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0216] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus, characterized in that: include: Obtain multiple neuraminidase virus pairs of H3N2 influenza virus and multiple effective antigen features processed by sequence screening dual optimization and feature selection; quantifying the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features, and constructing an antigen relationship prediction data set based on the plurality of effective antigen feature values ​​corresponding to each of the neuraminidase-virus pairs after the feature quantification; Performing model training on a pre-built two-layer prediction model according to the antigen relationship prediction dataset to obtain an antigen similarity prediction model; Inputting the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigen relationship prediction; Constructing an antigen correlation network based on the antigen relationship prediction results, and performing network clustering on the antigen correlation network to obtain antigen evolution analysis results of the H3N2 influenza virus; The two-layer prediction model includes a first-layer basic prediction model and a second-layer comprehensive prediction model; the first-layer basic prediction model includes multiple different prediction sub-models; the pre-constructed two-layer prediction model is trained according to the antigen relationship prediction data set to obtain an antigen similarity prediction model, including: Dividing the antigen relationship prediction data set into an antigen relationship prediction training set and an antigen relationship prediction test set according to a preset division ratio; Performing model training on each of the prediction sub-models using the antigen relationship prediction training set; Inputting the prediction results of each prediction sub-model into the second-layer comprehensive prediction model for model training based on comprehensive prediction; During the model training process, cross-validation is used to perform layered training on the prediction results of each prediction sub-model, and the model parameters of the two-layer comprehensive prediction model and each prediction sub-model are optimized respectively through grid parameter adjustment and random parameter adjustment; The antigen relationship prediction test set is used to perform model evaluation on the trained two-layer comprehensive prediction model and each trained prediction sub-model, and a final antigen similarity prediction model is output.

2. The method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to claim 1, wherein: The method of obtaining multiple neuraminidase virus pairs and multiple effective antigenic features of H3N2 influenza virus includes: Collect neuraminidase sequence data and neuraminidase inhibition test data of H3N2 influenza virus; performing sequence screening and double optimization on the neuraminidase sequence data to obtain optimized sequence data, and performing feature selection based on the optimized sequence data to extract multiple effective antigen features; The neuraminidase inhibition test data are first deduplicated and then subjected to antigen distance quantification processing to obtain multiple neuraminidase-virus pairs.

3. The method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to claim 2, wherein: The step of performing sequence screening and double optimization on the neuraminidase sequence data to obtain optimized sequence data comprises: Performing uniformly distributed random sampling on the neuraminidase sequence data to obtain sampled neuraminidase sequence data; Performing a multiple sequence alignment on the sampled neuraminidase sequence data, performing sequence alignment based on the multiple sequence alignment results, and screening out low-quality neuraminidase sequences with gaps exceeding a preset gap threshold during the sequence alignment process; Select at least one target neuraminidase sequence with a gap from all retained neuraminidase sequences; For each target neuraminidase sequence, determine the amino acid vacancy site of the target neuraminidase sequence, and record the sequence context amino acids upstream and downstream of the amino acid vacancy site; Screening out, from all retained complete neuraminidase sequences, at least one target complete neuraminidase sequence having an amino acid similarity match degree with the sequence context higher than a preset match degree threshold; Based on the at least one target complete neuraminidase sequence, predicting gap filling values ​​for the amino acid gaps by Bayesian interpolation, and dynamically filling the target neuraminidase sequence based on the gap filling values ​​to obtain a filled neuraminidase sequence; All the filled-in neuraminidase sequences and all the complete neuraminidase sequences were integrated, and the H3N2 recommended vaccine strains from previous years were added to obtain optimized sequence data.

4. The method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to claim 2, wherein: The feature selection based on the optimized sequence data to extract multiple effective antigen features includes: Using a variety of different prediction tools to perform structural analysis on each neuraminidase sequence in the optimized sequence data, and recording all residues predicted to be antigenic epitopes; Calculating the number of times each of the residues is predicted to be an antigenic epitope, and taking the residues whose antigenic epitope prediction number is greater than or equal to a preset number threshold as potential structural epitopes; Performing cluster analysis based on all the potential structural epitopes, determining the number of clusters according to the Euclidean distance relationship and silhouette coefficient between each of the potential structural epitopes, and obtaining five antigenic epitope features; screening hydrophobic residues, charge distribution residues, charge polarity residues, molecular volume residues and solvent accessible surface area residues from each neuraminidase sequence in the optimized sequence data; Based on all the hydrophobic residues, the charge distribution residues, the charge polarity residues, the molecular volume residues and the solvent accessible surface area residues, feature selection is performed using a random forest model, and five physicochemical property features representing hydrophobicity, charge, polarity, volume and accessible surface area are screened out according to feature importance scores; Predicting the N-glycosylation site of each neuraminidase sequence in the optimized sequence data to obtain N-glycosylation site characteristics; Based on data collection and analysis, the enzyme catalytic site characteristics of H3N2 influenza virus neuraminidase were determined.

5. The method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to claim 4, wherein: The step of quantifying the features of the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features to construct an antigen relationship prediction dataset comprises: For each antigenic epitope feature, comparing whether the amino acids of each site in the corresponding epitope region of the neuraminidase virus pair are different, and taking the number of differences as the antigenic epitope feature value of the neuraminidase virus pair in the antigenic epitope feature; For each physicochemical property characteristic, calculate the difference value of the neuraminidase virus pair at each site, and select the three largest difference values ​​to solve the average value as the physicochemical property characteristic value of the neuraminidase virus pair in the physicochemical property characteristic; Calculating the number of differences in the N-glycosylation sites of the neuraminidase virus pair as the N-glycosylation characteristic value of the neuraminidase virus pair at the N-glycosylation site; Calculating the Euclidean distance value from each differential site to the enzyme catalytic site in the neuraminidase virus pair, and selecting the three largest Euclidean distance values ​​to calculate the average value as the enzyme catalytic site characteristic value of the neuraminidase virus pair at the enzyme catalytic site characteristic; The antigenic epitope characteristic values, physicochemical property characteristic values, N-glycosylation characteristic values ​​and enzyme catalytic site characteristic values ​​of all the neuraminidase virus pairs were integrated to construct an antigenic relationship prediction data set.

6. The method for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus according to claim 1, wherein: The antigenic relationship prediction result is that the H3N2 neuraminidase virus pair in the neuraminidase inhibition test data is antigenically similar, or antigenically dissimilar; constructing an antigen correlation network based on the antigenic relationship prediction result, and performing network clustering on the antigen correlation network to obtain an antigenic evolution analysis result of the H3N2 influenza virus, including: Performing network visualization on the antigen relationship prediction results; Each virus strain in the neuraminidase sequence data to be tested is used as a node, and the antigenic similarity relationship between the virus strains is used as a node edge, and an antigenic correlation network of the H3N2 influenza virus is constructed based on each of the nodes and each of the node edges; constructing a phylogenetic tree of H3N2 influenza virus based on the antigen correlation network, wherein the phylogenetic tree is used to reflect the genetic continuity and evolution process between virus strains; Performing network clustering on the antigen correlation network using Markov clustering to identify virus clusters with similar antigens in the antigen correlation network, thereby obtaining a plurality of antigen clusters; Antigen evolution analysis of the H3N2 influenza virus is performed based on the multiple antigen clusters to obtain corresponding antigen evolution analysis results.

7. A device for analyzing the dynamic evolution of antigenicity of H3N2 influenza virus, characterized in that: include: A data acquisition unit, for acquiring multiple neuraminidase virus pairs of H3N2 influenza virus and multiple effective antigen features processed by sequence screening dual optimization and feature selection; a feature quantification unit, configured to perform feature quantification on the plurality of neuraminidase-virus pairs based on the plurality of effective antigen features, and construct an antigen relationship prediction data set based on the plurality of effective antigen feature values ​​corresponding to each of the neuraminidase-virus pairs after feature quantification; A model training unit, configured to perform model training on a pre-built two-layer prediction model based on the antigen relationship prediction data set to obtain an antigen similarity prediction model; an antigenic relationship prediction unit, configured to input the neuraminidase sequence data to be tested into the antigen similarity prediction model to perform antigenic relationship prediction; An antigen evolution analysis unit is used to construct an antigen correlation network based on the antigen relationship prediction results, and perform network clustering on the antigen correlation network to obtain antigen evolution analysis results of the H3N2 influenza virus; The two-layer prediction model includes a first-layer basic prediction model and a second-layer comprehensive prediction model; the first-layer basic prediction model includes multiple different prediction sub-models; and the model training unit includes: A data set division unit, configured to divide the antigen relationship prediction data set into an antigen relationship prediction training set and an antigen relationship prediction test set according to a preset division ratio; A prediction sub-model training unit, configured to perform model training on each of the prediction sub-models using the antigen relationship prediction training set; A comprehensive prediction model training unit, configured to input the prediction results of each prediction sub-model into the second-layer comprehensive prediction model for model training based on comprehensive prediction; A model parameter optimization unit is used to perform layered training on the prediction results of each prediction sub-model using cross-validation during the model training process, and to optimize the model parameters of the two-layer comprehensive prediction model and each prediction sub-model respectively through grid parameter adjustment and random parameter adjustment; The model evaluation unit is used to use the antigen relationship prediction test set to perform model evaluation on the trained two-layer comprehensive prediction model and each trained prediction sub-model, and output the final antigen similarity prediction model.

8. An electronic device, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity according to any one of claims 1 to 6 according to the instructions in the program code.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the method for analyzing the dynamic evolution of H3N2 influenza virus antigenicity according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for rapidly identifying antigenicity of H9N2 subtype avian influenza strain

    CN117912568A

  • Novel method and device for rapidly identifying antigenicity of coronavirus strain

    CN118782148A