Rapid antigenicity discrimination method, device, terminal and medium for influenza B virus
By extracting the viral characteristics of influenza B virus and establishing an antigenic prediction model, the problem of long and low efficiency of antigenic identification of influenza B virus in the prior art is solved, and the rapid and accurate identification of antigenicity of influenza B virus strains is achieved, and the effectiveness of influenza vaccines is improved.
Patent Information
- Application Number
- CN202410989243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-23
AI Technical Summary
The existing antigenic identification methods for influenza B viruses are time-consuming and inefficient, and are not suitable for large-scale antigenic identification of influenza strains.
By obtaining the HA1 sequence sample data and sample antigenic data of the Victoria lineage and Yamagata lineage of influenza B virus, viral characteristics, including antigenic epitope characteristics, physical and chemical properties characteristics, nitrogen glycosylation sites and receptor binding region characteristics, and establishing antigenic prediction models PREDAC-BV and PREDAC-BY to achieve rapid identification of antigenicity of the identified strains.
It improves the efficiency of antigenic identification of influenza B virus, can quickly and accurately identify the antigenicity of influenza B virus strains, and helps improve the effectiveness and production efficiency of influenza vaccines.
Smart Images

Figure CN118942538B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of virus antigen research, and particularly to a method, device, terminal and medium for rapidly identifying the antigenicity of influenza B virus. Background Art
[0002] Seasonal influenza is a huge threat to human health. Currently, influenza A and influenza B are the main types prevalent in the population. Among them, influenza B virus includes two lineages, Victoria and Yamagata, accounting for about 1 / 3 of the global influenza disease burden, and usually has a higher infection rate and mortality rate in children, posing a huge threat to human health.
[0003] Vaccination is currently the most effective way to prevent influenza virus infection and related deaths. However, influenza viruses continuously accumulate mutations through surface proteins, leading to changes in their antigenicity, thereby reducing the effectiveness of existing vaccines. We need to promptly replace the vaccine that matches the antigenicity of the currently prevalent influenza strains for effective prevention and control.
[0004] Therefore, accurately and rapidly identifying the antigenicity of influenza viruses is very important for improving the effectiveness of influenza vaccines. Currently, the gold standard for identifying the antigenicity of influenza virus strains is the HI test (Hemagglutination Inhibition Assay), but this identification method requires a huge amount of manpower, material resources and financial resources, and has technical problems such as long time consumption and low efficiency, and is not suitable for large-scale identification of the antigenicity of influenza virus strains. Researchers have developed some prediction tools that can rapidly and efficiently identify the antigenicity of influenza A virus, but there is currently no prediction tool for identifying the antigenicity of influenza B virus.
[0005] Therefore, there is an urgent need for a technology that can accurately and rapidly identify the antigenicity of influenza B virus, so as to promptly evaluate the antigenicity of the strains of the two lineages of currently prevalent influenza B, which is of great significance for improving the effectiveness of influenza vaccines, the R & D and production efficiency, and influenza prevention and control. Summary of the Invention
[0006] This application provides a method, device, terminal and medium for rapidly identifying the antigenicity of influenza B virus, which is used to solve the technical problems of the existing method for identifying the antigenicity of influenza B virus being restricted by unstable factors such as experimental conditions and technical levels, and having long time consumption and low efficiency.
[0007] To solve the above technical problems, the first aspect of this application provides a method for rapidly identifying the antigenicity of influenza B virus, including:
[0008] Obtain the sample data of the HA1 sequences and the sample antigenicity data of the preset Victoria lineage and Yamagata lineage of influenza B virus respectively, wherein the sample antigenicity data respectively include the antigenic relationships between pairwise strains within the Victoria lineage and the Yamagata lineage;
[0009] Extract the virus characteristics of the HA1 sequence sample data of the Victoria lineage and Yamagata lineage of influenza B virus respectively, and respectively obtain the virus characteristics related to antigen change corresponding to the sample data sets of the Victoria lineage and Yamagata lineage. The virus characteristics specifically include: epitope characteristics, physicochemical property characteristics, N-glycosylation sites, and receptor-binding region characteristics;
[0010] Using the virus characteristics and the sample antigenicity data as model training data, through model training, respectively obtain the antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the Victoria lineage and Yamagata lineage of influenza B virus;
[0011] Obtain the HA1 amino acid sequence of the influenza B virus strain of the Victoria lineage or Yamagata lineage to be identified, and input the HA1 amino acid sequence of the influenza B virus strain into the antigenicity prediction model of the corresponding lineage, so as to determine the antigenicity identification result of the influenza B virus strain according to the output result of the antigenicity prediction model.
[0012] Preferably, the extraction method of the epitope characteristic value specifically includes:
[0013] According to the preset amino acid site screening conditions, determine the amino acid sites that meet the amino acid site screening conditions in the HA1 protein sequences of the influenza B virus samples of the Victoria lineage and Yamagata lineage as the antigen sites of each lineage;
[0014] Cluster the antigen sites, and divide the antigen sites into five epitopes;
[0015] Construct a PSSM matrix containing each of the antigen sites through the PSI-BLAST algorithm;
[0016] According to a pair of sample data of the same lineage of influenza B virus selected, respectively take the sum of the absolute values of the differences in the PSSM scores of the two amino acid residues corresponding to each site on the epitope as the characteristic value of the epitope. When the characteristic values of all epitopes are obtained, output the epitope characteristic value.
[0017] Preferably, the extraction method of the physicochemical property characteristic value specifically includes:
[0018] According to the amino acid index database, determine the physicochemical property indexes corresponding to the Victoria lineage and the Yamagata lineage. Among them, the physicochemical property indexes include: hydrophobicity, volume, charge, polarity, and accessible surface area;
[0019] According to a pair of B-type influenza virus sample data of the same selected lineage and any selected physicochemical property index, calculate the average value of the differences in the corresponding physicochemical property scores at all amino acid-different sites of the pair of strains as the characteristic value of the physicochemical property. When there are more than three different amino acid sites in a pair of strains, then take the average value of the top three sites with the largest differences in the corresponding physicochemical property scores at all amino acid-different sites as the characteristic value of the physicochemical property. After obtaining the characteristic values of all physicochemical property indexes, output the characteristic values of the physicochemical properties.
[0020] Preferably, the extraction method of the receptor-binding region characteristics specifically includes:
[0021] According to a pair of B-type influenza virus sample data of the same selected lineage, calculate the average value of the shortest Euclidean distances from the amino acid-different sites on the HA1 protein sequence of the pair of strains to the receptor-binding region characteristics as the receptor-binding region characteristics.
[0022] Preferably, the training method of the antigenicity prediction model includes:
[0023] Based on the HA1 sequence data of the B-type influenza Victoria lineage and Yamagata lineage virus samples and the sample antigenicity data, combined with a variety of preset initial algorithm models, respectively use the HA1 sequence data of the Victoria lineage and Yamagata lineage samples and the sample antigenicity data as the input quantities of each initial algorithm model, train each initial algorithm model, and perform five-fold cross-validation;
[0024] According to the obtained cross-validation results, determine the optimal initial algorithm model corresponding to each lineage to obtain the best antigenicity prediction models for different lineages.
[0025] Preferably, the types of the initial algorithm models include: XGBoost, LightGBM, RF, and SVM.
[0026] Preferably, after determining the antigenicity discrimination result of the B-type influenza virus strain, it further includes:
[0027] Predict the antigenic relationship between all known Victoria lineage or Yamagata lineage strains based on the constructed prediction model, and construct an antigenic similarity network between strains. Use the Markov clustering method to cluster the antigenic similarity network of strains in the same lineage to identify different antigen classes. Assist in vaccine strain recommendation by analyzing the antigenic similarity rate between strains of the same lineage prevalent in the current epidemic season and the antigenic clusters of the same lineage that dominated the epidemic in the past. When the antigenic similarity rate is low, it is recommended to replace the vaccine strain of the same lineage, otherwise the vaccine strain is not replaced.
[0028] Meanwhile, the second aspect of the present application provides a rapid B-type influenza virus antigenicity identification device, including:
[0029] A sample data acquisition unit for respectively acquiring sample data of HA1 sequences and sample antigenicity data of the preset Victoria lineage and Yamagata lineage of B-type influenza virus, wherein the sample antigenicity data respectively includes the antigenic relationship between pairwise strains within the Victoria lineage and the Yamagata lineage;
[0030] An antigenic feature extraction unit for respectively extracting virus features from the HA1 sequence sample data of the Victoria lineage and Yamagata lineage of the B-type influenza virus, and respectively obtaining virus features related to antigen change corresponding to the sample data sets of the Victoria lineage and Yamagata lineage. The virus features specifically include: antigenic epitope features, physicochemical property features, N-glycosylation sites, and receptor binding region features;
[0031] An antigenicity prediction model training unit for using the virus features and the sample antigenicity data as model training data, and respectively obtaining antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the Victoria lineage and Yamagata lineage of the B-type influenza virus through model training;
[0032] An antigenicity identification unit for obtaining the HA1 amino acid sequence of the B-type influenza virus Victoria lineage or Yamagata lineage strain to be identified, and inputting the HA1 amino acid sequence of the B-type influenza virus strain into the antigenicity prediction model of the corresponding lineage, so as to determine the antigenicity identification result of the B-type influenza virus strain according to the output result of the antigenicity prediction model.
[0033] The third aspect of the present application provides a rapid B-type influenza virus antigenicity identification terminal, including: a memory and a processor;
[0034] The memory is used for storing program codes corresponding to a rapid B-type influenza virus antigenicity identification provided in the first aspect of the present application;
[0035] The processor is used to read and execute the program code.
[0036] A fourth aspect of the present application provides a computer-readable storage medium, in which program code corresponding to a rapid antigenicity discrimination of influenza B virus provided in the first aspect of the present application is stored.
[0037] It can be seen from the above technical solutions that the present application has the following advantages:
[0038] The technical solution provided by the present application is based on the HA1 protein sequence samples and antigenicity sample data of the Victoria lineage and Yamagata lineage of influenza B virus. Through feature extraction, virus characteristics corresponding to the sample data sets of the two lineages of influenza B virus are obtained. According to the virus characteristics and sample antigenicity data as model training data, through model training, antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the two lineages of influenza B virus are respectively obtained. Then, the Victoria lineage or Yamagata lineage strain of the influenza B virus to be discriminated is input into the above antigenicity prediction model, and according to the output result of the antigenicity prediction model, the discrimination result of the strain is determined. The present application establishes an antigenicity prediction model for the two lineages Victoria and Yamagata in influenza B virus based on machine learning algorithms, and then realizes rapid antigenicity prediction for the two lineages of influenza B virus respectively based on the constructed antigenicity prediction model, thereby improving the efficiency of influenza B virus antigenicity discrimination. At the same time, the present application clusters all the Victoria lineage strains of influenza B virus based on the constructed prediction model to obtain different antigen classes, and assists in vaccine strain recommendation by calculating the antigen similarity rate between the strains in the current epidemic season and the antigen clusters that dominated the past epidemics, which helps to improve the effectiveness of influenza vaccines in preventing Victoria lineage infections. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of an embodiment of a method for rapid antigenicity discrimination of influenza B virus provided by the present application.
[0041] Figure 2 It is an overall logical block diagram of a method for rapid antigenicity discrimination of influenza B virus provided by the present application.
[0042] Figure 3 This is a flowchart for determining the antigenic epitopes of the Victoria lineage and Yamagata lineage of influenza B virus in this application.
[0043] Figure 4 This is the effect diagram of the recommended results of the antigen prediction model assisting the Victoria lineage vaccine strain in this application.
[0044] Figure 5 This is a schematic structural diagram of an embodiment of a rapid antigenicity identification device for influenza B virus provided in this application.
[0045] Figure 6 This is a schematic structural diagram of an embodiment of a rapid antigenicity identification terminal for influenza B virus provided in this application. Detailed implementation manners
[0046] The embodiments of this application provide a rapid antigenicity identification method, device, terminal and medium for influenza B virus, which are used to solve the technical problems of long time consumption and low efficiency existing in the existing influenza B virus identification methods restricted by unstable factors such as experimental conditions and technical levels.
[0047] To make the invention purpose, features and advantages of this application more obvious and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described below are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0048] First, a detailed description of an embodiment of a rapid antigenicity identification method for influenza B virus provided in this application is as follows:
[0049] Please refer to Figure 1 , a rapid antigenicity identification method for influenza B virus provided in this embodiment includes:
[0050] Step 101: Obtain the HA1 sequence sample data and sample antigenicity data on the HA (Haemagglutinin) protein of the preset Victoria lineage and Yamagata lineage of influenza B virus respectively.
[0051] Among them, the sample antigenicity data respectively include the antigenic relationships between pairwise strains within the Victoria lineage and the Yamagata lineage.
[0052] It should be noted that the solution provided in this example first obtains the HA1 protein sequence sample data of two preset B influenza virus lineages Victoria and Yamagata, as well as the antigenic data of these B influenza virus sample data. The sample antigenic data can be obtained by collecting existing hemagglutination inhibition test data on the Victoria lineage and the Yamagata lineage to determine the antigenic relationship between some strains. The specific implementation method can refer to the following example:
[0053] Data collection and processing of virus samples: Download the amino acid sequences of HA proteins of Victoria lineage and Yamagata lineage from the global influenza data, and then perform the following processing on the downloaded sequence data:
[0054] 1) First, the sequences were quality controlled and the sequences that met the following criteria were excluded: the sequences whose names did not start with B and whose length was within the box-type Figure 4 Sequences outside the quantile interval of 1.5 times and sequences with lineage mismatches.
[0055] 2) Then, MAFFT (Multiple Alignment using Fast Fourier Transform) was used to align the HA sequences by multiple sequence alignment method. After multiple sequence alignment, the HA1 sequence was extracted from the aligned sequences.
[0056] 3) Finally, the missing and abnormal amino acids (i.e., "-", "X") at the beginning (first 40 amino acids) and end (last 40 amino acids) of the sequence were filled with the mode amino acid of the 10 most similar sequences measured by Hamming distance at the same site, because mutations rarely occur at these positions.
[0057] HI test data collection and processing: Download the Victoria and Yamagata lineage HI test data from the World Influenza Center Laboratory (WIC), and then perform the following processing on the downloaded HI data:
[0058] 1) Based on the results of the HI test, the antigenic relationship between strains a and b was determined using the following formula:
[0059] H aβ =log2T aβ -log2T aβ (1)
[0060] Among them, H aβ represents the HI titer of strain a produced by titrating the antiserum against reference strain b, T aβIndicates the maximum dilution value of antiserum β against strain b that can effectively prevent cell agglutination caused by strain a. When the absolute value of H between strain a and strain b is equal to or greater than 2, the two are defined as antigenically dissimilar; otherwise, they are defined as antigenically similar. aβ When the absolute value is equal to or greater than 2, the two are defined as antigenically dissimilar; otherwise, they are defined as antigenically similar.
[0061] 2) To obtain a more reliable data set, this application only includes strain pairs with two or more HI test results. For strain pairs with two HI test results, only the strain pairs with consistent antigenic relationships determined by formula 1 are retained; for strain pairs with more than two HI test results, their antigenic relationships are determined as the mode of the antigenic relationships determined by formula 1.
[0062] Match the processed HA1 amino acid sequences and the strain pairs with antigenic relationships determined by HI tests according to the strain names as the data set for constructing the antigenic relationship prediction model. Among them, there are 895 pairs in the Victoria lineage (301 pairs are antigenically similar and 594 pairs are antigenically dissimilar), and 751 pairs in the Yamagata lineage (322 pairs are antigenically similar and 429 pairs are antigenically dissimilar).
[0063] In addition, in addition to the above method of obtaining data from public databases, it is also possible to obtain the sample antigenicity data of these B-type influenza virus sample data by performing hemagglutination inhibition tests on the B-type influenza virus sample data and recording the test results.
[0064] Step 102: Extract virus characteristics from the HA1 sequence sample data of the B-type influenza virus Victoria lineage and Yamagata lineage respectively, and obtain the virus characteristics related to antigen change corresponding to the sample data sets of the Victoria lineage and Yamagata lineage respectively.
[0065] Among them, the virus characteristics specifically include: epitope characteristics, physicochemical property characteristics, N-glycosylation sites, and receptor binding region characteristics.
[0066] It should be noted that based on the HA1 amino acid sequence data of the B-type influenza virus Victoria lineage and Yamagata lineage virus samples obtained in the previous step, the virus characteristics of the two lineages are extracted from the sample data of the two lineages respectively. The virus characteristics mentioned in this embodiment generally refer to the characteristics that can affect the antigenicity of the influenza virus HA protein, specifically including: epitope characteristics, physicochemical property characteristics, N-glycosylation sites, and receptor binding region characteristics.
[0067] Among them, the extraction method of epitope characteristics includes:
[0068] 1) According to the preset amino acid site screening conditions, amino acid sites that meet the amino acid site screening conditions are respectively determined from the HA1 protein sequences of B-type influenza virus Victoria lineage and Yamagata lineage virus samples as the antigen sites of each lineage;
[0069] It should be noted that the screening targets corresponding to the amino acid site screening conditions in this embodiment are: a) Select surface-exposed amino acid residues; b) Select potential antigen sites predicted using ScanNet; c) Select amino acid residues located in the head domain of the HA protein.
[0070] More specifically, for screening target (a), in this embodiment, according to the crystal structures of B / Victoria (B / Brisbane / 60 / 2008, pdb 4FQM) and B / Yamagata (B / Yamanashi / 166 / 1998, pdb 4M40), the FreeSASA software is used to calculate the relative solvent accessibility (RSA) of each residue to determine whether the residue is exposed. Residues with an RSA greater than threshold A (exposed sites) are selected as preliminary candidate antigen sites, and the reference value of threshold A is 15%.
[0071] For screening target (b), in this embodiment, the ScanNet prediction algorithm is used to predict the probability that each residue is an antigen site. According to the output probability results, the top N residues with the highest probability are selected as potential antigen sites, and the reference value of N is 150.
[0072] For screening target (c), in this embodiment, the head domain of the HA protein of the virus sample (amino acid sites 52 - 281 for B / Victoria and 52 - 279 for B / Yamagata) is considered the main antigen region of influenza, so residues located in the head domain are spatially marked.
[0073] In summary, when the amino acid residue at a certain site simultaneously satisfies that the RSA is greater than threshold A, the probability result output by the ScanNet algorithm ranks among the top 150 with the highest probability, and the site is located within the head domain of the corresponding lineage, it indicates that the amino acid residue corresponding to this site can be used as an antigen site. Then, according to the same screening criteria, the union of the amino acid sites of chains A, C, and E that meet the above three conditions is used as the finally determined antigen site.
[0074] 2) Cluster the antigen sites and divide the antigen sites into five antigenic epitopes;
[0075] It should be noted that in step e), the K-Means method is used to cluster the antigen sites in situ to obtain the antigenic epitope regions. According to the spatial positions of the residues in Chain A, the above-determined antigen sites are divided into five epitopes by K-means. The epitopes of B / Victoria and B / Yamagata are spatially corresponded by TM-align. The finally determined epitopes are shown in Table 1.
[0076] Table 1 Distribution of antigenic epitopes and antigen sites of Victoria lineage and Yamagata lineage of influenza B virus
[0077]
[0078] 3) By using the PSI-BLAST algorithm, a PSSM matrix (Position-Specific Scoring Matrix) containing each antigen site is constructed.
[0079] 4) Taking the Victoria lineage as an example, according to a selected pair of sample data of the Victoria lineage of influenza B virus, the sum of the absolute values of the differences in the PSSM scores of the two amino acid residues corresponding to each site position on the antigenic epitope is used as the characteristic value of the antigenic epitope. When the characteristic values of all antigenic epitopes are obtained, the antigenic epitope characteristic values are output. The method for calculating the antigenic epitope characteristics of the Yamagata lineage is the same.
[0080] It should be noted that after the definition of the antigenic epitope is completed, the calculation of the characteristic values related to the epitope is carried out next. According to the HA1 sequences of the Victoria lineage and the Yamagata lineage, PSI-BLAST is used to generate the 346×20 position-specific scoring matrices (PSSMs) of the two lineages respectively to reflect the distribution of 20 amino acids at 346 sites. The PSSM matrix scores the 20 amino acids at each position according to the distribution frequency of different amino acids at each residue position.
[0081] For a pair of strains, the antigenicity score (characteristic value) of each epitope is calculated according to the sum of the absolute values of the differences in the PSSM matrix scores of the two compared residues at each residue position on the epitope. If a position involves a deletion or insertion, the difference in the PSSM matrix score at that position is defined as the maximum value of all absolute differences between the target residue and all 20 amino acids. For strains A and B, the t characteristic value of epitope e is calculated using the following formula 2:
[0082] pssm(A,B,e)=∑|m(p,e,A p )-m(p,e,B p )| (2)
[0083] where m(p, e, A p ) represents the amino acid A at position p belonging to epitope e p 's PSSM matrix score. m(p, e, B p ) represents the amino acid B at position p belonging to epitope e p 's PSSM matrix score.
[0084] The extraction methods for the physicochemical property eigenvalues include:
[0085] 1) According to the amino acid index database, determine the physicochemical property indicators corresponding to the Victoria lineage and the Yamagata lineage. Among them, the physicochemical property indicators include: hydrophobicity, volume, charge, polarity, and accessible surface area;
[0086] 2) Taking the Victoria lineage as an example, according to a selected pair of sample data of the Victoria lineage of influenza B virus and a selected physicochemical property indicator, calculate the average value of the differences in the corresponding physicochemical property scores at all amino acid-different sites of a pair of strains as the eigenvalue of the physicochemical property. When there are more than three different amino acid sites in a pair of strains, then take the average value of the top three sites with the largest differences in the corresponding physicochemical property scores at all amino acid-different sites as the eigenvalue of the physicochemical property. After obtaining the eigenvalues of all physicochemical property indicators, output the physicochemical property eigenvalues. The method for calculating the physicochemical properties of the Yamagata lineage is the same.
[0087] It should be noted that this application collected fifty-six amino acid indices reported to potentially affect the antigenicity change of influenza virus from the amino acid index database (AAindex). They represent five physicochemical properties, including hydrophobicity, volume, charge, polarity, and accessible surface area. Each time, one amino acid index is selected from each group of physicochemical properties to form a combination (consisting of 5 amino acid indices) for antigen relationship prediction, and the combination with the highest AUC (Area Under Curve, the area enclosed by the ROC curve and the coordinate axes) value is selected for model construction. Finally, five amino acid indices with IDs FAUJ880109, FAUJ8880103, ZIMJ680104, ZIMJ680103, CHOC760101 are used to construct the antigen prediction model for the Victoria lineage, and five amino acid indices with IDs FAUJ880109, PONJ960101, ZIMJ680104, CHAM820101, CHOC760102 are used to construct the antigen prediction model for the Yamagata lineage.
[0088] The extraction method for N-glycosylation sites is:
[0089] Predict the N-glycosylation sites of each HA1 sequence using NetNGlyc. For a pair of strains, the characteristic value of the N-glycosylation sites is the number of different N-glycosylation sites in the HA1 sequences of the two strains.
[0090] The methods for extracting the characteristics of the receptor-binding region include:
[0091] For a pair of strains, the characteristic value of the receptor-binding region is the average of the shortest Euclidean distances from the sites where the amino acids of the HA1 proteins of the two strains are different to the receptor-binding region (140-loop, 190-helix, 240-loop). The Euclidean distance between two residues is calculated based on their C-α atoms. Using the pdb structure as a template (Victoria: pdb 4FQM, Yamagata: pdb 4M40), the shortest Euclidean distance from a residue to the receptor-binding region is the shortest Euclidean distance from the residue to all residues in the receptor-binding region. If the amino acids at more than three sites change, only the shortest three Euclidean distances are considered in the calculation.
[0092] So far, the datasets for constructing the antigenicity prediction models of the Victoria lineage and Yamagata lineage of influenza B virus have been prepared, including a total of 12 characteristics and 1 target predicted by the model - the antigenic relationship between pairs of strains. The characteristics specifically include 5 epitope-related characteristics, 5 amino acid physicochemical property-related characteristics, 1 N-glycosylation-related characteristic, and 1 receptor-binding region-related characteristic.
[0093] Step 103: Use the virus characteristics and sample antigenicity data as model training data, and through model training, obtain the antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the Victoria lineage and Yamagata lineage of influenza B virus respectively.
[0094] More specifically, the training methods of the antigenicity prediction models include:
[0095] Based on the sample datasets of the Victoria lineage and Yamagata lineage of influenza B virus and the sample antigenicity data, combined with a variety of preset initial algorithm models, respectively use the HA1 sequence data and sample antigenicity data of the Victoria lineage and Yamagata lineage samples as the input quantities of each initial algorithm model, train each initial algorithm model, and perform five-fold cross-validation.
[0096] According to the obtained cross-validation results, determine the optimal initial algorithm model corresponding to each lineage for training to obtain the best antigenicity prediction models for different lineages.
[0097] It should be noted that based on the 4 groups and 12 types of features related to antigen changes constructed above, this application tested the performance of four machine learning methods in predicting the antigenic relationships defined by HI tests between strains within the Victoria lineage and the Yamagata lineage of influenza B virus, including: XGBoost (eXtreme Gradient Boosting, an optimized distributed Gradient Boosting framework), LightGBM (Light Gradient Boosting Machine, an optimized distributed Gradient Boosting framework), RF (Random Forest), and SVM (Support Vector Machine). Among them, when constructing the model, SMOTE (Synthetic Minority Over-sampling Technique) was used to balance positive and negative samples, and then the dataset was randomly divided into a training set and a test set at a ratio of 4:1. Model evaluation was carried out through cross-validation and independent testing, and accuracy, precision, recall, AUC, and F1 score were used as evaluation indicators. The calculation formulas are as follows:
[0098]
[0099] Among them, TP represents true positive samples, TN represents true negative samples, FP represents false positive samples, and FN represents false negative samples.
[0100] Then, select the machine learning method with the best performance (in the order of AUC, accuracy, precision, recall, F1 score) in cross-validation to establish the final antigenicity models PREDAC-BV and PREDAC-BY. The performance of the independent set test of the models is shown in Table 2.
[0101] Table 2 Performance of the independent set test of the antigen prediction model
[0102]
[0103] Step 104: Obtain the HA1 amino acid sequence of the influenza B virus strain of the Victoria lineage or the Yamagata lineage to be identified, and input the HA1 amino acid sequence of the influenza B virus strain into the antigenicity prediction model corresponding to the lineage, so as to determine the antigenicity identification result of the influenza B virus strain according to the output result of the antigenicity prediction model.
[0104] In addition, regarding the constructed antigenicity prediction models PREDAC-BV and PREDAC-BY for the Victoria lineage and Yamagata lineage of influenza B virus, they have the following functions:
[0105] 1) Construct antigen relationship networks. Use PREDAC-BV and PREDAC-BY to predict the antigen relationships between all pairwise strains within the Victoria and Yamagata lineages respectively. Then, select all strain pairs predicted to be antigenically similar, and construct the antigen relationship networks ACnet-BV and ACnet-BY for the Victoria lineage and Yamagata lineage respectively through Cytoscape. The nodes of the network are strains, and the edges are antigenic similarities. The antigenic similarity is the logarithm ratio of the probability that two strains are predicted to be antigenically similar to the probability that they are not antigenically similar. This network can vividly display the antigen relationships among all influenza strains in the Victoria lineage and Yamagata lineage, facilitating the study of antigenic similarities of influenza strains at the network level.
[0106] 2) Antigen cluster division. Use MCL to cluster the strains in the above antigen relationship networks to determine different antigen clusters, where the inflation parameter of MCL is selected according to the highest modularity of the network. In this application, a total of 7 antigen clusters were identified in the Victoria lineage, named BJ87, HK01, BR08, CO17, WT19, New, and AU21 respectively. Among them, the antigen cluster New is the antigen cluster identified in this application that does not contain vaccine strains, and the remaining 6 antigen clusters all contain vaccine strains that have been used in influenza vaccines, and are named after the abbreviation of the earliest used vaccine strain they contain. It can be considered that the antigen relationships among strains within the same antigen cluster are relatively similar, while the antigen relationships among strains belonging to different antigen clusters are relatively dissimilar. Through antigen cluster division, the antigen evolution processes of the Victoria lineage and Yamagta lineage can be revealed at the network level.
[0107] 3) Analyze the antigen evolution laws of the Victoria lineage and Yamagata lineage of influenza B virus. The antigen evolution of the two lineages of influenza B virus is a process in which new antigen clusters continuously replace old antigen clusters. Combining the antigen cluster division, regional and temporal distribution of strains, the evolutionary processes of the Victoria lineage and Yamagata lineage of influenza B virus in the Northern and Southern Hemispheres, different continents, and even different countries can be traced back, and the antigen evolution maps of the two lineages can be drawn.
[0108] 4) Auxiliary vaccine strain recommendation. Based on the antigenic relationships among the strains predicted by the antigen prediction model, we can calculate the proportion of antigenically similar strain pairs between all strains in each epidemic season in the Northern and Southern Hemispheres and the previously dominant antigenic clusters. When this proportion is low, it indicates that there have been significant antigenic changes between the strains prevalent in the current epidemic season and the previously dominant antigenic clusters, suggesting that the vaccine strain should be replaced at this time to improve the protective effect of the influenza vaccine. Figure 4 The antigenic similarity ratios between the strains of the Victoria lineage in each epidemic season in the Northern and Southern Hemispheres calculated based on the above-mentioned antigen prediction model PREDAC-BV and the previously dominant antigenic clusters. By comparing with the vaccine strains recommended by the World Health Organization (WHO), antigenic changes can be monitored earlier, and the replacement of influenza vaccine strains can be prompted earlier. Since the Yamagata lineage has not been prevalent in the population since March 2020, this invention does not provide auxiliary vaccine recommendations for this lineage.
[0109] The above is a detailed description of an embodiment of a method for rapidly identifying the antigenicity of influenza B virus provided by this application. The following is a detailed description of a device for rapidly identifying the antigenicity of influenza B virus provided by this application.
[0110] Please refer to Figure 5 , this embodiment provides a device for rapidly identifying the antigenicity of influenza B virus, including:
[0111] A sample data acquisition unit 201, configured to respectively acquire the HA1 sequence sample data and sample antigenicity data of the preset influenza B virus Victoria lineage and Yamagata lineage, wherein the sample antigenicity data respectively includes the antigenic relationships between pairwise strains within the Victoria lineage and the Yamagata lineage;
[0112] An antigen feature extraction unit 202, configured to respectively perform virus feature extraction on the HA1 sequence sample data of the influenza B virus Victoria lineage and Yamagata lineage, and respectively obtain the virus features related to antigenic changes corresponding to the sample data sets of the Victoria lineage and the Yamagata lineage. The virus features specifically include: antigenic epitope features, physicochemical property features, N-glycosylation sites, and receptor binding region features;
[0113] An antigenicity prediction model training unit 203, configured to use the virus features and sample antigenicity data as model training data, and respectively obtain the antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the influenza B virus Victoria lineage and Yamagata lineage through model training;
[0114] The antigenicity discrimination unit 204 is used to obtain the HA1 amino acid sequence of the influenza B virus influenza Victoria lineage or Yamagata lineage strain to be discriminated, and input the HA1 amino acid sequence of the influenza B virus strain into the antigenicity prediction model corresponding to the lineage, so as to determine the antigenicity discrimination result of the influenza B virus strain according to the output result of the antigenicity prediction model.
[0115] In addition to the above detailed description of the embodiments of the rapid antigenicity discrimination device for influenza B virus, the present application also provides a detailed description of an embodiment of a rapid antigenicity discrimination terminal for influenza B virus and a computer-readable storage medium.
[0116] Please refer to Figure 6 , a rapid antigenicity discrimination terminal for influenza B virus provided in this embodiment, the types of the terminal include but are not limited to: personal computers, industrial computers, servers, and embedded intelligent devices. The main components of the terminal include: a memory 33 and a processor 31, and the memory 33 and the processor 31 can be connected through a communication bus 34;
[0117] The memory 33 is used to store program codes corresponding to the rapid antigenicity discrimination of influenza B virus provided in the above embodiments;
[0118] The processor 31 is used to read and execute the program codes.
[0119] A computer-readable storage medium provided in this embodiment stores program codes corresponding to the rapid antigenicity discrimination of influenza B virus in the above embodiments.
[0120] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described terminal, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0121] In several embodiments provided by the present application, it should be understood that the disclosed terminal, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0122] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0123] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0124] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0125] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0126] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for rapid identification of influenza B virus antigenicity, characterized in that: include: Obtaining preset HA1 sequence sample data and sample antigenicity data of the Victoria lineage and the Yamagata lineage of influenza B virus, respectively, wherein the sample antigenicity data respectively include antigenic relationships between strains within the Victoria lineage and the Yamagata lineage; Extracting virus features from the Victoria lineage and Yamagata lineage HA1 sequence sample data of influenza B virus, respectively, and obtaining virus features related to antigenic changes corresponding to the Victoria lineage and Yamagata lineage sample data sets, respectively, wherein the virus features specifically include: antigen epitope feature values, physicochemical property feature values, nitrogen glycosylation sites, and receptor binding region features; Using the virus characteristics and the sample antigenicity data as model training data, through model training, antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to Victoria lineage and Yamagata lineage of influenza B virus are obtained respectively; Obtaining an HA1 amino acid sequence of an influenza B virus strain of the Victoria lineage or the Yamagata lineage to be identified, and inputting the HA1 amino acid sequence of the influenza B virus strain into an antigenicity prediction model of the corresponding lineage, so as to determine an antigenicity identification result of the influenza B virus strain according to an output result of the antigenicity prediction model; The method for extracting the antigen epitope characteristic value specifically includes: According to the preset amino acid site screening conditions, the amino acid sites that meet the amino acid site screening conditions are determined from the HA1 protein sequences of the influenza B Victoria lineage and the Yamagata lineage virus samples as the antigenic sites of the two lineages respectively, wherein the amino acid site screening conditions include: selecting surface exposed amino acid residues; selecting potential antigenic sites predicted by ScanNet; selecting amino acid residues located in the head domain of the HA protein; Clustering the antigenic sites to divide the antigenic sites into five antigenic epitopes; Using the PSI-BLAST algorithm, a PSSM matrix containing each of the antigenic sites is constructed; According to a pair of influenza B virus sample data of the same lineage, the sum of the absolute values of the differences between the PSSM scores of the two amino acid residues corresponding to each site on the antigenic epitope is taken as the characteristic value of the antigenic epitope, and when the characteristic values of all antigenic epitopes are obtained, the characteristic values of the antigenic epitope are output; The method for extracting the characteristic value of the physical and chemical properties specifically includes: According to the amino acid index database, the physical and chemical property indicators corresponding to the Victoria lineage and the Yamagata lineage are determined, wherein the physical and chemical property indicators include: hydrophobicity, volume, charge, polarity and accessible surface area; According to a pair of influenza B virus sample data of the same lineage and a selected physicochemical property index, the average value of the corresponding physicochemical property score difference at all amino acid sites of the pair of influenza B viruses is calculated as the characteristic value of the physicochemical property. When a pair of influenza B viruses has more than three different amino acid sites, the average value of the first three sites with the largest corresponding physicochemical property score difference at all amino acid sites is taken as the characteristic value of the physicochemical property. After the characteristic values of all physicochemical property indexes are obtained, the characteristic value of the physicochemical property is output.
2. A method for rapid identification of influenza B virus antigenicity according to claim 1, characterized in that: The method for extracting the receptor binding region features specifically includes: According to a pair of selected influenza B virus sample data of the same lineage, the average value of the shortest Euclidean distance from the sites with different amino acids on the HA1 protein sequences of the pair of influenza B viruses to the receptor binding region characteristics is calculated as the receptor binding region characteristics.
3. A method for rapid identification of influenza B virus antigenicity according to claim 1, characterized in that: The training method of the antigenicity prediction model includes: Based on the influenza B Victoria lineage and Yamagata lineage virus sample HA1 sequence data and the sample antigenicity data, combined with a plurality of preset initial algorithm models, the Victoria lineage and Yamagata lineage sample HA1 sequence data and the sample antigenicity data are used as inputs of each initial algorithm model, each initial algorithm model is trained, and a five-fold cross validation is performed; Based on the cross-validation results obtained, the optimal initial algorithm model corresponding to each lineage is determined to obtain the best antigenicity prediction model for different lineages.
4. A method for rapid identification of influenza B virus antigenicity according to claim 3, characterized in that: The types of initial algorithm models include: XGBoost, LightGBM, RF and SVM.
5. A method for rapid identification of influenza B virus antigenicity according to claim 1, characterized in that: After determining the antigenicity identification result of the influenza B virus strain, the method further comprises: Based on the constructed Victoria lineage antigenicity prediction model, the antigenic relationship between all Victoria lineage strains of known sequences is predicted, and an antigenic similarity network between strains is constructed; the Victoria lineage strain antigenic similarity network is clustered using the Markov clustering method to identify different antigenic classes; the vaccine strain recommendation is assisted by analyzing the antigenic similarity between the Victoria lineage strains prevalent in the current epidemic season and the Victoria lineage antigen clusters that dominated the epidemic in the past. When the antigenic similarity is low, it is recommended to replace the Victoria lineage vaccine strain, otherwise the vaccine strain should not be replaced.
6. A device for rapid identification of influenza B virus antigenicity, characterized in that: include: A sample data acquisition unit, used to respectively acquire HA1 sequence sample data and sample antigenicity data of the preset Victoria lineage and Yamagata lineage of influenza B virus, wherein the sample antigenicity data respectively include the antigenic relationship between strains within the Victoria lineage and the Yamagata lineage; An antigen feature extraction unit is used to extract virus features from the Victoria lineage and Yamagata lineage HA1 sequence sample data of the influenza B virus, respectively, to obtain virus features related to antigenic changes corresponding to the Victoria lineage and Yamagata lineage sample data sets, respectively, wherein the virus features specifically include: antigen epitope feature values, physical and chemical property feature values, nitrogen glycosylation sites, and receptor binding region features; An antigenicity prediction model training unit, used to obtain antigenicity prediction models PREDAC-BV and PREDAC-BY corresponding to the Victoria lineage and the Yamagata lineage of influenza B virus respectively through model training based on the virus characteristics and the sample antigenicity data as model training data; An antigenicity identification unit, used to obtain the HA1 amino acid sequence of the influenza B virus Victoria lineage or Yamagata lineage strain to be identified, input the HA1 amino acid sequence of the influenza B virus strain into the antigenicity prediction model of the corresponding lineage, and determine the antigenicity identification result of the influenza B virus strain according to the output result of the antigenicity prediction model; The method for extracting the antigen epitope characteristic value specifically includes: According to the preset amino acid site screening conditions, the amino acid sites that meet the amino acid site screening conditions are determined from the HA1 protein sequences of the influenza B Victoria lineage and the Yamagata lineage virus samples as the antigenic sites of the two lineages respectively, wherein the amino acid site screening conditions include: selecting surface exposed amino acid residues; selecting potential antigenic sites predicted by ScanNet; selecting amino acid residues located in the head domain of the HA protein; Clustering the antigenic sites to divide the antigenic sites into five antigenic epitopes; Using the PSI-BLAST algorithm, a PSSM matrix containing each of the antigenic sites is constructed; According to a pair of influenza B virus sample data of the same lineage, the sum of the absolute values of the differences between the PSSM scores of the two amino acid residues corresponding to each site on the antigenic epitope is taken as the characteristic value of the antigenic epitope, and when the characteristic values of all antigenic epitopes are obtained, the characteristic values of the antigenic epitope are output; The method for extracting the characteristic value of the physical and chemical properties specifically includes: According to the amino acid index database, the physical and chemical property indicators corresponding to the Victoria lineage and the Yamagata lineage are determined, wherein the physical and chemical property indicators include: hydrophobicity, volume, charge, polarity and accessible surface area; According to a pair of influenza B virus sample data of the same lineage and a selected physicochemical property index, the average value of the corresponding physicochemical property score difference at all amino acid sites of the pair of influenza B viruses is calculated as the characteristic value of the physicochemical property. When a pair of influenza B viruses has more than three different amino acid sites, the average value of the first three sites with the largest corresponding physicochemical property score difference at all amino acid sites is taken as the characteristic value of the physicochemical property. After the characteristic values of all physicochemical property indexes are obtained, the characteristic value of the physicochemical property is output.
7. A terminal for rapid antigenicity identification of influenza B virus, characterized in that: include: Memory and processor; The memory is used to store a program code corresponding to a rapid antigenic identification of a type B influenza virus according to any one of claims 1 to 5; The processor is used for reading and executing the program code.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes corresponding to the rapid antigenic identification of type B influenza virus according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for rapidly identifying antigenicity of H9N2 subtype avian influenza strain
CN117912568A