A method for predicting future mutations of influenza virus

By defining and predicting the surface antigen region of the influenza virus capsid protein HA, the problem of predicting influenza virus mutations has been solved, enabling early prediction of future influenza virus mutations and improving the specificity and protective efficacy of vaccines.

CN117352053BActive Publication Date: 2026-02-13FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210735837.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2026-02-13
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

Due to the rapid mutation of the influenza virus, existing vaccines are ineffective because the vaccine components do not match the circulating strains. They cannot effectively predict future mutations, thus affecting the protective efficacy of the vaccines.

Method used

By constructing a method for delineating the surface antigen region of the influenza virus capsid protein HA, and utilizing multiple sequence alignment, phylogenetic tree and immunogenic region delineation, combined with mutation frequency and three-dimensional spatial analysis, we can predict future influenza virus mutations, generate potential mutation spectra, and screen out high-probability mutation sequences.

Benefits of technology

It can capture more future mutations from fewer candidate sequences, predict dominant mutations that are prevalent several flu seasons in advance, improve the specificity and protective efficacy of vaccines, and save vaccine development time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117352053B_ABST
    Figure CN117352053B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biology, in particular to a method for predicting influenza virus mutation, further relates to a method for demarcating the HA surface antigen region of influenza virus coat protein and the role in predicting influenza virus mutation. The method for predicting influenza virus mutation comprises: immunogenic region demarcation, prediction template selection, simulation mutation, screening and sequencing. The present application provides the feasibility of predicting the future mutation of influenza, and proposes a practical model to solve this problem. The present application can help to develop preventive therapeutic drugs in advance to better cope with the threat of influenza.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular to a method for predicting influenza virus mutation, further relates to a method for demarcating the HA surface antigen region of influenza virus coat protein and its role in predicting influenza virus mutation. BACKGROUND

[0002] Seasonal influenza virus infects 5-15% of the global population each year and causes more than 500,000 deaths. As a typical RNA virus, seasonal influenza virus constantly mutates on the surface coat protein, such as hemagglutinin (HA) antigen, to evade the immunity induced by previous infection or vaccination. Due to the important influence of HA protein in host-virus interaction, it has been used as the main component of various vaccines.

[0003] Each year, the World Health Organization (WHO) assesses the protection ability of the existing vaccine, and if the protection efficacy is significantly decreased, a new vaccine component will be recommended. After the new vaccine component is determined, it takes at least several months to produce, package and sell the vaccine. This obvious time lag may cause the mismatch between the HA vaccine component inoculated several months later and the prevalent strain, resulting in occasional vaccine failure.

[0004] For example, during 2004-2005, the effectiveness of the vaccine recommended by WHO was only 10% (95% confidence interval [CI]). The reason is that the main community prevalent strain A / California / 7 / 2004-like (H3N2) is inconsistent in antigen with the WHO recommended vaccine strain A / Fujian / 411 / 2002-like (H3N2), which is no longer prevalent in the next influenza season. Similarly, the 2014-2015 influenza vaccine also did not provide effective protection against the dominant A / H3N2 influenza virus, because the vaccine A / Texas / 50 / 2012-like (H3N2) had disappeared in the population.

[0005] In fact, the time lag from vaccine screening to population immunization not only leads to vaccine failure, but also leads to vaccine effectiveness far below expectations. According to the CDC's annual comprehensive assessment, the effectiveness of influenza vaccine fluctuated between 10% and 60% in the 16 years from 2004 to 2020, with a median of only 40.5%, which shows that there is still much room for improvement in this field. In other words, due to the rapid and continuous variation of the virus, this vaccine strategy based on post-strain seems to be unable to provide sufficient protection against future mutant strains. Only by vaccinating against the future variation of influenza can the effectiveness of the vaccine be effectively improved.

[0006] Ongoing global influenza surveillance has accumulated a large number of historical sequences. Many evolutionary studies focus on the tracing of influenza sequences rather than predicting future variations. Predicting the future antigen mutation spectrum of seasonal influenza helps us better develop proactive measures to prevent influenza. SUMMARY

[0007] The purpose of the present application is to provide a method for delineating the HA surface antigen region of the influenza virus coat protein and a method for predicting the future mutation of the influenza virus, which helps to develop preventive therapeutic drugs in advance to better cope with the threat of influenza.

[0008] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows.

[0009] In a first aspect, the present application provides a method for delineating the HA surface antigen region of the influenza virus coat protein.

[0010] A method for delineating the HA surface antigen region of the influenza virus coat protein, comprising the following steps:

[0011] S101 influenza virus dataset preprocessing:

[0012] Collect a set of influenza HA nucleic acid sequences D from a public database, align and extract the 3Ps-2 to 3Pe positions in the HA nucleotide sequence corresponding to the HA1 region using a multiple sequence alignment tool, obtain a set of HA1 nucleic acid sequences S, and translate a set of protein sequences P according to the set of HA1 nucleic acid sequences S;

[0013] Wherein, HA1 is a domain fragment of HA protein, Ps is the starting position of HA1 protein reference sequence, Pe is the termination position of HA1 protein reference sequence, the starting position Ps and the termination position Pe of HA reference sequence HA1 of different subtypes of influenza are different (search for the specific sites of Ps and Pe of HA1 protein of different subtypes of influenza in uniprot database with the keyword influenza HA1 protein);

[0014] S102 sequence phylogenetic tree construction:

[0015] Based on the principle of space-time conversion, a representative sequence set T is formed to construct a phylogenetic tree and obtain a parent-child node sequence set PT;

[0016] S103 prediction of immunogenic region division:

[0017] According to the formula MF i = M i / A, the mutation frequency of each site in the set of HA1 protein sequences P is calculated to obtain a set of high-frequency mutation sites F; in the formula, MF i represents the mutation frequency of the i-th site in the HA1 protein sequence; Mi The number of sequences at site i where the mutation occurs; A represents the total number of HA1 sequences.

[0018] The set containing all high-frequency mutation sites is defined as the high-frequency mutation site set F;

[0019] Preferably, MF i Sites with a value >0.05 are defined as high-frequency mutation sites;

[0020] Preferably, MF i Sites with a value >0.1 are defined as high-frequency mutation sites;

[0021] Potential immunogenic sites for experimental validation were collected from literature and databases to obtain experimental site set A;

[0022] Define F∪A as the set I of potential immunogenic sites for HA1. Calculate the distance matrix D between each site in the set I on the three-dimensional protein structure of HA1. dist The site distance matrix D dist The calculation formula is D dist ={d ij |i,j∈I,d ij ∈{minimum / average / central carbon atom distance between amino acid i and amino acid j}};

[0023] According to the clustering formula Cluster={c i c j |D min (c i c j ) <N C c i c j ∈D dist , i, j∈I}, where N C The values ​​range from 4 to 10 angstroms. In three-dimensional space, the site set I is clustered into n predicted immunogenicity regions C = {C1, C2, ..., C} according to distance. n}, and make M the number of amino acid sites in each predicted immunogenicity region. a Meets the condition 5≤M a ≤18.

[0024] Preferably, in step S101: the influenza HA nucleic acid sequence set D can be represented in the following form: D = {d} i |i=1~N d}, where N d d represents the total number of HA records collected. i The nucleic acid sequence representing the i-th HA;

[0025] The HA1 nucleic acid sequence set S can be expressed in the following form: wherein N s is Ps x 3-2, N n is Pe x 3, s i represents the i-th nucleic acid sequence;

[0026] The translated protein sequence set P can be expressed in the following form: p i represents the i-th protein sequence.

[0027] Preferably, in the step S102, the time-space conversion principle is that: the time is to select N m (5-100) nucleic acid sequences respectively in each influenza season; and the space is that the N m nucleic acid sequences are respectively from different regions and / or countries.

[0028] Preferably, the representative sequence set T can be expressed in the following form: t i represents the i-th nucleic acid sequence; and the parent-child node sequence set PT can be expressed in the following form: wherein N p is the total number of the representative sequence set T, and pt i represents the i-th pair of nucleic acid sequences.

[0029] Preferably, in the step S103, the D dist is the minimum atomic distance, which is calculated according to the formula D dist (a i , a j )(a i , a j ∈{atomic coordinates}) to calculate the minimum atomic distance D dist between each site in the site set I.

[0030] Preferably, in the step S103, the D dist is the central carbon atom distance, which is calculated according to the formula D dist (a i , a j )(a i , a j ∈{central carbon atom coordinates}) to calculate the central carbon atom distance D dist between each site in the site set I.

[0031] Preferably, in the step S103, the D distThe average atomic distance is calculated according to formula D. dist =AVE{d(a i a j )}(a i a j ∈{atomic coordinates}), calculate the average atomic distance D between each point in the site set I. dist .

[0032] Preferably, in step S103, the C region located in the predicted immunogenicity region... i and C j For the edge point x, according to the principle of minimization, x∈D dist (C i C j (C) i C j ∈{potential immunogenicity region set C}), and divide site x into a predicted immunogenicity region containing fewer sites.

[0033] Secondly, the present invention provides a method for predicting future mutations of influenza viruses.

[0034] A method for predicting future mutations of influenza viruses, comprising the following steps:

[0035] S201 Immunogenic Region Division: According to the method of the first aspect of the present invention, the set of high-frequency mutation sites and potential immunogenic sites in the influenza HA1 protein is divided into n predicted immunogenic regions.

[0036] S202 Prediction Template Selection:

[0037] The abundance of each sequence in the HA1 protein sequence set P or HA1 nucleic acid sequence set S is calculated using the formula ∑Ik(x=y)(I, k∈protein sequence set P or nucleic acid sequence set S) to generate an abundance descending set B; the pre-defined predicted immunogenicity regions are then mapped according to the mapping relationship F. m Mapped onto the HA1 sequence; the top N abundances in descending order set B are selected. e Using the strip as a template, a template set E is generated for each potentially immunogenic region;

[0038] S203 simulated mutation:

[0039] Calculate the nucleotide mutation frequency and generate a nucleotide mutation probability matrix Mn, or generate an amino acid mutation probability matrix Mp based on the amino acid mutation frequency.

[0040] The nucleotide mutation probability matrix Mn can be represented as follows: Mn={N ab |a, b∈{A,T,C,G},N ab{probability of nucleotide a mutating into nucleotide b}; the amino acid mutation probability matrix Mp can be expressed as: Mp = {M ab | a, b e {20 types of amino acids}, M ab {probability of amino acid a mutating into amino acid b} can be obtained.

[0041] wherein M ab is calculated as follows: first, the frequency of the triplets encoding the amino acid Res is counted according to the formula i A weight matrix Wtr i = {W1, W2, …, W Ntri} of the triplets is obtained; in the formula, Otr i represents the number of times that the triplet i encoding the amino acid Res i appears in the nucleic acid sequence set S, Ntr i represents the total number of triplets encoding the amino acid Res i , represents the total number of times that all triplets encoding the amino acid Res i appear in S; then the value of M i is obtained according to the accumulation formula P(Res j → Res tri ) = W trj × W i ×∑(P(nuc j → nuc k )×P(nuc l → nuc m )×P(nuc n → nuc ab )); wherein R1, R2 e {20 types of amino acids}, nuc i , nuc j , nuc k , nuc l , nuc m , nuc n e {A, T, C, G}, the sequence (nuc i , nuc k , nuc m ) represents a triplet encoding Res i , and the sequence (nuc j , nuc l , nuc n ) represents a triplet encoding Res j .

[0042] When the template is a nucleic acid sequence, the mutation simulation is performed according to the formula R=Mn x E, and finally a potential mutation spectrum R is obtained, which can be expressed in the following form: wherein V is the total number of mutated sequences.

[0043] When the template is a protein sequence, the mutation simulation is performed according to the formula R=Mp x E, and finally a potential mutation spectrum R is obtained, which can be expressed in the following form: wherein V is the total number of mutated sequences.

[0044] S204 screening and sorting:

[0045] According to the simulated host immune barrier formula: The difficulty of each subsequence in the potential mutation spectrum to break through the immune barrier is measured, and a potential mutation spectrum Q is obtained after host immune screening; in the formula, DSIB is the difficulty of breaking through the immune barrier, N i is the total number of amino acid sites in all antigen regions, F i is whether the mutated amino acid has appeared at position i in history, and 0 represents appearance and 1 represents non-appearance;

[0046] The abundance of each subsequence in the potential mutation spectrum Q is calculated according to the formula ∑Ik(x=y) (I, k∈potential mutation spectrum Q), and is sorted in descending order to generate a potential mutation spectrum U; or the frequency of each mutation in the historical mutation spectrum is calculated according to the formula ∑mut(mut∈potential mutation spectrum Q), and is sorted in descending order to generate a potential mutation spectrum U, and the first N w sequences are selected to construct the final predicted mutation spectrum W.

[0047] In a specific embodiment, in the step S202:

[0048] In the formula ∑IK(x=y), when x=y, it is 1, otherwise it is 0;

[0049] The abundance descending set B can be expressed in the following form: or wherein N b represents the number of protein sequence set P or nucleic acid sequence set S after deduplication; wherein the abundance of each sequence in the HA1 protein sequence set P or HA1 nucleic acid sequence set S is sorted in descending order;

[0050] The mapping relationship F m can be expressed in the following form: F m ={a i →bi , a i ∈ potential immunogenic site set I, b i ∈ abundance descending set B} are mapped to the HA1 sequence, wherein, F m represents a potential immunogenic region m;

[0051] The template set E generated for each potential immunogenic region is expressed in the following form: Wherein, the number of templates N e is in the range of 1-10.

[0052] In a specific embodiment, in step S204, the measurement standard of the difficulty of each subsequence sequence in the potential mutation spectrum to break through the immune barrier is that when the DSIB of the subsequence sequence is greater than I (0.05-0.8), it is removed, and the potential mutation spectrum Q is obtained after host immune screening;

[0053] The calculation formula of the predicted mutation spectrum W is

[0054] The value of N w is in the range of 2≤N w ≤500.

[0055] In view of the high mutation rate of the antigen region, it is very technically challenging to predict the upcoming mutations for influenza. The method for delineating the new antigen region of the influenza virus coat protein HA surface provided by the present application and the method for predicting future mutations of the influenza virus can re-delineate the antigen region, obtain multiple templates from the epidemic strain for each influenza season, and capture more future dominant mutations in fewer candidate sequences.

[0056] More importantly, the method of the present application can predict the upcoming new dominant mutation that will be prevalent several influenza seasons in advance.

[0057] The method provided by the present application can prevent the threat of influenza according to the currently prevalent strains of different subtypes.

[0058] In the method for predicting future mutations of the influenza virus provided by the present application, the screened mutations are sorted according to the simulated abundance, reflecting the potential possibility of prevalence in the community. This strategy can be extended to other viruses.

[0059] In summary, the present application provides a feasible method for predicting future mutations of influenza and proposes a practical model to solve this difficult problem. With antigenicity evaluation, the present application can help to develop preventive therapeutic drugs or vaccines in advance to better cope with the threat of influenza.

[0060] The method for demarcating the surface new antigen region of the influenza virus coat protein HA provided by the application and the method for predicting future mutations of the influenza virus have the following advantages:

[0061] (1) More future mutations can be captured in fewer candidate sequences.

[0062] (2) Potential dominant mutations can be ranked in the front.

[0063] (3) Dominant mutations in different antigen regions in the future can be predicted years in advance, thereby saving time for developing targeted influenza prevention measures. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 An H1N1_development tree provided for Embodiment 1 of the application.

[0065] Figure 2 An H3N2_development tree provided for Embodiment 2 of the application.

[0066] Figure 3 H1N1 prediction performance (strain coverage, actually observed mutations, and predicted mutations) results provided for Embodiment 3 of the application.

[0067] In the figure, the dotted line represents the strain coverage of each antigen region, the column chart represents the number of actually observed mutations in each antigen region, and the gray part represents the number of predicted mutations. Different letters represent different antigen regions, and A to G represent antigen regions A to G, respectively.

[0068] Figure 4 H3N2 prediction performance (strain coverage, actually observed mutations, and predicted mutations) results provided for Embodiment 3 of the application.

[0069] In the figure, the dotted line represents the strain coverage of each antigen region, the column chart represents the number of actually observed mutations in each antigen region, and the gray part represents the number of predicted mutations. Different letters represent different antigen regions, and A to G represent antigen regions A to G, respectively. DETAILED DESCRIPTION

[0070] In order for those skilled in the art to better understand the technical solutions of the application, the following embodiments further describe the application in detail. The following embodiments are only used to illustrate the application, but not to limit the scope of the application.

[0071] Embodiment 1

[0072] A method for demarcating an H1N1 influenza virus coat protein HA surface antigen region, comprising the following steps:

[0073] Collect 22919 HA nucleic acid sequences of H1N1 influenza from public database (collected from NCBI Influenza Virus Resources database), form H1N1_HA nucleic acid sequence set D = {AACATCCA...ATAAGGAATAAA; AACCTCTA...ATAAGGACCAAA;...; AACATCTT...ATAAGGACCAAT} (data source: NCBI Influenza Virus Resources database) ;

[0074] Submit H1N1_HA nucleic acid sequence set D to Cluster omega for multiple sequence alignment, form H1N1_HA1 nucleic acid sequence set S = {AACATCCA...ATAAGGAATAAA; AACCTCTA...ATAAGGACCAAA;...; AACATCTT...ATAAGGACCAAT} (data source: NCBI Influenza Virus Resources database) ;

[0075] Translate H1N1_HA1 nucleic acid sequence set S to form H1N1_HA1 protein sequence set P = {EDLPGNDNSTAT...NMRNVPEKQT; QDLPGNDNSTAT...GMRNVPEKQT;... QDLPGNDNSAAT...AMRNVPEKAT} (data source: NCBI Influenza Virus Resources database) ;

[0076] Based on the principle of space-time conversion, time is to select N m (5-100) nucleic acid sequences in each influenza season, respectively, and space is that the N m nucleic acid sequences are from different regions and / or countries; form H1N1 representative sequence set T = {AACATCCA...ATAAGGAATAAA; AACCTCTA...ATAAGGACCAAA;...; AACATCTT...ATAAGGACCAAT} (H1N1 representative sequence set T is shown in Table 1) ; In this embodiment, N m is 20, that is, 20 sequences are selected each year, and a total of 200 sequences are collected in 10 years (1998-2008) ;

[0077] Table 1 H1N1 representative sequence set T

[0078]

[0079] The phylogenetic tree was constructed, and the H1N1_HA1 parent-child node sequence set PT = AACATCCA...ATAAGGAATAAA: AACCTCTA...ATAAGGACCAAA; AACCTCTA...ATAAGGACCAAA: AACATCTT...ATAAGGACCAAT;...; AACATCTT...ATAAGGACCAAT: AACATCTT...ATAAGGACCAAT} was obtained (the results are shown in Figure 1). Figure 1

[0080] According to the formula MF i = M i The mutation frequency of each site in the H1N1_HA1 protein sequence set P was calculated, wherein MF i represents the mutation frequency of the i-th site in the H1N1_HA1 protein sequence; M i represents the number of sequences with mutations at site i; A represents the total number of H1N1 influenza HA1 nucleic acid sequences, which is 22919 here. Sites with MFi>0.1 were defined as high-frequency mutation sites;

[0081] The set containing all high-frequency mutation sites was defined as the high-frequency mutation site set F, and the high-frequency mutation site set F of the H1N1 influenza HA1 protein sequence was {35, 36, 43, 54, 69, 71, 82, 84, 85, 120, 129, 130, 133, 139, 146, 153, 183, 185, 186, 209, 216, 222, 252, 253, 260, 267, 271, 273, 274, 283, 295, 306, 308, 310}.

[0082] ​According to Predicting the Mutating Distribution at Antigenic Sites of the Influenza Virus (Scientific Reports, 2016 Feb 3|6:20239|DOI:10.1038 / srep20239), the experimentally verified immunogenic sites are collected to form an experimental site set A of the H1N1 type influenza HA1 protein sequence, A = {73, 74, 75, 76, 77, 78, 127, 128, 140, 141, 142, 143, 144, 145, 156, 157, 158, 159, 160, 162, 163, 164, 165, 166, 167, 169, 170, 171, 172, 173, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 206, 207, 208, 224, 225, 238, 239, 240},

[0083] The H1N1_HA1 potential immunogenic site set I is defined as F U A = {35, 36, 43, 54, 69, 71, 73, 74, 75, 76, 77, 82, 84, 85, 120, 127, 128, 129, 130, 133, 139, 140, 141, 142, 143, 144, 145, 146, 153, 156, 157, 158, 159, 160, 162, 163, 164, 165, 166, 167, 169, 170, 171, 172, 183, 185, 186, 187, 188, 189, 190, 191, 193, 194, 195, 196, 198, 206, 207, 208, 209, 216, 222, 224, 225, 238, 239, 252, 253, 260, 267, 271, 273, 274, 283, 295, 306, 308, 310}.

[0084] The minimum distance between the amino acid residues of each site in the H1N1_HA1 potential immunogenic site set I is calculated based on the three-dimensional spatial protein structure of HA1 (H1N1:3m6s.pdb) to form a matrix D dist . For example: the minimum atomic distance between site 35 and site 36 is 3.8, the minimum atomic distance between site 36 and site 85 is 28.7, and the minimum atomic distance between site 267 and site 283 is 8.9;

[0085] According to the clustering formula Cluster = {c i , c j | D min(c i , c j )<N c , c i , c j ∈D dist , i,j∈I}, cluster the potential immunogenicity site set I of H1N1_HA1. Wherein, N c can be 4-10 angstroms. When N c = 4 angstroms, in three-dimensional space, the site set I can be gathered into 7 predicted immunogenicity regions C = {C1, C2,..., C7} according to distance, and the number of amino acids M in each predicted immunogenicity region is 7-14;

[0086] For the sites at the edge of the predicted immunogenicity region, according to the principle of tending to small i∈D dist (C i , C j )(C i , C j ∈{potential immunogenicity region set C}), it is classified into the region containing fewer sites. For H1N1 type influenza HA1 protein, sites 153, 186, 187, 190 are at the edge of AR D and AR E, and since AR D contains fewer sites, sites 153, 186, 187, 190 are divided into AR D.

[0087] H1N1_HA1 is divided into the following 7 predicted immunogenicity regions:

[0088] AR A(35, 36, 85, 267, 283, 295, 306, 308, 310),

[0089] AR B(73, 74, 75, 76, 77, 169, 170, 171, 172, 260),

[0090] AR C(43, 54, 69, 71, 82, 84, 271, 273, 274),

[0091] AR D(127, 130, 153, 183, 186, 187, 190, 191, 216, 222, 224, 225),

[0092] AR E(156, 157, 158, 159, 160, 185, 188, 189, 193, 194, 195, 196),

[0093] AR F(162, 163, 164, 165, 166, 167, 198, 206, 207, 208, 209, 238, 239)

[0094] ARG (120, 128, 129, 133, 139, 140, 141, 142, 143, 144, 145, 146, 252, 253).

[0095] Example 2

[0096] A method for demarcating a region of HA surface antigen of H3N2 influenza virus coat protein, comprising the following steps:

[0097] Collecting 28183 HA nucleic acid sequences of H3N2 influenza from public database (collected from NCBI Influenza Virus Resources database) to form H3N2_HA nucleic acid sequence set D={AACATCCA…ATAAGGAATAAA; AACCTCTA…ATAAGGACCAAA; ……; AACATCTT…ATAAGGACCAAT} (data source: NCBI Influenza Virus Resources database);

[0098] Submitting H3N2_HA nucleic acid sequence set D to Cluster omega for multiple sequence alignment to form H3N2_HA1 nucleic acid sequence set S={AACATCCA…ATAAGGAATAAA; AACCTCTA…ATAAGGACCAAA; ……; AACATCTT…ATAAGGACCAAT} (data source: NCBI Influenza Virus Resources database);

[0099] Translating H3N2_HA1 nucleic acid sequence set S to form H3N2_HA1 protein sequence set P={EDLPGNDNSTAT…NMRNVPEKQT; QDLPGNDNSTAT…GMRNVPEKQT; ……QDLPGNDNSAAT…AMRNVPEKAT} (data source: NCBI Influenza Virus Resources database);

[0100] Based on the principle of space-time conversion, time is to select N m (5-100) nucleic acid sequences in each influenza season respectively; space is that these N m nucleic acid sequences are from different regions and / or countries respectively; to form H3N2 representative sequence set T{AACATCCA…ATAAGGAATAAA; AACCTCTA…ATAAGGACCAAA; ……; AACATCTT…ATAAGGACCAAT} (H3N2 representative sequence set T is shown in Table 2); in this example, Nm 20, i.e. 20 were selected each year, collected for 10 years (1998-2008), and a total of 200;

[0101] A phylogenetic tree was constructed, and the H3N2_HA1 parent-child node sequence set PT = {AACATCCA...ATAAGGAATAAA: AACCTCTA...ATAAGGACCAAA; AACCTCTA...ATAAGGACCAAA: AACATCTT...ATAAGGACCAAT;...; AACATCTT...ATAAGGACCAAT: AACATCTT...ATAAGGACCAAT} was obtained (the results are shown in Figure 1). Figure 2

[0102] According to the formula MF i = M i The mutation frequency of each site in the H3N2_HA1 protein sequence set P was calculated, where MF i represents the mutation frequency of the i-th site in the H3N2_HA1 protein sequence; M i represents the number of sequences with mutations at site i; and A represents the total number of H3N2 influenza HA1 nucleic acid sequences, which is 28183 here. Sites with MFi>0.1 were defined as high-frequency mutation sites.

[0103] Table 2 H3N2 representative sequence set T

[0104]

[0105] The set containing all high-frequency mutation sites was defined as the high-frequency mutation site set F, and the high-frequency mutation site set F of the H3N2 influenza HA1 protein sequence was {30, 31, 32, 33, 45, 48, 91, 92, 121, 128, 135, 138, 140, 142, 159, 165, 167, 171, 173, 198, 210, 212, 225, 261, 271, 273, 311, 312}.

[0106] ​According to Predicting the Mutating Distribution at Antigenic Sites of the Influenza Virus (Scientific Reports, 2016 Feb 3|6:20239|DOI: 10.1038 / srep20239), the experimentally verified immunogenic sites are collected to form an experimental site set A of the H3N2 type influenza HA1 protein sequence, A = {50, 53, 54, 62, 75, 82, 83, 122, 124, 131, 133, 137, 143, 144, 145, 146, 155, 156, 158, 160, 164, 172, 174, 188, 189, 193, 196, 197, 201, 207, 213, 217, 230, 244, 260, 262, 276, 278},

[0107] The H3N2_HA1 potential immunogenic site set I is defined as F U A = {30, 31, 32, 33, 45, 48, 50, 53, 54, 62, 75, 82, 83, 91, 92, 121, 122, 124, 128, 131, 133, 135, 137, 138, 140, 142, 143, 144, 145, 146, 155, 156, 158, 159, 160, 164, 165, 167, 171, 172, 173, 174, 188, 189, 193, 196, 197, 198, 201, 207, 210, 212, 213, 217, 225, 244, 260, 261, 262, 271, 273, 276, 278, 311, 312};

[0108] Based on the three-dimensional spatial protein structure of HA1 (H3N2: 4we8.pdb), the minimum distance between the amino acid residues of each site in the H3N2_HA1 potential immunogenic site set I is calculated to form a matrix D dist . For example, the minimum atomic distance between site 30 and site 31 is 1.3, the minimum atomic distance between site 53 and site 32 is 39.0, and the minimum atomic distance between site 121 and site 45 is 42.9;

[0109] According to the clustering formula Cluster = {c i , c j | D min (c i , c j ) < N c , c i , c j ∈ D distClustering is performed on the set I of potential immunogenic sites of H3N2_HA1, where N ∈ I, i,j∈I. c A value of 4 to 10 angstroms is acceptable. When N... c When the distance is 4 angstroms, in three-dimensional space, the site set I can be clustered into 6 predicted immunogenic regions C = {C1, C2, ..., C6} according to the distance, and the number of amino acids M in each predicted immunogenic region is 7 to 14.

[0110] For sites located at the edge of the predicted immunogenicity region, according to the principle of minimization, i∈D dist (C i C j (C) i C j ∈{potential immunogenic region set C}), and it is classified into the region containing fewer sites. For the H3N2 influenza HA1 protein, site 128 is located on the edge of ARD and ARE. Since ARE contains fewer sites, site 128 is classified into ARE.

[0111] H3N2_HA1 was divided into the following 6 potentially immunogenic regions:

[0112] AR A(30, 31, 32, 33, 45, 311, 312),

[0113] AR B (82, 83, 121, 122, 124, 171, 172, 173, 174, 260, 261, 262),

[0114] AR C (48, 50, 53, 54, 62, 91, 92, 271, 273, 276, 278),

[0115] AR D (131, 155, 156, 158, 159, 160, 188, 189, 193, 196, 197, 198, 217),

[0116] AR E(128, 165, 164, 167, 201, 207, 210, 212, 213, 244),

[0117] AR F (75, 133, 135, 137, 138, 140, 142, 143, 144, 145, 146, 225).

[0118] Example 3

[0119] Predicting future mutations of the H1N1 influenza virus involves the following steps:

[0120] Immunogenic region division: divide the HA1 protein of H1N1 influenza into 7 potential immunogenic regions according to the method described in Example 1;

[0121] Template selection: calculate the abundance of each sequence in the HA1 protein sequence set P according to the formula ∑Ik(x=y)(I, k∈protein sequence set P), and map the divided potential immunogenic region to the HA1 sequence according to the mapping relationship F m , that is, mark the position of the immunogenic region on the HA1 sequence; select the top 3 with the largest number in the data set as templates, 3 templates are generated for each potential immunogenic region, and finally generate the template set E; take AR A as an example, E={DKSIEVYKT, DKSIEIYKT, DSPTQVYRT}, and the H1N1 template set E is shown in Table 3;

[0122] Table 3 H1N1_template set E

[0123]

[0124]

[0125] Simulation mutation: calculate the mutation frequency of amino acids, generate an amino acid probability matrix Mp={S->T:0.08, S->F:0.13,…, W->E:0.02}, and the H1N1 amino acid probability matrix Mp is shown in Table 4.

[0126] Screening and sorting: according to the formula R=Mp×E, simulate the mutation, and according to the formula: measure the difficulty of each offspring sequence in the potential mutation spectrum to break through the immune barrier, and according to the formula ∑Ik(x=y)(I, k∈potential mutation spectrum Q), calculate the number of occurrences of each offspring sequence in the potential mutation spectrum and sort in descending order, select the top 90 to construct the final predicted mutation spectrum W, W={KKSIEVYKA; DESIEVCKW;…; KKGIEVYKA}, and the H1N1 predicted mutation spectrum W is shown in Table 5.

[0127] The H1N1 mutation prediction performance results are shown in Tables 6 and Figure 3 .

[0128] Table 4 H1N1_amino acid probability matrix Mp

[0129]

[0130] Table 5 H1N1_predicted mutation spectrum W

[0131]

[0132]

[0133]

[0134] Table 6 H1N1 prediction performance results

[0135]

[0136]

[0137]

[0138] In the table, a OT represents the observed mutation type; b OTT represents the observed mutation type predicted; c SC represents the strain coverage; d TC represents the type coverage; f TCN represents the type coverage of the newly emerging antigen region; g SCN represents the strain coverage of the newly emerging antigen region.

[0139] For H1N1, from 2007 to 2020, 5-63 mutations were observed during 13 influenza seasons, and the method could predict 61% of the mutation types (type coverage) on average in 7 antigen regions of H1N1, with an average strain coverage of 87%.

[0140] Example 4

[0141] Predicting future mutations of H3N2 influenza virus, comprising the following steps:

[0142] Immunogenic region division: divide the HA1 protein of H3N2 influenza into 6 potential immunogenic regions according to the method described in Example 2;

[0143] Template selection: calculate the abundance of each sequence in the H3N2_HA1 protein sequence set P according to the formula ∑Ik(x=y)(I, k∈protein sequence set P), and map the divided potential immunogenic region to the HA1 sequence according to the mapping relationship F m Mapping to the HA1 sequence, that is, annotating the site of the immunogenic region on the HA1 sequence; select the top 3 with the largest number in the data set as the template, 3 templates for each potential immunogenic region, and finally generate the template set E. Taking ARA as an example, E={TNDRNQS, TNDRNHS, TNDRNQN}; see Table 7 for the H3N2 template set E.

[0144] Table 7 H3N2_template set E

[0145]

[0146]

[0147] Simulate mutations: Calculate the mutation frequency of each amino acid, generate the amino acid probability matrix Mp = {S->T:0.08, S->F:0.13, …, W->E:0.02}, see Table 8 for H3N2 amino acid probability matrix Mp;

[0148] Screen and sort: Perform mutation simulation according to the formula R = Mp x E, and sort the simulated host immune barrier formula: Measure the difficulty of each offspring sequence in the potential mutation spectrum to break through the immune barrier, and calculate the number of occurrences of each offspring sequence in the potential mutation spectrum according to the formula ∑Ik(x=y)(I, k∈potential mutation spectrum Q) and sort in descending order, select the top 90 to construct the final predicted mutation spectrum W, W = {KKSIEVYKA; DESIEVCKW; …; KKGIEVYKA}, see Table 9 for H3N2 predicted mutation spectrum W.

[0149] Table 8 H3N2_amino acid probability matrix Mp

[0150]

[0151] Table 9 H3N2_predicted mutation spectrum W

[0152]

[0153]

[0154] H3N2 mutation prediction performance results are shown in Table 10 and Figure 4 .

[0155] Table 10 H3N2 prediction performance results

[0156]

[0157]

[0158]

[0159] In the table, a OT represents the observed mutation type; b OTT represents the observed mutation type predicted; c SC represents the strain coverage; d TC represents the type coverage; f TCN represents the type coverage of the newly emerging antigen region; g SCN represents the strain coverage of the newly emerging antigen region.

[0160] For H3N2, from 2007 to 2020, 3-67 mutations were observed in each influenza season, and the method could predict 63% of the mutation types (type coverage) and 89% of the strains on average in 6 antigen regions of H3N2.

[0161] The results of the above examples show that the method can accurately predict most of the mutations (type coverage) that occur in each influenza season, and can predict the more important (high prevalence) mutations (strain coverage) in each influenza season.

[0162] Example 5

[0163] The results of the comparison between the prediction method of future mutations of influenza viruses provided by the present application and the "mutation-selection-ranking" strategy in the prior art are as follows:

[0164] Comparison of prediction performance:

[0165] 1. The method provided by the present application can predict more mutants in fewer candidate sequences

[0166] The "mutation-selection-ranking" model provides 100 candidate sequences, and the average strain coverage from 2015 to 2019 is 87.8%, and the average type coverage is 37.8%;

[0167] The method provided by the present application provides 90 candidate sequences, and the average strain coverage from 2015 to 2019 is 97.1%, and the average type coverage is 66.3%.

[0168] Compared with the "mutation-selection-ranking" model, the average strain coverage of the method provided by the present application is increased by 10.6%, and the average type coverage is increased by 75.4%.

[0169] 2. The method provided by the present application has a higher capture rate for dominant mutants (the top 3 mutants with the highest abundance in each influenza season)

[0170] In the "mutation-selection-ranking" model, an average of 69.9% of the top 3 dominant mutations can be predicted during the period from 2015 to 2019;

[0171] In the method provided by the present application, an average of 93.8% of the top 3 dominant mutations can be predicted during the period from 2015 to 2019.

[0172] Compared with the "mutation-selection-ranking" model, the average capture rate for the top 3 dominant mutations is increased by 34.2%.

[0173] 3. The method provided by the present application can predict new mutants (non-top 3 mutants of the previous influenza season) in the future, while the "mutation-selection-sorting" strategy does not have this performance

[0174] In the method provided by the present application, the average strain coverage of new mutants during 2015-2019 is 78.4%, and the average type coverage of new mutants is 62.3%. More importantly, according to the existing data, the method provided by the present application can predict new dominant mutations in the future 2-8 years in advance, while the "mutation-selection-sorting" model does not have this prediction performance.

[0175] Among them, the type coverage type coverage: TC = T top90 / T total , wherein T top90 represents the type of the first 90 sequences of the mutant spectrum, T total represents the total number of mutant types observed in each influenza season.

[0176] Strain coverage: SC = S top90 / S total , wherein S top90 represents the first 90 sequences of the mutant spectrum, S total represents the total number of strains observed in each influenza season.

[0177] Type coverage and strain coverage examples:

[0178]

[0179] The above describes the preferred embodiments of the present application, but the present application is not limited to the specific details in the above embodiments, and various modifications can be made to the technical solutions of the present application within the technical concept of the present application, and these simple modifications all belong to the protection scope of the present application.

[0180] In addition, it should be noted that each specific technical feature and step described in the above specific embodiments can be combined in any appropriate manner without contradiction, and in order to avoid unnecessary repetition, the present application will not be described again.

[0181] In addition, various different embodiments of the present application can also be combined in any appropriate manner, as long as it does not deviate from the idea of the present application, it should also be considered as disclosed by the present application.

Claims

1. A method for delineating the HA surface antigen region of influenza virus capsid protein, characterized in that: The method includes the following steps: S101 influenza virus dataset preprocessing: Collect influenza HA nucleic acid sequence set D from public databases, use multiple sequence alignment tools to align and extract the Ps×3-2 to Pe×3 positions corresponding to the HA1 region in the HA nucleotide sequence to obtain HA1 nucleic acid sequence set S, and translate protein sequence set P based on HA1 nucleic acid sequence set S; where Ps is the start position of HA1 protein reference sequence and Pe is the end position of HA1 protein reference sequence. Phylogenetic tree construction of S102 sequence: Based on the spatiotemporal transformation principle, a representative sequence set T is constructed, which is then used to build a phylogenetic tree to obtain the parent-child node sequence set PT. The spatiotemporal transformation principle is as follows: the time period is within each flu season, and N sequences are selected respectively. m Nucleic acid sequences, spatially defined as N m The nucleic acid sequences are from different regions and / or countries, and the N m The value range is 5 to 100; S103 predicts the segmentation of immunogenic regions: According to the formula MF i =M i / A calculates the mutation frequency at each site in the HA1 protein sequence set P to obtain the high-frequency mutation site set F; where MF i M represents the mutation frequency at position i in the HA1 protein sequence; i The number of sequences at site i where the mutation occurs; A represents the total number of HA1 sequences. The set containing all high-frequency mutation sites is defined as the high-frequency mutation site set F, where the high-frequency mutation sites refer to MF. i Sites >0.05; Potential immunogenic sites for experimental validation were collected from literature and databases to obtain experimental site set A; Define F∪A as the set I of potential immunogenic sites for HA1. Calculate the distance matrix D between each site in the set I on the three-dimensional protein structure of HA1. dist The site distance matrix D dist The calculation formula is D dist ={d ij |i,j∈I,d ij ∈{minimum / average / central carbon atom distance between amino acid i and amino acid j}}; According to the clustering formula Cluster={c i c j |D min (c i c j ) <N C c i c j ∈D dist , i, j∈I}, where N C The values ​​range from 4 to 10 angstroms. In three-dimensional space, the site set I is clustered into n predicted immunogenicity regions C={C1, C2, ..., C...} based on distance. n }, and make the number of amino acid sites M in each predicted immunogenicity region a Meets the condition 5≤M a ≤18.

2. The method for delineating the HA surface antigen region of influenza virus capsid protein according to claim 1, characterized in that, In step S103: the high-frequency mutation site refers to MF i Sites >0.

1.

3. The method for delineating the HA surface antigen region of influenza virus capsid protein according to claim 1, characterized in that, In step S103: the D dist The minimum atomic distance is calculated using the formula D. dist (a i a j (a) i a j ∈{atomic coordinates}), calculate the minimum atomic distance D between each point in the site set I. dist .

4. The method for delineating the HA surface antigen region of influenza virus capsid protein according to claim 1, characterized in that, In step S103: the D dist The distance to the central carbon atom is calculated using formula D. dist (a i a j (a) i a j ∈{center carbon atom coordinates}), calculate the distance D between the center carbon atoms of each point in the site set I. dist .

5. The method for delineating the HA surface antigen region of influenza virus capsid protein according to claim 1, characterized in that, In step S103: the D dist The average atomic distance is calculated according to formula D. dist =AVE{d(a i a j )}(a i a j ∈{atomic coordinates}), calculate the average atomic distance D between each point in the site set I. dist .

6. The method for delineating the HA surface antigen region of influenza virus capsid protein according to claim 1, characterized in that, In step S103: For C located in the predicted immunogenicity region i and C j For the edge point x, according to the principle of minimization, x∈ D dist (C i C j (C) i C j ∈{predicted immunogenicity region set C}), and divide site x into a predicted immunogenicity region containing fewer sites.

7. A method for predicting future mutations of influenza viruses, characterized in that: The method includes the following steps: S201 Immunogenic Region Division: The method according to claim 1 divides the set of high-frequency mutation sites and potential immune prototype sites in the HA1 protein into n predicted immunogenic regions; S202 Prediction Template Selection: The abundance of each sequence in the HA1 protein sequence set P or the HA1 nucleic acid sequence set S is calculated using the formula ∑Ik(x=y)(I, k∈protein sequence set P) to generate an abundance descending set B; the pre-defined potential immunogenic regions are then mapped according to the mapping relationship F. m Map to HA1 sequences; select the top N from the descending abundance set B. e Using the strip as a template, a template set E is generated for each potentially immunogenic region; S203 simulated mutation: Calculate the nucleotide mutation frequency and generate a nucleotide mutation probability matrix Mn, or generate an amino acid mutation probability matrix Mp based on the amino acid mutation frequency. The nucleotide mutation probability matrix Mn can be represented as follows: Mn={N ab |a, b∈{A,T,C,G},N ab ∈{the probability of nucleotide a mutating into nucleotide b}}; the amino acid mutation probability matrix Mp can be represented as follows: Mp={M ab |a, b∈{20 amino acid types}, M ab The amino acid probability matrix Mp can be obtained by finding the probability of amino acid a mutating into amino acid b. Among them, M ab The calculation method is as follows: First, according to the formula The frequency of the triplet codons encoding amino acids is statistically analyzed, and the Res codons for each amino acid are counted. i Obtain the weight matrix Wtr of a series of triple codons i ={W1,W2,…,W Ntri }; where Otr i This indicates that the amino acid Res is encoded. i The number of times the triplet codon i appears in the nucleic acid sequence set S, Ntr i This indicates that the amino acid Res is encoded. i The total number of triplet codons, This indicates that the amino acid Res is encoded. i The total number of times all triplet codons appear in S; Then, according to the accumulation formula P(Res) i →Res j )=W tri ×W trj × ∑(P(nuc i →nuc j )×P(nuc k →nuc l )×P(nuc m →nuc n Obtain M ab The value of R1, R2 ∈ {20 amino acid types}, nuc i ,nuc j ,nuc k ,nuc l ,nuc m ,nuc n ∈{A,T,C,G}, sequence (nuc i ,nuc k ,nuc m ) indicates encoding Res i The triple codon, sequence (nucleus) j ,nuc l ,nuc n ) indicates encoding Res j The triplet codon; When the template is a nucleic acid sequence, mutation simulation is performed according to the formula R=Mn×E to obtain the potential mutation spectrum R. The potential mutation spectrum R can be expressed in the following form: , Where V is the total number of mutant sequences; When the template is a protein sequence, mutation simulation is performed according to the formula R=Mp×E to obtain the potential mutation spectrum R. The formula for calculating the potential mutation spectrum R is as follows: , Where V is the total number of mutant sequences; S204 Filtering and Sorting: According to the simulated host immune barrier formula: The potential mutation spectrum Q is obtained by measuring the ease with which each progeny sequence in the potential mutation spectrum can overcome the immune barrier, after host immune screening; where DSIB represents the difficulty of overcoming the immune barrier, and N represents the potential mutation spectrum Q. i F represents the total number of amino acid sites across all antigen regions. i This indicates whether the mutated amino acid has appeared at position i in the past; if it has, it is recorded as 0, and if it has not, it is recorded as 1. The abundance of each progeny sequence in the potential mutation spectrum Q is calculated using the formula ∑Ik(x=y) (I, k∈potential mutation spectrum Q), and then sorted in descending order to generate the potential mutation spectrum U; or the frequency of each mutation predicted in the historical mutation spectrum is counted using the formula ∑mut (mut∈potential mutation spectrum Q), and then sorted in descending order to generate the potential mutation spectrum U, selecting the top N. w The final predicted mutation spectrum W is constructed.

8. The method for predicting future mutations of influenza virus according to claim 7, characterized in that, In step S202: According to the formula ∑IK(x=y), the value is 1 when x=y, and 0 otherwise; The abundance descending set B can be represented in the following form: or , In the formula N b This represents the number of protein sequence sets P or nucleic acid sequence sets S after redundancy removal; where the abundance of each sequence in the HA1 protein sequence set P or nucleic acid sequence set S is sorted in descending order. The mapping relationship F m It can be represented in the following form: F m ={a i →b i a i ∈ Potential immunogenic site set I, b i ∈ the abundance descending set B} is mapped to the HA1 sequence, where F m m represents the potentially immunogenic region; The template set E, generated for each potential immunogenic region, can be represented in the following form: Where the number of templates N e The value range is 1 to 10.

9. The method for predicting future mutations of influenza virus according to claim 7, characterized in that, In step S204: The criterion for measuring the ease with which each progeny sequence in the potential mutation spectrum can break through the immune barrier is that if the DSIB of the progeny sequence is greater than I (0.05~0.8), it is removed, and the potential mutation spectrum Q is obtained after host immune screening. The formula for calculating the predicted mutation spectrum W is as follows: ; The N w The range of values ​​is 2 ≤ N w ≤500.

Citation Information

Patent Citations

  • Method for predicting flu antigen through model and application thereof

    CN101847179A

  • Measurement and prediction of virus genetic mutation patterns

    CN112313748A