Amplicon sequencing technology-based whole genome sequencing primer for fiveleaf virus GI.2 and application of primer
By designing amplicon sequencing primer sets for whole genome sequencing such as virus GI.2, the complex, time-consuming and cost-effective sequencing in the existing technology is solved, and fast and accurate whole genome sequencing and viral typing are achieved to support epidemic control.
Patent Information
- Application Number
- CN202510756726.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The prior art is difficult to efficiently and at low cost to perform whole genome sequencing of virus GI.2. Conventional methods are complex and time-consuming and difficult to accurately reflect viral mutation information. Metagenome sequencing is costly and has limited data.
Design a primer set based on amplicon sequencing technology to achieve high coverage and low cost whole genome sequencing through specific PCR amplification and tiled covering of the whole genome of the virus, combined with the second-generation/third-generation sequencing platform.
The rapid and accurate sequencing of the entire genome of Zaru virus GI.2 has been achieved, and the sequencing coverage and detection accuracy have been improved. It can reflect the mutation characteristics of the virus epidemic strain, providing tools for epidemic tracing and prevention.
Smart Images

Figure CN120249568A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of viral gene sequencing, and particularly relates to a primer for whole-genome sequencing of sapovirus GI.2 based on amplicon sequencing technology and its application. Background Art
[0002] Genetic detection and accurate typing of viral epidemic strains are effective means to achieve precise treatment and effective control of the epidemic, and the same is true for sapovirus. Sapovirus is an important cause of acute gastroenteritis (AGE) worldwide, which can cause outbreaks and sporadic cases of acute gastroenteritis, ranking second in the pathogen spectrum of viral acute gastroenteritis. The transmission routes of sapovirus are diverse, with strong environmental resistance, rapid virus mutation, and short immune protection time, having high infectivity and rapid transmission ability. It has been detected in various places such as kindergartens, schools, long-term care institutions, hospitals, restaurants, hotels, and cruise ships.
[0003] Sapovirus is susceptible to people of all age groups, with the highest incidence rate in children ≤5 years old, accounting for about 50% of the total population of all age groups. Currently, there are 17 strains of sapovirus epidemic and sporadic strains that can infect humans, among which the main genotypes are GI.2 and GII.3, accounting for about 70% of all sapovirus infections. Genetic detection and accurate typing of sapovirus epidemic strains are the basis for effective prevention and control of the epidemic. However, the whole-genome research on sapovirus in China is almost blank, without relevant technical methods. There are very few articles on conventional sapovirus detection and whole-genome analysis of sapovirus included in PUBMED, and they are all the first-generation sanger sequencing method. Currently, there are only 41 whole-genome sequences of GI.2 included in NCBI globally.
[0004] Those skilled in the art know that obtaining the whole genome by first-generation sequencing requires more than 10 experiments, which is complex and time-consuming. Moreover, since only one sequence is obtained, information such as virus mutations is easily masked. Not only that, metagenomic sequencing based on fecal samples or anal swab samples contains a lot of intestinal microorganisms and human-derived samples, resulting in less effective data. Sufficient sequencing depth is required, the cost is extremely high, and it is very difficult to obtain the target sequence. Therefore, developing a simple and low-cost second-generation / third-generation sequencing method for sapovirus has become an urgent technical problem to be solved.
[0005] Amplicon sequencing is a highly targeted method for analyzing gene variations in specific genomic regions. Amplicon sequencing mainly includes 16S rDNA sequencing, 18S rDNA sequencing, ITS sequencing, and target region amplicon sequencing, etc. As a complementary technology to whole-genome sequencing, amplicon capture sequencing can greatly simplify the experimental process and analysis target, is a rapid and effective technology, and plays a unique role in the new generation of high-throughput sequencing.
[0006] Patent document CN119913242A discloses a method for whole-genome sequencing of sapovirus, and the method includes: extracting viral genes, performing reverse transcription, then using an amplicon sequencing primer set for targeted amplification, constructing a library, and performing bioinformatics analysis after sequencing. This technical solution provides a primer set (10 groups) for whole-genome detection of sapovirus, but it does not explain the genotypes of sapovirus to which this primer set is applicable, or whether it is applicable to whole-genome sequencing of all genotypes of sapovirus, and does not verify the number of reaction systems for amplification required for the provided primer set, the applicable sample concentration, as well as its feasibility and detection accuracy.
[0007] Based on this, the present invention provides a primer set for whole-genome sequencing of sapovirus GI.2 and a simplified amplicon-based enrichment sequencing method. Using this primer set for whole-genome sequencing can avoid interference from other microorganisms during sequencing, ensure a higher sequencing depth of the GI.2 target gene, has the advantages of a wider whole-genome coverage, high detection accuracy, and strong specificity. Moreover, the amplicon enrichment method is convenient to operate, has low requirements for sample concentration, and the sequencing results can reflect the mutation characteristics of the main epidemic strains of sapovirus, providing a powerful tool for the rapid diagnosis and full prevention of sapovirus. Summary of the Invention
[0008] Based on this, based on amplicon sequencing technology, the present invention provides a method for whole-genome detection of sapovirus that is not for the purpose of disease diagnosis and treatment. The essence of this detection method is to perform specific PCR amplification on sapovirus through a primer set, and the amplicons generated by the amplification cover the whole-genome sequence of sapovirus in a tiled manner, thereby achieving deep sequencing of the sequence of the epidemic strain GI.2 of sapovirus. Secondly, the present invention provides a primer set as described above. On the other hand, the present invention provides an application of the primer set in the preparation of a product for whole-genome sequencing of sapovirus GI.2; the present invention also provides a kit including the primer set; The object of the present invention is achieved through the following technical solutions: In the first aspect of the present invention, the present invention provides a method for whole-genome sequencing of sapovirus GI.2 that is not for the purpose of disease diagnosis and treatment, and the method includes the following steps: (1) Extract nucleic acid from the sample to be detected and reverse transcribe to obtain a cDNA strand; (2) First-round PCR amplification Using the cDNA strand obtained in step (1) as a template, perform specific PCR amplification using the primer set and collect the amplification products; (3) Purification and splicing of the PCR amplification products; (4) Establish a library; (5) Sequencing on a machine; (6) Bioinformatics analysis.
[0009] The primer set described in step (2) consists of primer pairs 1 - 18 in Table 1, and this primer set is used for the whole genome sequencing of Saffold virus GI.2.
[0010] Specifically, primer pair 1 consists of the forward amplification primer (F) shown in SEQ ID NO.1 and the reverse primer (R) shown in SEQ ID NO.2; primer pair n consists of the forward amplification primer shown in SEQ ID NO. (2n - 1) and the reverse primer shown in SEQ ID NO.2n, where n is an integer between 2 and 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18.
[0011] In a specific embodiment of the present invention, the primer set is in two primer pools respectively. Among them, primer pairs 1, 3, 5, 7, 9, 11, 13, 15, and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16, and 18 are in primer pool 2.
[0012] Preferably, the primer concentration in primer pool 1 and primer pool 2 is 10 - 50 µM.
[0013] More preferably, the primer concentration in primer pool 1 and primer pool 2 is 10 µM.
[0014] In a specific embodiment of the present invention, after the PCR product is purified, the library is constructed using the Nextera® XT Library Prep Kit, sequenced with Miniseq, and the downloaded sequences are assembled using the CLC Genomics Workbench 23.0 software with the sequence of Genbank accession number MG515477.1 (GI.2) as the reference sequence.
[0015] Preferably, the sample to be detected described in step (1) includes but is not limited to blood, throat swabs, saliva, and infected tissues.
[0016] In the second aspect of the present invention, the present invention provides a set of primer sets, which are used for the whole genome sequencing of Saffold virus GI.2, and are characterized in that the primer set consists of primer pairs 1 - 18.
[0017] Specifically, primer pair 1 consists of a forward amplification primer (F) shown in SEQ ID NO.1 and a reverse primer (R) shown in SEQ ID NO.2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n - 1) and a reverse primer shown in SEQ ID NO.2n, where n is an integer selected from 2 to 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18.
[0018] As shown in Table 1, the primer set consists of primer pair 1 shown in SEQ ID NO.1 - 2, primer pair 2 shown in SEQ ID NO.3 - 4, primer pair 3 shown in SEQ ID NO.5 - 6, primer pair 4 shown in SEQ ID NO.7 - 8, primer pair 5 shown in SEQ ID NO.9 - 10, primer pair 6 shown in SEQ ID NO.11 - 12, primer pair 7 shown in SEQ ID NO.13 - 14, primer pair 8 shown in SEQ ID NO.15 - 16, primer pair 9 shown in SEQ ID NO.17 - 18, primer pair 10 shown in SEQ ID NO.19 - 20, primer pair 11 shown in SEQ ID NO.21 - 22, primer pair 12 shown in SEQ ID NO.23 - 24, primer pair 13 shown in SEQ ID NO.25 - 26, primer pair 14 shown in SEQ ID NO.27 - 28, primer pair 15 shown in SEQ ID NO.29 - 30, primer pair 16 shown in SEQ ID NO.31 - 32, primer pair 17 shown in SEQ ID NO.33 - 34, primer pair 18 shown in SEQ ID NO.35 - 36.
[0019] In a specific embodiment of the present invention, the primer set is in two primer pools respectively. Among them, primer pairs 1, 3, 5, 7, 9, 11, 13, 15, and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16, and 18 are in primer pool 2.
[0020] In a third aspect of the present invention, the present invention provides an application of the primer set described in the first aspect of the present invention in the preparation of a product for the whole - genome sequencing of Zhalu virus GI.2.
[0021] The product includes but is not limited to reagents, kits, chips, test strips, membrane strips, or detection platforms.
[0022] In a fourth aspect of the present invention, the present invention provides a kit, characterized in that the kit includes the primer set described in the first aspect of the present invention, or the kit includes a buffer containing the primer set described in the first aspect of the present invention.
[0023] Furthermore, the kit further includes reverse transcriptase, PCR reaction premix, and sequencing adapters.
[0024] The sequencing adapters are common sequencing adapters known to those skilled in the art and applicable to second-generation / third-generation sequencing platforms, including but not limited to illumina, Ion, or MGI. The sequencing adapters can be obtained by purchasing through commercially available kits.
[0025] The PCR reaction premix includes nuclease-free water, buffer, DNA polymerase, Mg 2+ , dNTPs, and the PCR reaction premix can be obtained by purchasing through commercially available channels.
[0026] The technical solution provided by the present invention has the following advantages: 1) After optimization, the present invention provides 18 pairs of primer pairs designed based on the overlapping principle for the gene sequence characteristics of sapovirus GI.2. The primer pairs are divided into two primer pools for PCR amplification, which can effectively shorten the amplification time and difficulty.
[0027] 2) Compared with the 1 - 1.2-fold coverage of first-generation sequencing and the inability of metagenomic sequencing to effectively obtain all targeted sequences, the optimal primer set obtained by the present invention can effectively improve the coverage of primer amplicons, ensuring a 1.5 - 1.8-fold coverage of the viral sequence, and achieving the acquisition of a fast and full-coverage whole-genome sequence.
[0028] 3) The high-fidelity primers designed based on amplicon sequencing technology provided by the present invention can ensure the authenticity of sequence mutations and have a high accuracy in virus typing detection. Description of the Drawings
[0029] Figure 1 It is a flow chart of the whole-genome sequencing of sapovirus GI.2 based on amplicon sequencing; Figure 2 It is the concentration of the purified PCR amplification products of the two primer sets screened by the present invention; Figure 3 It is the position of the primer set provided by the present invention relative to the target gene and the amplicon coverage area; Figure 4 It is the distribution map of the whole-genome sequence products of sapovirus GI.2 clinical samples. Detailed Embodiments
[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0031] Figure 1 Flow chart of the whole genome sequencing of human bocavirus GI.2 based on amplicon sequencing
[0032] Example 1: Design of primers for the whole genome sequencing of human bocavirus GI.2 The process of designing primers based on amplicon sequencing technology in the present invention is as follows: Based on the whole genome sequence of human bocavirus GI.2 included in NCBI, a candidate primer set covering the whole genome was generated using multiple PCR primer design software, and optimization and screening were carried out through the following steps: First, primers containing repetitive sequences were excluded; subsequently, primers with lengths or GC contents not meeting the standards (15 - 25bp, 35% - 65%) were filtered out; then, through genome alignment, primers targeting hypervariable regions or cross - gene regions were excluded, and primers for the beginning and end of the sequence were designed manually to ensure full genome coverage.
[0033] The primer combinations finally screened out in the present invention need to meet the following conditions: the amplified fragment < 800bp, the length difference within the group < 200bp, the genome coverage ≥ 1.5 - fold, the annealing temperature difference ≤ 10°C, and the formation of dimers is avoided, so as to obtain two efficient primer combinations suitable for dual - tube amplification, as shown in Table 1 and Table 2 respectively. In subsequent experiments, the primer sets shown in Table 1 and Table 2 were used for amplification.
[0034] Table 1 Primers for amplicon sequencing of human bocavirus GI.2 。
[0035] Table 2 Primers for amplicon sequencing of human bocavirus GI.2 。
[0036] The primer pools shown in Table 1 and Table 2 were used for dual - tube mixed PCR amplification and purification respectively. The statistical results of the concentrations of the purified PCR products are as Figure 2 shown. It can be seen that the product recovery efficiency of the primer pool shown in Table 1 is higher than that of Table 2, and the statistical test shows that the difference is statistically significant (p = 0.024). Therefore, the present invention preferably selects the 18 - pair primer set shown in Table 1 as the best primer set screened in the present invention. The positions and sequence coverage regions of the 18 - pair primers are as Figure 3 shown, where red represents primers and pink - purple represents the sequence acquisition regions.
[0037] Example 2: Method for the whole genome sequencing of human bocavirus GI.2 Step 1: Extract nucleic acids from the sample to be tested Extract the RNA in the sample to be tested using a commercially available kit.
[0038] Step 2: Reverse transcription Use the RNA obtained in Step 1 as a template and add reverse transcriptase to obtain cDNA.
[0039] The reverse transcription system includes: 8 µl of RNA template and 2 µl of reverse transcriptase (5X RT SuperMix).
[0040] The reverse transcription reaction conditions are: 25°C for 2 min; 55°C for 20 min; 95°C for 1 min; store at 4°C.
[0041] Step 3: Multiplex PCR amplification .
[0042] Use the cDNA obtained in Step 2 as a template and perform specific PCR amplification in Primer Pool 1 and Primer Pool 2 respectively. The amplification system includes: 5 µl of cDNA product, 15 µl of 2X high-fidelity enzyme; 3 µl each of 10 µmol Primer Pool 1 / Primer Pool 2; 7 µl of ddH2O.
[0043] The PCR amplification program is as follows: .
[0044] After multiplex PCR amplification, purify and splice the amplification products, then use the Nextera® XT Library Prep Kit to construct a library for sequencing on Miniseq. Use the sequence with Genbank accession number MG515477.1 (GI.2) as the reference sequence for the downloaded sequences and use the CLC Genomics Workbench 23.0 software for splicing.
[0045] Concentration optimization process of Primer Pool 1 and Primer Pool 2: Synthesize the dry powders of Primer Pool 1 and Primer Pool 2 and dilute them to 50 µmol, 20 µmol, 10 µmol, and 5 µmol with ddH2O respectively. Perform PCR amplification according to the matrix concentration combination method shown in Table 3, and count the concentration of the purified PCR products. The results are shown in Table 3.
[0046] Table 3 Matrix paired multiplex PCR purified concentration of Primer Pool 1 and Primer Pool 2 (ng / µl) .
[0047] From the results in the above table, it can be seen that when the concentrations of primer pool 1 and primer pool 2 are between 10-50 µmol, the concentrations of the products after PCR amplification and purification are all above 52 ng / µl. Considering the detection cost, the preferred primer pool concentration of the present invention is 10 µmol.
[0048] Example 3: Clinical application verification Three stool specimens of Sazavirus GI.2 with Ct values of 19.99, 22.03 and 28.15 detected by fluorescence PCR were respectively taken for two clinical validations, and whole genome sequencing was performed according to the method provided in Example 2 of the present invention.
[0049] Results: After specific PCR amplification by the method provided in Example 2 of the present invention, the concentrations of the PCR products of GI.2 were 53.00 ng / µl and 48.05 ng / µl, 40.85 ng / µl and 15.00 ng / µl, and 29.50 ng / µl and 16.80 ng / µl, respectively. After the second-generation sequencing library was built, the concentrations were 4.13 and 4.82 ng / µl, 4.94 and 5.69 ng / µl, and 1.75 and 2.56 ng / µl, respectively, which met the requirements for the second-generation sequencing products.
[0050] After whole genome sequencing of the test sample according to the method provided by the present invention, the virus is typed according to the sequencing result, and then the genome of the test sample is tested using traditional methods. The virus typing results after the test are shown in Table 4.
[0051] Table 4 GI.2 typing specificity based on amplicon whole genome sequence typing .
[0052] Three stool specimens of Sarcoma virus GI.2 covered the two dominant gene subtypes prevalent in Beijing from 2021 to 2022. The whole genome sequence was obtained by two rounds of PCR amplification, purification, library construction, and sequencing by the method provided by the present invention. The length of the whole genome sequence GI.2 was 7352-7470 bp, and the sequence distribution was shown in Figure 4 (Results of two repeated measurements). The sequence information and gene subgroup distribution are shown in Table 5.
[0053] Table 5 GI.2 sequence information and gene subgroup distribution .
[0054] The above data can confirm that the primer set for whole genome sequencing of Sazavirus GI.2 provided by the present invention and the sequence obtained by the sequencing method can be successfully used for virus genotyping, and the typing results are completely consistent with the typing results of the traditional typing method.
[0055] The primers and sequencing method for the complete genome sequencing of sapovirus GI.2 based on amplicon sequencing technology provided by the present invention can replace the traditional sequencing method to accurately genotype sapovirus. Moreover, it can further be used for virus recombination analysis. Considering that traditional genotyping only relies on the partial gene region of VP1 with a few hundred bases, the complete genome sequence covering more than seven thousand bases can identify whether recombination occurs with other genomes. The genomic sequence results obtained in the present invention are used for recombination identification, and no recombinant genotypes are found.
[0056] Meanwhile, the complete genome can also be further used for epidemic tracing and virus evolution analysis. According to the results in Table 5, it can be seen that GI.2 clade2 and clade 3 have a similarity of 99.80% to 99.82% in nucleic acid sequence and a similarity of 99.63% in amino acid sequence. There are a total of 15 nucleotide mutations and 9 amino acid mutations, and these variations are mainly concentrated in the open reading frame 1 gene region. Both branches show the highest nucleic acid sequence similarity with the strain detected in Shenzhen in 2015 (MG515477.1): the similarity of clade2 is 98.91% to 98.92%, and the similarity of clade 3 is 99.02%.
[0057] Therefore, the primers, kits and sequencing method based on amplicon sequencing technology provided by the present invention provide a powerful tool for the rapid diagnosis and full prevention of sapovirus.
[0058] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for whole-genome sequencing of sapovirus GI.2 that is not for the purpose of disease diagnosis and treatment, the method comprising the following steps: (1) Extract nucleic acid from the sample to be detected and reverse transcribe to obtain cDNA strands; (2) First-round PCR amplification Using the cDNA strands obtained in step (1) as a template, perform specific PCR amplification using a primer set and collect the amplification products; (3) Purification and splicing of the PCR amplification products; (4) Library construction; (5) Sequencing on a machine; (6) Bioinformatics analysis; The primer set described in step (2) consists of primer pairs 1-18 in Table 1, and the primer set is used for whole-genome sequencing of sapovirus GI.2; Specifically, primer pair 1 consists of a forward amplification primer (F) shown in SEQ ID NO.1 and a reverse primer (R) shown in SEQ ID NO.2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n - 1) and a reverse primer shown in SEQ ID NO.2n, where n is an integer selected from 2 to 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18.
2. The method according to claim 1, characterized in that, The primer set described in step (2) is in two primer pools respectively. Among them, primer pairs 1, 3, 5, 7, 9, 11, 13, 15, and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16, and 18 are in primer pool 2.
3. The method according to claim 2, wherein The primer concentrations in primer pool 1 and primer pool 2 are 10 - 50 µM.
4. The method according to claim 1, wherein The sample to be detected described in step (1) includes blood, throat swabs, saliva, and infected tissues.
5. Primer set, which is used for the whole genome sequencing of Sapporo virus GI.2, characterized in that, The primer set consists of primer pairs 1-18.
6. The primer set according to claim 5, characterized in that, Primer pair 1 consists of a forward amplification primer (F) shown in SEQ ID NO.1 and a reverse primer (R) shown in SEQ ID NO.2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n - 1) and a reverse primer shown in SEQ ID NO.2n, where n is an integer selected from 2 to 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18.
7. Use of the primer set according to claim 5 in the preparation of a product for whole-genome sequencing of sapovirus GI.
2.
8. The application according to claim 7, wherein The product includes reagents, reagent kits, chips, test strips, membrane strips, or detection platforms.
9. Kit, characterized in that, The reagent kit includes the primer set according to claim 5.
10. The kit according to claim 9, wherein, The reagent kit further includes reverse transcriptase, PCR reaction premix, and sequencing adapters.
Citation Information
Patent Citations
Amplification primers and amplification method for GI.1 type sapovirus genome
CN108893467A
Virus whole genome sequencing method
CN119913242A
Primer set for detecting sapovirus and detection method using the same
KR102744746B1
Cited By
Method, primer group and kit for capturing whole genome of fiveleaf virus
CN121183041A