Primers for whole genome sequencing of Saruvirus GI.2 based on amplicon sequencing technology and their application

Through amplicon sequencing technology, the specific primer set and optimized library sequencing process are solved, and the efficient and low-cost problem of whole-genome sequencing of Zaru Virus GI.2 is achieved, and the sequencing results with high coverage and accuracy are achieved, supporting viral typing and recombination analysis.

CN120249568BActive Publication Date: 2025-08-29BEIJING CENT FOR DISEASE PREVENTION & CONTROL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510756726.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-29
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and at low cost to perform whole genome sequencing of virus GI.2, and conventional methods take time and easily conceal viral mutation information. Metagenome sequencing is costly and difficult to obtain effective data.

Method used

Using amplicon sequencing technology, a specific primer set was designed for PCR amplification, combined with Nextera® XT Library PrepKit for library construction and Miniseq sequencing, and splicing using CLC Genomics Workbench to optimize primer concentration and amplification conditions to achieve whole genome coverage and high accuracy sequencing.

Benefits of technology

The whole genome sequencing of Zaru Virus GI.2 is achieved, which improves sequencing depth and coverage, ensures the accuracy of sequencing results and the accuracy of viral typing, can reflect the characteristics of viral mutations, and provides a powerful tool for rapid diagnosis and prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120249568B_ABST
    Figure CN120249568B_ABST
Patent Text Reader

Abstract

The present invention discloses primers for sequencing the whole genome of Saruvirus GI.2 based on amplicon sequencing technology and their applications, belonging to the field of viral gene sequencing technology. Based on this technology, the present invention preferably obtains a primer set consisting of 18 primer pairs. This primer set can effectively improve primer amplicon coverage, ensuring coverage of 1.5-1.8 times the viral sequence, shortening amplification time and difficulty, and improving primer amplicon coverage and whole-genome capture efficiency. This is of great significance for tracing the source and spread of Saruvirus GI.2 epidemics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of viral gene sequencing, and specifically relates to a primer for sequencing the whole genome of the Saruvirus GI.2 based on amplicon sequencing technology and its application. Background Art

[0002] Genetic testing and accurate typing of circulating viral strains are effective means for targeted treatment and effective epidemic control, including for Sapovirus. Sapovirus is a major cause of acute gastroenteritis (AGE) worldwide, causing outbreaks and sporadic episodes of the disease, ranking second among viral causes of AGE. Sapoviruses have diverse transmission pathways, strong environmental resilience, rapid viral mutations, and short-lived immune protection, making them highly contagious and capable of rapid transmission. They have been detected in a variety of settings, including kindergartens, schools, long-term care facilities, hospitals, restaurants, hotels, and cruise ships.

[0003] Sarcoviruses are susceptible to people of all ages, with the highest incidence in children ≤5 years, who account for approximately 50% of all age groups. Currently, there are 17 epidemic and sporadic strains of Sarcoviruses that can infect humans, with the predominant genotypes being GI.2 and GII.3, which account for approximately 70% of all Sarcovirus infections. Genetic testing and accurate typing of epidemic Sarcovirus strains are essential for effective epidemic prevention and control. However, Sarcovirus whole-genome research in China is virtually nonexistent, with no relevant technical methods. Conventional Sarcovirus testing and Sarcovirus whole-genome analysis published in PUBMED are very rare, and most of these articles use first-generation Sanger sequencing. Currently, there are only 41 GI.2 whole-genome sequences indexed by NCBI worldwide.

[0004] Those skilled in the art know that first-generation sequencing requires more than 10 experiments to obtain a whole genome, which is complex and time-consuming. Moreover, since only one sequence is obtained, it is easy to mask information such as viral mutations. In addition, metagenomic sequencing based on stool samples or anal swab samples contains many intestinal microorganisms and human samples, and the effective data obtained is relatively small. Sufficient sequencing depth is required, the cost is extremely high, and it is difficult to obtain the target sequence. Therefore, the development of a simple and low-cost second-generation / third-generation sequencing method for Zaru virus has become a technical problem that urgently needs to be overcome.

[0005] Amplicon sequencing is a highly targeted method used to analyze genetic variation in specific genomic regions. Amplicon sequencing primarily includes 16S rDNA sequencing, 18S rDNA sequencing, ITS sequencing, and targeted region amplicon sequencing. Amplicon capture sequencing, as a complementary technology to whole-genome sequencing, can significantly simplify experimental workflows and analytical objectives. It is a rapid and effective technology that plays a unique role in next-generation high-throughput sequencing.

[0006] Patent document CN119913242A discloses a method for whole-genome sequencing of saproviruses. The method comprises extracting viral genes, performing targeted amplification using an amplicon sequencing primer set after reverse transcription, constructing a library, and performing bioinformatics analysis after sequencing. This technical solution provides a set of 10 primer sets for whole-genome sequencing of saproviruses. However, it does not specify whether these primer sets are suitable for sequencing the whole genome of saproviruses or for all genotypes of saproviruses. Furthermore, it does not verify the number of reaction systems required for amplification using the provided primer sets, the applicable sample concentration, or their feasibility and detection accuracy.

[0007] Based on this, the present invention provides a whole-genome sequencing primer set for Zaruvirus GI.2 and a simplified version of the amplicon-based enrichment sequencing method. The use of the primer set for whole-genome sequencing can avoid interference with sequencing by other microorganisms, ensure a higher sequencing depth of the GI.2 target gene, and have the advantages of wider whole-genome coverage, high detection accuracy, and strong specificity. The amplicon enrichment method is easy to operate and has low sample concentration requirements. The sequencing results can reflect the mutation characteristics of the main epidemic strains of Zaruvirus, providing a powerful tool for the rapid diagnosis and adequate prevention of Zaruvirus. Summary of the Invention

[0008] Based on this, and based on amplicon sequencing technology, the present invention provides a method for detecting the whole genome of the Saruvirus that is not intended for disease diagnosis and treatment. The essence of the detection method is to perform specific PCR amplification of the Saruvirus using a set of primers. The amplicons generated by the amplification cover the whole genome sequence of the Saruvirus in a shingled manner, thereby achieving deep sequencing of the sequence of the epidemic Saruvirus strain GI.2. Secondly, the present invention provides a primer set as described above. On the other hand, the present invention provides a use of the primer set in the preparation of a product for sequencing the whole genome of the Saruvirus GI.2; the present invention also provides a kit comprising the primer set;

[0009] The purpose of the present invention is achieved through the following technical solutions:

[0010] In a first aspect of the present invention, a method for sequencing the whole genome of a Sativa virus GI.2 not for the purpose of disease diagnosis and treatment is provided, the method comprising the following steps:

[0011] (1) Extract nucleic acid from the sample to be tested and reverse transcribe to obtain cDNA chain;

[0012] (2) First round of PCR amplification

[0013] Using the cDNA chain obtained in step (1) as a template, perform specific PCR amplification using a primer set and collect the amplified product;

[0014] (3) Purification and splicing of PCR amplification products;

[0015] (4) Establish a library;

[0016] (5) Sequencing on a machine;

[0017] (6) Bioinformatics analysis.

[0018] The primer set described in step (2) consists of primer pair 1 to primer pair 18 in Table 1, and the primer set is used for whole genome sequencing of Saprovirus GI.2.

[0019] Specifically, primer pair 1 consists of a forward amplification primer (F) shown in SEQ ID NO. 1 and a reverse primer (R) shown in SEQ ID NO. 2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n-1) and a reverse primer shown in SEQ ID NO. 2n, where n is selected from an integer between 2 and 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, and 18.

[0020] In a specific embodiment of the present invention, the primer groups are in two primer pools respectively, wherein primer pairs 1, 3, 5, 7, 9, 11, 13, 15 and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16 and 18 are in primer pool 2.

[0021] Preferably, the primer concentrations in primer pool 1 and primer pool 2 are 10-50 μM.

[0022] More preferably, the primer concentration in the primer pool 1 and the primer pool 2 is 10 μM.

[0023] In a specific embodiment of the present invention, the PCR product was purified and library construction was performed using the Nextera® XT Library Prep Kit. Sequencing was performed using Miniseq, and the resulting sequence was spliced ​​using CLC Genomics Workbench 23.0 software, using Genbank accession number MG515477.1 (GI.2) as the reference sequence.

[0024] Preferably, the sample to be tested in step (1) includes but is not limited to blood, throat swab, saliva, and infected tissue.

[0025] In the second aspect of the present invention, the present invention provides a primer set for whole genome sequencing of Saprovirus GI.2, characterized in that the primer set consists of primer pair 1 to primer pair 18.

[0026] Specifically, primer pair 1 consists of a forward amplification primer (F) shown in SEQ ID NO. 1 and a reverse primer (R) shown in SEQ ID NO. 2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n-1) and a reverse primer shown in SEQ ID NO. 2n, where n is selected from an integer between 2 and 18, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, and 18.

[0027] As shown in Table 1, the primer set consists of primer pair 1 shown in SEQ ID NOs. 1-2, primer pair 2 shown in SEQ ID NOs. 3-4, primer pair 3 shown in SEQ ID NOs. 5-6, primer pair 4 shown in SEQ ID NOs. 7-8, primer pair 5 shown in SEQ ID NOs. 9-10, primer pair 6 shown in SEQ ID NOs. 11-12, primer pair 7 shown in SEQ ID NOs. 13-14, primer pair 8 shown in SEQ ID NOs. 15-16, primer pair 9 shown in SEQ ID NOs. 17-18, primer pair 10 shown in SEQ ID NOs. 19-20, primer pair 11 shown in SEQ ID NOs. 21-22, primer pair 12 shown in SEQ ID NOs. 23-24, primer pair 13 shown in SEQ ID NOs. 25-26, primer pair 14 shown in SEQ ID NOs. 27-28, primer pair 15 shown in SEQ ID NOs. 29-30, primer pair 16 shown in SEQ ID NOs. It consists of primer pair 16 shown in SEQ ID NOs. 31-32, primer pair 17 shown in SEQ ID NOs. 33-34, and primer pair 18 shown in SEQ ID NOs. 35-36.

[0028] In a specific embodiment of the present invention, the primer groups are in two primer pools respectively, wherein primer pairs 1, 3, 5, 7, 9, 11, 13, 15 and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16 and 18 are in primer pool 2.

[0029] In the third aspect of the present invention, the present invention provides a use of the primer set described in the first aspect of the present invention in preparing a product for whole genome sequencing of Saprovirus GI.2.

[0030] The products include but are not limited to reagents, test kits, chips, test strips, membrane strips or detection platforms.

[0031] In a fourth aspect of the present invention, the present invention provides a kit, characterized in that the kit comprises the primer set described in the first aspect of the present invention, or the kit comprises a buffer solution containing the primer set described in the first aspect of the present invention.

[0032] Furthermore, the kit also includes reverse transcriptase, PCR reaction premix, and sequencing adapter.

[0033] The sequencing adapter is a universal sequencing adapter known to those skilled in the art and applicable to second-generation / third-generation sequencing platforms, including but not limited to Illumina, Ion or MGI, and the sequencing adapter can be purchased through commercially available kits.

[0034] The PCR reaction premix includes nuclease-free water, buffer, DNA polymerase, Mg 2+ , dNTPs, and the PCR reaction premix can be purchased from commercial sources.

[0035] The technical solution provided by the present invention has the following advantages:

[0036] 1) After optimization, the present invention provides a set of 18 primer pairs designed using the shingled principle based on the sequence characteristics of the Saprovirus GI.2 gene. The primer pairs are divided into two primer pools for PCR amplification, which can effectively shorten the amplification time and difficulty.

[0037] 2) Compared with the 1-1.2x coverage of first-generation sequencing and the inability of metagenomic sequencing to effectively obtain all targeted sequences, the optimal primer set obtained by the present invention can effectively improve the coverage of primer amplicon, ensuring coverage of 1.5-1.8 times the viral sequence, and achieving rapid and full coverage of the whole genome sequence.

[0038] 3) The high-fidelity primers designed based on amplicon sequencing technology provided by the present invention can ensure the authenticity of sequence mutations and have a high accuracy rate for virus typing detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flowchart for sequencing the whole genome of Saruvirus GI.2 based on amplicon sequencing;

[0040] Figure 2 The concentrations of the purified PCR amplification products of the two primer sets screened in the present invention are:

[0041] Figure 3 The position of the primer set provided by the present invention relative to the target gene and the area covered by the amplicon;

[0042] Figure 4 This is the distribution map of the whole genome sequence products of Sarovirus GI.2 clinical samples. DETAILED DESCRIPTION

[0043] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.

[0044] Figure 1 Flow chart for whole genome sequencing of Saruvirus GI.2 based on amplicon sequencing.

[0045] Example 1: Design of primers for whole genome sequencing of Sativa virus GI.2

[0046] The process of designing primers based on amplicon sequencing technology in the present invention is as follows:

[0047] Based on the full genome sequence of Szabovirus GI.2 included in NCBI, multiplex PCR primer design software was used to generate candidate primer sets covering the entire genome, and optimization screening was performed through the following steps: first, primers containing repetitive sequences were eliminated; then, primers whose length or GC content did not meet the standards (15-25bp, 35%-65%) were filtered out; then, through genome alignment, primers targeting highly variable regions or cross-gene regions were excluded, and sequence head and end primers were manually designed to ensure full genome coverage.

[0048] The primer combinations ultimately selected by the present invention must meet the following conditions: amplified fragment <800 bp, intra-group length variation <200 bp, genome coverage ≥1.5-fold, annealing temperature difference ≤10°C, and avoidance of dimer formation. This yielded two sets of highly efficient primer combinations suitable for double-tube amplification, as shown in Tables 1 and 2, respectively. In subsequent experiments, the present invention used the primer sets shown in Tables 1 and 2 for amplification.

[0049] Table 1 Sequencing primers for the Zaru virus GI.2 amplicon

[0050]

[0051] Table 2 Sequencing primers for the Zaru virus GI.2 amplicon

[0052]

[0053] The primer pools shown in Table 1 and Table 2 were used for double-tube PCR amplification and purification, and the statistical results of the purified PCR product concentration were as follows: Figure 2As shown, it can be seen that the product recovery efficiency of the primer pool shown in Table 1 is higher than that in Table 2, and the statistical test shows that the difference is statistically significant (p=0.024). Therefore, the present invention preferably uses the 18 primer pairs shown in Table 1 as the best primer set screened by the present invention. The positions and sequence coverage areas of the 18 primer pairs are shown in Table 1. Figure 3 As shown, red represents primers and pink and purple represent the acquired sequence regions.

[0054] Example 2: Whole-genome sequencing of Sativa virus GI.2

[0055] Step 1: Extract nucleic acid from the sample to be tested

[0056] RNA was extracted from the samples using a commercially available kit.

[0057] Step 2: Reverse transcription

[0058] Using the RNA obtained in step 1 as a template, reverse transcriptase was added to obtain cDNA.

[0059] The reverse transcription system includes: 8µl of RNA template and 2µl of reverse transcriptase (5X RT SuperMix).

[0060] The reverse transcription reaction conditions were as follows: 25°C, 2 min; 55°C, 20 min; 95°C, 1 min; and storage at 4°C.

[0061] Step 3: Multiplex PCR amplification

[0062] .

[0063] Using the cDNA obtained in step 2 as a template, specific PCR amplification was performed in primer pool 1 and primer pool 2, respectively. The amplification system included: 5 µl of cDNA product, 15 µl of 2X high-fidelity enzyme, 3 µl of 10 µmol primer pool 1 / primer pool 2, and 7 µl of ddH2O.

[0064] The PCR amplification procedure is as follows:

[0065] .

[0066] After multiplex PCR amplification, the amplified products were purified and spliced. Libraries were constructed using the Nextera® XT Library Prep Kit for Miniseq sequencing. Sequences were assembled using CLC Genomics Workbench 23.0 software, using Genbank accession number MG515477.1 (GI.2) as a reference sequence.

[0067] Concentration optimization process of primer pool 1 and primer pool 2:

[0068] Primer pool 1 and primer pool 2 were synthesized into dry powders and diluted with ddH2O to 50 µmol, 20 µmol, 10 µmol, and 5 µmol, respectively. PCR amplification was performed according to the matrix concentration combination method shown in Table 3. The concentration of the purified PCR products was statistically analyzed, and the results are shown in Table 3.

[0069] Table 3 Concentrations of primer pools 1 and 2 after purification (ng / µl) for matrix-paired multiplex PCR

[0070] .

[0071] As can be seen from the results in the table above, when the concentrations of primer pools 1 and 2 were between 10 and 50 µmol, the concentrations of the purified PCR products were both above 52 ng / µl. Considering the cost of testing, the preferred primer pool concentration for this invention is 10 µmol.

[0072] Example 3: Clinical application verification

[0073] Three stool specimens with Saruvirus GI.2 Ct values ​​of 19.99, 22.03, and 28.15 detected by fluorescence PCR were collected for two clinical validations, and whole genome sequencing was performed according to the method provided in Example 2 of the present invention.

[0074] Results: After specific PCR amplification using the method provided in Example 2 of the present invention, the concentrations of the PCR products of GI.2 were 53.00 ng / µl and 48.05 ng / µl, 40.85 ng / µl and 15.00 ng / µl, and 29.50 ng / µl and 16.80 ng / µl, respectively. After the second-generation sequencing library was constructed, the concentrations were 4.13 and 4.82 ng / µl, 4.94 and 5.69 ng / µl, and 1.75 and 2.56 ng / µl, respectively, meeting the requirements for the second-generation sequencing product.

[0075] After whole genome sequencing of the test sample according to the method provided by the present invention, the virus is typed according to the sequencing results, and then the genome of the test sample is tested using traditional methods. The virus typing results after detection are shown in Table 4.

[0076] Table 4 GI.2 amplicon-based whole genome sequence typing specificity

[0077]

[0078] Three stool specimens of Saprovirus GI.2 covered the two dominant genotypes prevalent in Beijing from 2021 to 2022. The whole genome sequence was obtained by PCR amplification, purification, library construction, and sequencing twice using the method provided by the present invention. The length of the whole genome sequence GI.2 was 7352-7470 bp, and the sequence distribution is shown in Figure 2. Figure 4 (Results of two repeated measurements). The sequence information and gene subgroup distribution are shown in Table 5.

[0079] Table 5 GI.2 sequence information and gene subgroup distribution

[0080]

[0081] The above data can confirm that the primer set for whole genome sequencing of Saruvirus GI.2 provided by the present invention and the sequence obtained by the sequencing method can be successfully used for virus genotyping, and the typing results are completely consistent with the typing results of traditional typing methods.

[0082] The primers and sequencing method for whole-genome sequencing of Saruvirus GI.2, based on amplicon sequencing technology, provided by this invention, can replace traditional sequencing methods for accurate Saruvirus genotyping. Furthermore, they can be further used for viral recombination analysis. While traditional genotyping relies on a partial VP1 gene region of only a few hundred bases, a full genome sequence covering over 7,000 bases can identify recombination with other genomes. The genome sequence obtained in this invention was used for recombination identification, and no recombinant genotypes were found.

[0083] The whole genome can also be used for epidemic tracing and viral evolution analysis. Table 5 shows that GI.2 clade 2 and clade 3 share 99.80% to 99.82% nucleic acid sequence similarity and 99.63% amino acid sequence similarity. They also share 15 nucleotide mutations and 9 amino acid mutations, primarily concentrated in the open reading frame 1 gene region. Both clades show the highest nucleic acid sequence similarity with the strain detected in Shenzhen in 2015 (MG515477.1): clade 2 shares 98.91% to 98.92% similarity, and clade 3 shares 99.02% similarity.

[0084] Therefore, the primers, kits, and sequencing methods based on amplicon sequencing technology provided by the present invention provide powerful tools for the rapid diagnosis and adequate prevention of Sazavirus.

[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for whole-genome sequencing of a Sativa virus GI.2 not for the purpose of disease diagnosis and treatment, comprising the following steps: (1) Extract nucleic acid from the sample to be tested and reverse transcribe to obtain cDNA chain; (2) First round of PCR amplification Using the cDNA chain obtained in step (1) as a template, perform specific PCR amplification using a primer set and collect the amplified product; (3) Purification and splicing of PCR amplification products; (4) Establish a library; (5) Sequencing on a machine; (6) Bioinformatics analysis; The primer set described in step (2) consists of primer pair 1 to primer pair 18, and the primer set is used for whole genome sequencing of Sativa virus GI.2; primer pair 1 consists of the forward amplification primer shown in SEQ ID NO. 1 and the reverse primer shown in SEQ ID NO. 2; primer pair n consists of the forward amplification primer shown in SEQ ID NO. (2n-1) and the reverse primer shown in SEQ ID NO. 2n, and n is selected from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 and 18.

2. The method according to claim 1, characterized in that The primer sets in step (2) are respectively in two primer pools, wherein primer pairs 1, 3, 5, 7, 9, 11, 13, 15 and 17 are in primer pool 1, and primer pairs 2, 4, 6, 8, 10, 12, 14, 16 and 18 are in primer pool 2.

3. The method according to claim 2, characterized in that The primer concentrations in the primer pool 1 and the primer pool 2 are 10-50 μM.

4. The method according to claim 1, wherein The sample to be tested in step (1) is selected from blood, throat swab, saliva or infected tissue.

5. A primer set for whole genome sequencing of Sativa virus GI.2, characterized in that: The primer set consists of primer pair 1 to primer pair 18, wherein primer pair 1 consists of a forward amplification primer shown in SEQ ID NO.1 and a reverse primer shown in SEQ ID NO.2; primer pair n consists of a forward amplification primer shown in SEQ ID NO. (2n-1) and a reverse primer shown in SEQ ID NO.2n, and n is selected from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 and 18.

6. Use of the primer set according to claim 5 in preparing a product for whole genome sequencing of Saprovirus GI.

2.

7. The use according to claim 6, characterized in that The product is selected from reagents, test kits, chips, test papers, membrane strips or detection platforms.

8. A kit, characterized in that The kit comprises the primer set according to claim 5.

9. The kit according to claim 8, characterized in that The kit also includes reverse transcriptase, PCR reaction premix and sequencing adapter.

Citation Information

Patent Citations

  • Amplification primers and amplification method for GI.1 type sapovirus genome

    CN108893467A

  • Virus whole genome sequencing method

    CN119913242A