A method, primer set and kit thereof for adenovirus whole genome sequencing

By using third-generation sequencing technology and specific primer sets, rapid and accurate detection and typing of the entire adenovirus genome can be achieved, solving the problems of limited type coverage and human DNA contamination in existing technologies, and improving sequencing efficiency and accuracy.

CN119320847BActive Publication Date: 2026-06-02PEKING UNION MEDICAL COLLEGE HOSPITAL +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNION MEDICAL COLLEGE HOSPITAL
Filing Date
2024-11-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for whole-genome sequencing of adenoviruses suffer from problems such as limited type coverage, incomplete sequencing coverage of genomic regions, high data complexity, and severe contamination by human genomic DNA, resulting in low sequencing accuracy and efficiency.

Method used

Using a third-generation sequencing technology, a method was employed to perform multiplex amplification and sequencing adapter ligation using specific primer sets. Universal multiplex primer sequences were designed and combined with PacBio and Nanopore sequencers to achieve rapid detection and typing of the entire adenovirus genome.

Benefits of technology

Adenovirus whole genome detection can be completed within 10 hours, covering all types, solving human genomic DNA contamination, improving sequencing accuracy and efficiency, simplifying library preparation steps, and enhancing data coverage uniformity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119320847B_ABST
    Figure CN119320847B_ABST
Patent Text Reader

Abstract

The application discloses a method for adenovirus whole genome sequencing, a primer group and a kit thereof, and belongs to the technical field of DNA detection.The method for adenovirus whole genome sequencing comprises the following steps: S1, first round amplification is carried out by using first round amplification primers with sequences shown in SEQ ID NO.1 to SEQ ID NO.93 in Table 1; and S2, second round amplification is carried out by using second round amplification primers with sequences shown in SEQ ID NO.94 to SEQ ID NO.285 in Table 2.The average length of data sequenced by the method is 0.8 kb, compared with short read length of a second-generation platform, the uniformity of adenovirus coverage can be greatly improved, meanwhile, the adenovirus types are comprehensively covered, and the method can be widely promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of DNA detection technology, specifically relating to a method, primer set and kit for adenovirus whole genome sequencing. Background Technology

[0002] Adenoviruses can infect a variety of mammals, including humans. Adenovirus infection can occur through multiple routes of transmission, including respiratory, contact, fecal-oral, and ocular, leading to a range of illnesses. Respiratory transmission occurs via airborne droplets, causing respiratory infections such as colds, laryngitis, bronchitis, and pneumonia, with symptoms including fever, cough, sore throat, and runny nose. Contact transmission occurs through contact with contaminated objects or direct contact with an infected person. Fecal-oral transmission occurs through contaminated food or water, causing gastrointestinal infections, particularly in infants and young children, with symptoms including diarrhea, vomiting, and abdominal pain. Ocular transmission occurs through contact with contaminated water or objects, causing conjunctivitis, characterized by red eyes, eye pain, and tearing. Adenoviruses can also cause other types of infections, such as cystitis, hepatitis, and meningoencephalitis.

[0003] The adenovirus genome is a linear double-stranded DNA molecule, approximately 26-45 kb in length. Its structure includes terminal repeat sequences (ITRs), which play crucial roles in viral genome replication and packaging. The adenovirus genome is divided into early and late genes. Early genes are expressed in the early stages of infection and are primarily involved in viral replication and host cell regulation; late genes are expressed in the later stages of viral replication and mainly encode structural proteins. Key genes include the E1A and E1B genes in the E1 region, which regulate the cell cycle and transcription, promoting viral genome replication; the E2 region encodes DNA polymerase, terminase, and DNA-binding proteins, participating in viral DNA replication; the E3 region is mainly involved in host immune escape, encoding various proteins that regulate the host immune response; and the E4 region regulates viral gene expression and DNA repair, participating in the transcriptional regulation of the viral genome. The L region encodes viral structural proteins, including capsid proteins and other proteins involved in viral assembly.

[0004] Adenovirus replication and its life cycle include steps such as attachment and invasion, early gene expression, DNA replication, late gene expression, and viral release. The virus binds to host cell receptors via fibrin, enters the cell, and then enters the nucleus via endocytosis. After the viral genome is released into the nucleus, early gene transcription begins, producing proteins that regulate the host cell environment in preparation for viral replication. Subsequently, viral DNA polymerase and other related enzymes synthesize a new viral genome. Late gene transcription synthesizes viral structural proteins, and the virus assembles to form progeny viral particles, which are ultimately released through cell lysis or exocytosis to infect new cells.

[0005] With the development of genomics technology, sequencing technology is increasingly widely used in various aspects of infectious disease tracing, detection, typing, and drug resistance assessment, and is rapidly developing towards faster and more economical methods. Traditionally, Sanger sequencing has been widely used for sequencing and validating partial regions of the adenovirus genome due to its high accuracy, but its low throughput and high cost have limited the application of whole-genome sequencing. With the development of next-generation sequencing (NGS) technologies, such as the Illumina and MGI platforms, the efficiency and throughput of adenovirus whole-genome sequencing have been greatly improved. These platforms can generate large amounts of sequence data in a short time, and multiple samples can be analyzed simultaneously through parallel sequencing. This is of great significance for adenovirus genomic variation research, epidemiological surveys, and the identification of novel adenovirus strains.

[0006] The application of NGS technology has greatly promoted the development of adenovirus genomics research. For example, researchers have used NGS technology to perform systematic whole-genome sequencing on adenovirus samples isolated from different regions and populations, revealing the genetic diversity and evolutionary patterns of adenoviruses. By comparing and analyzing adenovirus genome data, new gene variations, recombination events, and genome rearrangements can be identified, which is of great reference value for understanding the pathogenic mechanisms of adenoviruses and vaccine development. In addition, NGS technology has also been used to monitor the epidemic trends and transmission routes of adenoviruses, providing data support for public health decision-making.

[0007] Despite the numerous advantages that NGS technology has demonstrated in adenovirus whole-genome sequencing, several challenges remain. The accuracy and integrity of high-throughput sequencing data depend on a variety of factors during sample preparation and data processing. To improve sequencing accuracy, researchers typically require high-coverage sequencing and rigorous quality control and subsequent analysis using bioinformatics tools. Adenoviruses have over 80 different serotypes, with significant differences in their genomic sequences. This diversity increases the complexity of whole-genome sequencing data. When using NGS, different adenovirus types require separate alignment with reference genomes and sequence assembly, demanding substantial computational resources and time. Furthermore, for novel or rare adenovirus types, suitable reference genomes may be lacking, further complicating sequencing and analysis. The high variability and frequent recombination events in adenovirus genomes complicate the assembly and interpretation of NGS data. NGS technology typically produces short reads (usually 150-300 bases), which may fail to accurately resolve highly variable regions and repetitive sequences in the genome, leading to assembly errors or data loss. For example, certain gene regions of adenoviruses have high GC content and repetitive sequences, which often exhibit uneven sequencing coverage and assembly difficulties in NGS. Furthermore, adenovirus infection is common in various populations and environments, with widespread infection rates. NGS technology has limitations when handling mixed samples (such as environmental samples or patient co-infection samples). For mixed samples, NGS technology needs to distinguish the sequences of different adenovirus strains, requiring higher sequencing depth and complex bioinformatics analysis workflows. Low sequencing depth may lead to cross-contamination between viral strains, affecting the accuracy of the results. In addition, host DNA and other microbial DNA in mixed samples may also interfere with the sequencing and analysis of the adenovirus genome, further increasing the difficulty of data processing.

[0008] For whole-genome sequencing of adenoviruses, the current common method is to use a second-generation sequencing platform for metagenomic sequencing. This platform has problems such as long sequencing time, short read length, large amount of human genomic DNA contamination in the sequencing results, high sequencing data volume, and difficulty in obtaining sequencing results that cover the entire genome.

[0009] With the rise of third-generation sequencing technologies (such as PacBio SMRT and Oxford Nanopore), the accuracy and efficiency of adenovirus whole-genome sequencing have been further improved. Third-generation sequencing technologies can read longer DNA fragments, reduce the complexity of sequence assembly, and have unique advantages in detecting genomic structural variations.

[0010] However, there are very few reports on the use of third-generation sequencing technology for adenovirus typing and origin tracing. Existing methods can only cover a few types, and the sequencing coverage of the genome is limited, failing to provide adequate coverage of the complete genome. Therefore, the field continues to develop sequencing methods for adenovirus typing and origin tracing that provide comprehensive type coverage, broader genome coverage, and better coverage of the complete genome. Summary of the Invention

[0011] To overcome the shortcomings and defects of existing technologies, such as limited coverage of types and genomic regions, and the inability to effectively cover the complete genome, this invention provides a method, primers, and kits for whole-genome sequencing of adenoviruses.

[0012] The technical solution of the present invention is as follows:

[0013] A method for whole-genome sequencing of adenoviruses, comprising the following steps:

[0014] S1. The first round of amplification was performed using the first-round amplification primers shown in SEQ ID NO.1 to SEQ ID NO.93 in Table 1;

[0015] S2. Second-round amplification was performed using the second-round amplification primers with the sequences shown in Table 2, SEQ ID NO.94 to SEQ ID NO.285.

[0016] The method for whole-genome sequencing of adenovirus further includes the following step: S3. Sequencing adapter ligation.

[0017] The method for whole-genome sequencing of adenovirus further includes the following step: S4. Sequencing on a sequencing machine.

[0018] The system for the first round of amplification included: 2.5 pg to 5 ng / μl template, 500 mM KCl, 100 mM Tris HCl, 250 mM MgCl2, 10 mM dNTP, 5 U / μl Taq polymerase, 10 μM first round amplification primers, and water;

[0019] Preferably, the template refers to adenovirus nucleic acid;

[0020] Preferably, the reaction program for the first round of amplification includes: 95℃ for 3 min; with 98℃ for 20 s, 60℃ for 2 min, and 65℃ for 3 min as one cycle, for a total of 30 cycles;

[0021] Preferably, the system for the second round of amplification includes: 0.0025–0.5 μl / μl of the first round amplification product, 500 mM MgCl2, 100 mM Tris HCl, 250 mM MgCl2, 10 mM dNTP, 5 U / μl Taq polymerase, 10 μM of the second round amplification primers, and water;

[0022] Preferably, the reaction program for the second round of amplification includes: 95℃ for 2 min; a cycle consisting of 98℃ for 20 s, 60℃ for 1 min, and 65℃ for 2 min, for a total of 6 cycles; and 72℃ for 5 min.

[0023] Preferably, the 5' end of the second round primer has a phosphorylation modification.

[0024] The sequencing adapter ligation includes: ligation reaction after magnetic bead purification;

[0025] Preferably, the ligation reaction system comprises: 1-10 ng / μl of second-round amplification product, 50 mM MgCl2, 5 mM ATP, 50 mM DTT, 250 mM Tris-HCl, 400 U / μl of T4 ligase, and 0.03 μl / μl of sequencing adapter;

[0026] Preferably, the connection reaction is performed at 25°C for 10 minutes.

[0027] The above refers to the process of placing the ligation product in a Nanopore sequencer for sequencing.

[0028] A primer set for whole-genome sequencing of adenovirus includes a first-round amplification primer and a second-round amplification primer; the first-round amplification primer has the sequences shown in SEQ ID NO.1 to SEQ ID NO.93; and the second-round amplification primer has the sequences shown in SEQ ID NO.94 to SEQ ID NO.285.

[0029] The 5' end of the second primer has a phosphorylation modification.

[0030] A kit for whole-genome sequencing of adenovirus, comprising the primer set described above for whole-genome sequencing of adenovirus.

[0031] The kit for whole-genome sequencing of adenovirus further includes: amplification reagents and ligation reagents;

[0032] Preferably, the amplification reagents include: KCl, Tris HCl, MgCl2, dNTPs, Taq polymerase, and water;

[0033] Preferably, the ligation reagents include: purified magnetic beads, MgCl2, ATP, DTT, Tris-HCl, T4 ligase, and sequencing adapters.

[0034] The beneficial effects of this invention are as follows:

[0035] This invention, based on third-generation sequencing technology, develops a set of primers and a kit for sequencing the entire adenovirus genome. First, sequencing technologies based on PacBio and Nanopore can directly read long sequence fragments of several thousand to tens of thousands of bases, significantly reducing the risk of assembly mismatches. These technologies perform particularly well in sequencing regions with high GC content and repetitive sequences, enabling more accurate revelation of complex structural variations in the adenovirus genome. Second, by using universal multiplex primer sequences designed for conserved regions of the adenovirus genome, the entire adenovirus genome sequence can be effectively enriched from clinical samples, effectively addressing the issues of adenovirus typing and tracing.

[0036] The method of the present invention has the following characteristics:

[0037] 1. A universal adenovirus whole genome-specific targeting primer design method based on various adenovirus types and an adenovirus whole genome sequencing kit developed based on the method, which can complete the detection of the adenovirus whole genome in 10 hours;

[0038] 2. When using multiple primers designed based on this method for detection, there is no need to remove human genomic DNA during sample extraction. The DNA and RNA extracted using conventional automated instruments such as magnetic beads can meet the requirements for subsequent library construction and sequencing.

[0039] 3. This method has extremely high sensitivity and can be applied to the whole genome sequencing of clinical samples with adenovirus Ct values ​​as high as 35;

[0040] 4. During amplification, the universal sequence is added to the target band by multiplex amplification with low concentration of specific primers in the first round. Then, universal barcode primers and common rapid amplification enzymes are added directly for amplification, which can quickly complete the construction of the adenovirus whole genome library.

[0041] 5. An additional phosphorylation modification was added to the 5' end of the barcode used. The rapid amplification enzyme used during amplification will add an additional A to the extension end. The resulting amplification product can be directly ligated to the sequencing adapter via TA, which greatly simplifies the library preparation steps and shortens the experimental time.

[0042] 6. By using a long-read third-generation sequencing platform, the amplified long fragment data can be simply spliced ​​to obtain the whole genome sequence of adenovirus, which can greatly improve the stability of adenovirus type determination.

[0043] This invention develops a method for sequencing all types of adenoviruses, which can complete the whole genome sequencing of adenoviruses within 10 hours and achieve full-coverage assembly, tracing, and typing of the virus.

[0044] The method of this invention has a wide coverage, and the primer sequences are designed for all adenovirus types, theoretically adaptable to all adenovirus types, all of which can be sequenced for whole genome. However, in more than 200 adenovirus-positive swab samples collected by the first applicant's institution over the past two years, it was found that the most common adenovirus types in my country are concentrated in categories 3, 7, 1, 2, 5, 21, 11, 14, and 55. Therefore, this invention has verified the detection efficacy of the method and primer set of this invention for all common adenovirus types in my country within the scope available in clinical practice. The results are shown in Table 7 of Experimental Example 4.

[0045] This method is applicable to various clinical sample types, such as routine swabs and bronchoalveolar lavage fluid. There is no need to remove human genomic DNA during sample extraction. The DNA extracted using conventional automated instruments such as magnetic beads can meet the requirements for subsequent library construction and sequencing.

[0046] This method optimizes and adjusts the multiplex amplification system, enabling whole-genome sequencing even when the Ct value of adenovirus in the original nucleic acid sample reaches 35. It is an ultra-sensitive genome sequencing method.

[0047] This invention designs specific primers for the detection target and adds a universal sequence to the 5' end of the primers. This allows for the addition of the universal sequence to the amplified target band while simultaneously amplifying and enriching adenovirus. The library for sequencing can then be obtained by directly adding a universal barcode primer and amplifying with a common rapid amplification enzyme.

[0048] This invention optimizes the multiplex amplification system conditions multiple times. With the optimized amplification primers, the proportion of adenovirus data in the sequencing data can reach more than 80% without removing human genomic DNA, thus solving the problem of human genomic DNA contamination in adenovirus whole genome detection.

[0049] This invention utilizes a long-read third-generation sequencing platform, where amplified long fragment data only requires simple splicing to directly obtain whole-genome data. The average length of the sequencing data obtained by this method is 0.8kb, which greatly improves the uniformity of adenovirus coverage compared to the short read length of the second-generation platform. Attached Figure Description

[0050] Figure 1 The present invention provides a flowchart illustrating the experimental principle of a method for whole-genome sequencing of adenoviruses, as a specific embodiment of the present invention.

[0051] Figure 2 This is a distribution diagram of the sequencing data of sample 1 in Experimental Example 1 of the present invention.

[0052] Figure 3 For Experimental Example 2 of this invention, the values ​​were 0.5... μ l、1 μ l, 2.5 μ l, 5 μ A genome coverage distribution map of a single round of sequencing under non-purified conditions, where A to D correspond to 0.5... μ 1 μl, 1 μl, 2.5 μl, 5 μl.

[0053] Figure 4 The figure shows the average coverage distribution of adenovirus under different conditions (Ct values ​​of 35, 32, 29, and 26) in Experiment Example 3 of the present invention, where A to D correspond to Ct values ​​of 35, 32, 29, and 26, respectively. Detailed Implementation

[0054] The following detailed description of the present invention, in conjunction with the accompanying drawings, specific experimental examples, and experimental cases, does not limit the scope of protection of the present invention.

[0055] Sources of biomaterials

[0056] All adenovirus-positive swab samples used in Experiments 1-4 were from Peking Union Medical College Hospital, Chinese Academy of Medical Sciences.

[0057] Group 1 Examples, Sequencing Method of the Present Invention

[0058] This set of embodiments provides a method for whole-genome sequencing of adenovirus. All embodiments in this set share the following common feature: the method for whole-genome sequencing of adenovirus includes the following steps:

[0059] S1. The first round of amplification was performed using the first-round amplification primers shown in SEQ ID NO.1 to SEQ ID NO.93 in Table 1;

[0060] S2. Second-round amplification was performed using the second-round amplification primers with the sequences shown in Table 2, SEQ ID NO.94 to SEQ ID NO.285.

[0061] Those skilled in the art, based on the descriptions and teachings of this invention, can perform whole-genome sequencing, typing, and tracing of respiratory syncytial virus (RSV) by employing steps S1 and S2 as described above, combined with common experimental procedures in the field of molecular biology, such as those described in *Molecular Cloning: A Laboratory Manual*, and known sequencing operations. Any act of sequencing, typing, and tracing the whole genome of RSV using the above steps falls within the scope of protection of this invention.

[0062] In a further embodiment, the method for whole-genome sequencing of adenovirus further includes the following step: S3. Sequencing adapter ligation.

[0063] In a further embodiment, the method for whole-genome sequencing of adenovirus further includes the following step: S4. Sequencing on a sequencing machine.

[0064] In a specific embodiment, the system for the first round of amplification includes: 2.5 pg to 5 ng / μ template, 500mM M KCl, 100mM Tris HCl, 250mM MgCl2, 10mM dNTP, 5U / μ Taq polymerase, 10 μM first-round amplification primers, and water;

[0065] Preferably, the template refers to adenovirus nucleic acid;

[0066] Preferably, the reaction program for the first round of amplification includes: 95℃ for 3 min; with 98℃ for 20 s, 60℃ for 2 min, and 65℃ for 3 min as one cycle, for a total of 30 cycles;

[0067] Preferably, the system for the second round of amplification comprises: 0.0025–0.5 μ l / μ The first round amplification product, 500mM M KCl, 100mM Tris HCl, 250mM MgCl2, 10mM dNTP, 5U / μl Taq polymerase, 10uM second round amplification primers and water;

[0068] Preferably, the reaction program for the second round of amplification includes: 95℃ for 2 min; a cycle consisting of 98℃ for 20 s, 60℃ for 1 min, and 65℃ for 2 min, for a total of 6 cycles; and 72℃ for 5 min.

[0069] Preferably, the 5' end of the second round primer has a phosphorylation modification.

[0070] In a specific embodiment, the sequencing adapter ligation includes: ligation reaction after magnetic bead purification;

[0071] Preferably, the ligation reaction system comprises: 1-10 ng / μl of second-round amplification product, 50 mM MgCl2, 5 mM ATP, 50 mM DTT, 250 mM Tris-HCl, 400 U / μl of T4 ligase, and 0.03 μl / μl of sequencing adapter;

[0072] Preferably, the connection reaction is performed at 25°C for 10 minutes.

[0073] In a more specific embodiment, the sequencing on the instrument refers to placing the ligation product in a Nanopore sequencer for sequencing.

[0074] Group 2 Examples, Primer Set of the Present Invention

[0075] This set of embodiments provides a primer set for adenovirus whole-genome sequencing. All embodiments in this set share the following common feature: the primer set for adenovirus whole-genome sequencing includes a first-round amplification primer and a second-round amplification primer; the first-round amplification primer has the sequences shown in SEQ ID NO.1 to SEQ ID NO.93; the second-round amplification primer has the sequences shown in SEQ ID NO.94 to SEQ ID NO.285.

[0076] Those skilled in the art, based on the descriptions and teachings of this invention, can achieve whole-genome sequencing, typing, and tracing of respiratory syncytial virus (RSV) using primers with sequences shown in SEQ ID NO. 1 to SEQ ID NO. 285, combined with common experimental procedures in the field of molecular biology, such as those described in *Molecular Cloning: A Laboratory Manual* and known sequencing operations. Any synthesis, amplification, reaction, preparation, or production of primers with sequences shown in SEQ ID NO. 1 to SEQ ID NO. 285 falls within the protection scope of this invention.

[0077] In a specific embodiment, the 5' end of the second round primer is phosphorylated.

[0078] Group 3 Examples, the reagent kit of the present invention

[0079] This set of embodiments provides a kit for adenovirus whole-genome sequencing. All embodiments in this set share the following common feature: the kit for adenovirus whole-genome sequencing includes a primer set for adenovirus whole-genome sequencing as described in any of the embodiments in Group 2.

[0080] In a further embodiment, the kit for adenovirus whole genome sequencing further includes: amplification reagents and ligation reagents;

[0081] Preferably, the amplification reagents include: KCl, Tris HCl, MgCl2, dNTPs, Taq polymerase, and water;

[0082] Preferably, the ligation reagents include: purified magnetic beads, MgCl2, ATP, DTT, Tris-HCl, T4 ligase, and sequencing adapters.

[0083] In one specific embodiment of the present invention, a set of primers applicable to the amplification of large fragment universal sequences for all adenovirus types was developed. These primers are used to specifically amplify and enrich adenovirus whole genome nucleic acid fragments derived from clinical samples such as swabs and bronchoalveolar lavage fluid. The assembly and typing of the adenovirus whole genome can be completed in 10 hours. Compared with conventional NGS detection technology, this method can effectively achieve adenovirus recombination, cover the entire adenovirus genome, and maintain good amplification uniformity.

[0084] This method eliminates the need for removal of human genomic DNA during extraction. DNA extracted using conventional automated instruments such as magnetic beads is sufficient for subsequent library construction and sequencing. It is particularly suitable for sample types with high human genomic DNA content, such as swabs and bronchoalveolar lavage fluid. Detection can be completed with a total DNA extraction volume of 50 pg or more. Compared with conventional adenovirus sequencing procedures, it can significantly improve the success rate of library construction for trace samples.

[0085] This invention designs specific primers for the detection target and adds a universal sequence to the 5' end of the primer, which can achieve simultaneous coverage of the entire adenovirus genome. By detecting samples with different Ct values, it can be surprisingly found that this method can also achieve good whole genome coverage for clinical samples with Ct values ​​as high as 35.

[0086] This invention optimizes the multiplex amplification system conditions through multiple rounds of testing. The optimized primers, without removing human genomic DNA, achieve an adenovirus content of over 90% in the sequencing data, solving the problem of human genomic DNA contamination in adenovirus whole-genome sequencing. Furthermore, by using a long-read third-generation sequencing platform, the amplified long-fragment data can be easily assembled to directly obtain the entire adenovirus genome sequence, which can be used for subsequent adenovirus typing and etiological analysis.

[0087] The average length of the sequencing data obtained using this method is 0.8k, which is more than 5 times longer than the PE150 standard used in second-generation mNGS platforms, greatly improving the accuracy and stability of pathogen identification. Furthermore, it maintains a high degree of consistency in target fragment size selection, enhancing the sensitivity of adenovirus whole-genome sequencing by controlling the amplified fragment length while considering fragment length.

[0088] The detailed implementation process of this method is as follows: Figure 1 As shown, the specific operation is as follows:

[0089] 1. Synthesis of adenovirus whole genome primers

[0090] Based on the 800bp targeted amplification fragment length, the adenovirus genome comprises 46 regions. The designed universal sequence and primer sequences are shown in Table 1 below:

[0091] Table 1

[0092]

[0093]

[0094]

[0095] The number following ADV corresponds to the adenovirus amplification region. The F-terminal and R-terminal primers following the number are a specific target primer pair used for specific amplification and enrichment of the corresponding target region. In specific experiments, one or more combinations can be selected for multiplex amplification depending on the experimental objectives. Primers can be synthesized by primer synthesis companies, such as Shanghai Sangon Biotech.

[0096] 2. Universal tag sequence primer design

[0097] The universal tag sequence consists of three parts: a tag sequence at the 5' end to identify the start sequence, a 24 bp tag sequence, and a universal primer sequence at the 3' end. The 5' end of the sequence undergoes additional phosphorylation modification, and the 96 tag sequences exhibit significant differences, facilitating segmentation and identification.

[0098] 3. Synthesis of universal tag sequence primers

[0099] The tag sequences designed based on the above primer design principles are shown in Table 2 below:

[0100] Table 2

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] The primer set following PMGIB, with its numerical code, is a group of tag sequence primers used to add tag sequences to specific amplification enrichment target regions. In specific experiments, one or more combinations can be selected for multiplex amplification depending on the experimental objective. Primers can be synthesized by primer synthesis companies, such as Shanghai Sangon Biotech.

[0108] 4. Nucleic acid extraction from clinical samples

[0109] The principle of this method is based on multiplex targeted enrichment, which can specifically increase the proportion of adenovirus in sequencing data. Therefore, it is not necessary to remove human genomic DNA during extraction. This method is suitable for methods that extract DNA and RNA simultaneously. We recommend using the magnetic bead method of the Mingde automated extraction instrument for nucleic acid extraction. The extraction method should follow the kit instructions and instrument manual. The extraction time is approximately 0.5 hours.

[0110] 5. First-round multiplex PCR amplification using adenovirus-specific primers

[0111] The specific primers for the first round of amplification consist of two parts. The first part is a universal sequence located at the 5' end. The introduction of this sequence is because the universal sequence is used as the primer binding site during the second round of amplification, so that all samples can be amplified using the same set of primers. The second part is an adenovirus-specific primer located at the 3' end. This primer specifically binds to a specific region of the adenovirus, which can specifically enrich this part of the target region to be sequenced.

[0112] The adenovirus primers were uniformly and quantitatively diluted to 10 μM, then mixed in equal volumes and used as adenovirus-specific primers for the first round of amplification.

[0113] The general preparation system for the first round of amplification reaction is shown in Table 3 below:

[0114] Table 3

[0115]

[0116] After mixing the prepared reaction mixture thoroughly, centrifuge briefly and then perform the amplification reaction on a PCR instrument:

[0117]

[0118] The first round of amplification takes approximately 180 minutes.

[0119] 6. Second round of PCR amplification reaction with tagged sequence

[0120] The amplification products from the first round were purified using 0.55X magnetic beads, and the purified products could be directly used for the second round of amplification.

[0121] Alternatively, the product from the first round of PCR reaction can be used directly as a template, taking 0.5-5... μ Add the prepared reaction reagents as shown in Table 4 below to perform the second round of tagged sequence amplification.

[0122] Table 4

[0123]

[0124] After mixing the prepared reaction mixture thoroughly, centrifuge briefly and then perform the amplification reaction on a PCR instrument:

[0125]

[0126] The second round of amplification takes approximately 30 minutes.

[0127] The tagged primers used in this procedure are all phosphorylated at the 5' end. The F end and R end of the tagged primer sequence form a pair. Each sample is amplified using a pair of tagged primers, and different samples use different tagged primers. During sequencing, different samples can be directly separated based on the tagged primers.

[0128] 7. Connect sequencing adapters.

[0129] The second-round amplification products were purified using standard 1X purification magnetic beads. The specific procedure was described in the instruction manual for the purification magnetic beads. The purification magnetic beads used were from Mindray Bio-Medical Electronics (device name: Nucleic Acid Extraction Kit (Magnetic Bead Method); specification: 96 samples / kit). After purification, the Qubit concentration was determined. Equal volumes of samples prepared for sequencing in the same batch were mixed, and the total mixed volume of 50–500 ng was sufficient for subsequent ligation reactions.

[0130] The configuration of the connection system is shown in Table 5 below:

[0131] Table 5

[0132]

[0133]

[0134] After mixing the prepared reaction mixture, briefly centrifuge and ligate at 25°C for 10 minutes. After the reaction, ligate the sequencing adapter. The ligation product is then processed using the Nanopore official sequencing protocol. The ligation and sequencing process takes approximately 0.5 hours.

[0135] 8. Sequencing and Data Analysis

[0136] The sequencing data is analyzed simultaneously using Mindray analysis software or other software. A complete adenovirus genome sequence can be detected after 2–8 hours of sequencing. Refer to the software's user manual for specific analysis procedures. Other software options include ARTIC software or MetaSPAdes.

[0137] Experiment Example 1: Sequencing of the entire adenovirus genome using magnetic beads for purification.

[0138] Three adenovirus-positive swab samples were collected. After one round of amplification, the amplification products were purified using 0.55X magnetic beads. Finally, library construction and sequencing were performed according to standard operating procedures. All three samples were assembled into complete adenovirus genome sequences. A distribution diagram of a specific sample is shown below. Figure 2 As shown.

[0139] Experimental Example 2: Directly add the first round of amplification products to perform tag sequence amplification and then perform adenovirus whole genome sequencing.

[0140] Two adenovirus-positive swab samples were collected. After the first round of amplification, without magnetic bead purification, 5 μl, 2.5 μl, 1 μl, and 0.5 μl of amplification product were directly added to the second round of amplification system for further amplification. Finally, library construction and sequencing were performed according to standard operating procedures. The complete adenovirus genome sequence was obtained from all three samples under different conditions. The average adenovirus coverage under the different conditions of 0.5 μl, 1 μl, 2.5 μl, and 5 μl are as follows. Figure 3 As shown in A-3D.

[0141] Experiment Example 3: Whole-genome sequencing of adenovirus under different Ct values

[0142] Two adenovirus-positive swab samples were collected. After adenovirus Ct value detection using qPCR, the samples were serially diluted to Ct values ​​of 26, 29, 32, and 35 before adenovirus whole-genome sequencing. Adenovirus whole-genome sequences were obtained from all three samples under different conditions. The average adenovirus coverage under different conditions (Ct values ​​of 35, 32, 29, and 26) is shown below. Figure 4 As shown in A to 4D.

[0143]

[0144] Experiment Example 4: Whole Genome Sequencing of Different Types of Adenovirus Samples

[0145] Nine adenovirus-positive swab samples of different types were collected. After adenovirus Ct value detection using qPCR, the samples underwent adenovirus whole-genome sequencing. Adenovirus whole-genome sequences were obtained from all nine samples of different types. The average adenovirus coverage under different conditions is shown in Table 7 below.

[0146] Table 7

[0147] Sample Name Classification Sequencing coverage depth ADV01 Type 3 343 ADV02 Type 21 382 ADV03 Type 5 273 ADV04 Type 2 371 ADV05 Type 1 267 ADV06 Type 7 315 ADV07 Type 11 421 ADV08 Type 14 368 ADV09 Type 55 435

Claims

1. A method for non-diagnostic purposes of whole genome tri-seq of adenovirus for adenovirus positive clinical samples, characterized in that, Includes the following steps: S1. Perform the first round of amplification using the first-round amplification primers shown in SEQ ID NO.1 to SEQ ID NO.93; S2. Perform a second round of amplification using the second-round amplification primers shown in SEQ ID NO.94 to SEQ ID NO.

285.

2. A method for non-diagnostic purposes of whole genome third generation sequencing of adenovirus for adenovirus positive clinical samples according to claim 1, characterized in that, It also includes the following step: S3. Sequencing adapter ligation.

3. A method for non-diagnostic purposes of whole genome third generation sequencing of adenovirus for adenovirus positive clinical samples according to claim 1, characterized in that, It also includes the following steps: S4. Sequencing.

4. A method for non-diagnostic purposes of whole genome third generation sequencing of adenovirus in adenovirus positive clinical samples according to claim 1, characterized in that, The system for the first round of amplification included: 2.5 pg ~ 5 ng / μl template, 500 mM M KCl, 100 mM Tris HCl, 250 mM MgCl2, 10 mM dNTP, 5 U / μl Taq polymerase, 10 μM first round amplification primers and water; And / or, the template refers to: adenovirus nucleic acid; And / or, the reaction program for the first round of amplification includes: 95°C for 3 min; with 98°C for 20 s, 60°C for 2 min, and 65°C for 3 min as one cycle, for a total of 30 cycles; And / or, the system for the second round of amplification includes: 0.0025~0.5 μl / μl of the first round amplification product, 500 mM KCl, 100 mM Tris HCl, 250 mM MgCl2, 10 mM dNTP, 5 U / μl Taq polymerase, 10 μM of the second round amplification primers, and water; And / or, the reaction program for the second round of amplification includes: 95°C for 2 min; 6 cycles consisting of 98°C for 20 s, 60°C for 1 min, and 65°C for 2 min; 72°C for 5 min; And / or, the 5' end of the second round of amplification primers has a phosphorylation modification.

5. A method for non-diagnostic purposes of whole genome third generation sequencing of adenovirus for adenovirus positive clinical samples according to claim 2, characterized in that, The sequencing adapter ligation includes: ligation reaction after magnetic bead purification; And / or, the conditions for the connection reaction are 25°C for 10 min.

6. A method for non-diagnostic purposes of whole genome third generation sequencing of adenovirus in adenovirus positive clinical samples according to claim 3, characterized in that, The above refers to the process of placing the ligation product in a Nanopore sequencer for sequencing.

7. A primer set for whole genome third generation sequencing of adenovirus in adenovirus positive clinical samples, characterized in that, It includes a first-round amplification primer and a second-round amplification primer; the first-round amplification primer has the sequences shown in SEQ ID NO.1 to SEQ ID NO.93; the second-round amplification primer has the sequences shown in SEQ ID NO.94 to SEQ ID NO.

285.

8. The primer set for whole genome third generation sequencing of adenovirus for adenovirus positive clinical samples according to claim 7, characterized in that, The 5' end of the primers for the second round of amplification is phosphorylated.

9. A kit for whole genome third generation sequencing of adenovirus for adenovirus positive clinical samples, characterized in that, Includes a primer set for third-generation sequencing of the whole genome of adenovirus for adenovirus-positive clinical samples, as described in claim 7 or 8.

10. The kit for whole genome third generation sequencing of adenovirus in adenovirus positive clinical samples according to claim 9, characterized in that, It also includes: amplification reagents and / or ligation reagents; The amplification reagents include: KCl, Tris HCl, MgCl2, dNTPs, Taq polymerase, and water; The ligation reagents include: purified magnetic beads, MgCl2, ATP, DTT, Tris-HCl, T4 ligase, and sequencing adapters.