Method for constructing macro virus group sequencing library for removing host ribosomal RNA (Ribosomal Ribonucleic Acid) by enzyme method and kit
By removing host ribosomal RNA through a combination of DNA probe hybridization and enzymatic digestion, a high-quality metaviromic sequencing library was constructed, which solved the problem of insufficient viral whole-genome sequencing coverage in existing technologies and achieved efficient pathogen identification and whole-genome assembly.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG CENT FOR DISEASE CONTROL & PREVENTION
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to efficiently and minimally enrich viral RNA from clinical samples with extremely low viral loads in the context of a high proportion of host nucleic acid, resulting in insufficient coverage and sensitivity of viral whole-genome sequencing, which fails to meet the needs of high-coverage whole-genome assembly and pathogen identification.
Host ribosomal RNA was removed by DNA probe hybridization combined with RNase H and DNase I enzymatic digestion. Subsequently, RNA fragmentation, reverse transcription, adapter ligation and PCR amplification were performed to construct a high-quality metaviromic sequencing library.
This technology enables efficient enrichment of viral RNA in complex clinical samples, constructs stable DNA libraries suitable for high-throughput sequencing, improves the coverage and identification accuracy of the viral whole genome sequence, and meets the needs of high-sensitivity pathogen identification and whole genome assembly for low-load samples.
Smart Images

Figure CN121951002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biotechnology, and in particular to a method and kit for constructing a metaviromic sequencing library by enzymatically removing host ribosomal RNA. Background Technology
[0002] With the overlapping global spread of SARS-CoV-2 and influenza viruses, the demand for routine monitoring continues to increase. Relying solely on traditional real-time quantitative PCR (qPCR) technology is insufficient to comprehensively address the needs for co-infection monitoring of multiple respiratory viruses, mutation tracking, and the discovery of low-abundance subtypes. Therefore, a metaviromic sequencing-based full pathogen spectrum + mutation monitoring strategy has been rapidly incorporated into key prevention and control technical solutions. This strategy can simultaneously and unbiasedly detect multiple respiratory viruses in a single test, directly obtaining their whole genome sequences and mutation information. This provides crucial evidence for accurately identifying transmission chains and predicting drug-resistant mutations. The core of this technology lies in its ability to efficiently obtain high-quality viral genome sequences from complex clinical samples.
[0003] Currently, high-throughput whole-genome sequencing of RNA viruses such as the novel coronavirus and influenza virus mainly relies on two technical approaches: For known virus species, PCR amplification is usually performed using specific primers designed for their genome, or hybridization capture is performed using pre-designed probes, followed by the construction of sequencing libraries; For the identification of unknown or untyped pathogens, non-targeted metaviromymic (or metaviromymic) sequencing technology is commonly used. The mainstream approach is as follows: After total RNA is extracted from the sample, biotin-labeled host (human) ribosomal RNA (rRNA) antisense DNA probes are hybridized with the sample at high temperature, and then the DNA / RNA hybridization complex is captured and removed by streptavidin-coated magnetic beads, thereby consuming the majority of the host rRNA in the sample; after the host RNA is removed, the remaining RNA is reverse transcribed, and libraries suitable for second-generation sequencing platforms are constructed by adapter ligation or transposase insertion.
[0004] Although the aforementioned probe hybridization-magnetic bead capture method for removing host rRNA is stable in cell lines or standards, it exhibits significant systemic drawbacks when dealing with real clinical samples, especially throat swabs with low viral load, partially degraded RNA, and potential inactivation: First, the magnetic bead capture process for rRNA removal involves significant physical loss of nucleic acids. When the proportion of viral RNA in the sample is extremely low and the total initial amount is minimal, the number of viral RNA molecules available for library construction after this step approaches the critical value, and the final library concentration is often below the detection limit. Second, to compensate for the reduced library yield, an additional 3-5 rounds of PCR amplification are required, directly resulting in a redundant read ratio exceeding 30% in the sequencing data and an extremely low effective sequence ratio of the target virus. Third, it can only barely meet the preliminary identification of virus species, while high-coverage viral typing and whole-genome assembly are often difficult to achieve due to uneven coverage or insufficient depth. Therefore, existing technologies cannot simultaneously meet the dual requirements of high-sensitivity pathogen identification and high-coverage whole-genome sequencing for clinical samples with extremely low viral load and extremely high host background in a single library preparation, and there is an urgent need to provide a new library preparation strategy. Summary of the Invention
[0005] In view of this, this application provides a method and kit for constructing a metaviromic sequencing library by enzymatically removing host ribosomal RNA. This method can efficiently and with low damage enrich viral RNA from clinical samples with extremely low viral load in a high proportion of host nucleic acid background, and successfully construct a metaviromic sequencing library that can be used to obtain the whole genome sequence of the virus. This allows for simultaneous completion of virus typing and whole genome assembly in a single library construction.
[0006] Specifically, this application is implemented through the following technical solution:
[0007] The first aspect of this application provides a method for constructing a metavinomic sequencing library by enzymatically removing host ribosomal RNA. The method is used to obtain the full-length viral genome sequence from respiratory clinical samples with a Ct value not exceeding 32 detected by real-time quantitative PCR. The method includes:
[0008] The total RNA of the respiratory clinical sample was hybridized with a DNA probe targeting host ribosomal RNA to form a DNA / RNA hybrid strand; before reverse transcription, the RNA in the hybrid strand was digested with RNase H, and the remaining DNA probe in the system was digested with DNase I; after purification, enriched viral RNA was obtained; wherein, the method does not include the step of removing linear RNA;
[0009] The viral RNA was fragmented.
[0010] The fragmented RNA was reverse transcribed to synthesize the first-strand cDNA. In the same reaction system, the second-strand cDNA was synthesized simultaneously using the first-strand cDNA as a template. End repair and dA tail addition were performed to obtain double-stranded cDNA.
[0011] The adapter with the dT tail is ligated to the double-stranded cDNA to obtain the ligation product;
[0012] The ligation product was amplified by PCR and purified to obtain a metaviromic sequencing library.
[0013] A second aspect of this application provides a kit for constructing a metavinomic sequencing library by enzymatically removing host ribosomal RNA, the kit comprising:
[0014] Remove host ribosomal RNA components, including DNA probes targeting host ribosomal RNA, RNase H, DNase I, and their respective reaction buffers;
[0015] Library preparation components include reagents for RNA fragmentation, cDNA synthesis, adapter ligation, and PCR amplification;
[0016] And, purification components.
[0017] The method and kit for constructing metavinomic sequencing libraries by enzymatically removing host ribosomal RNA provided in this application firstly replaces the traditional magnetic bead capture and removal method by employing a core strategy of enzymatic digestion with RNase H and DNase I sequentially after DNA probe hybridization. This allows for direct and specific degradation of host ribosomal RNA in solution and removal of residual probes, significantly reducing the loss of target viral RNA caused by physical adsorption and washing, and improving the recovery efficiency and enrichment specificity of viral nucleic acid in low-load samples. Secondly, the method integrates a complete library construction process from RNA fragmentation to PCR amplification. Through standardized steps including end repair and tailing, and adapter ligation, it ensures that the viral RNA enriched after the aforementioned enzymatic treatment can be efficiently and completely converted into a stable DNA library with a standardized structure suitable for high-throughput sequencing. Finally, this integrated method enables the construction of libraries from specific enrichment to full sequence coverage of trace viral RNA in complex clinical samples, providing a reliable and efficient preparation basis for sequencing-based virus identification and genome analysis. Attached Figure Description
[0018] Figure 1 A flowchart of Example 1 of the method for constructing a metaviromic sequencing library by enzymatic removal of host ribosomal RNA provided in this application;
[0019] Figure 2 This application shows a map of the depth distribution of the whole genome sequence of the influenza virus.
[0020] Figure 3 This application shows a depth distribution map of the whole genome sequence coverage of a novel coronavirus. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0022] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0023] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0024] The following specific embodiments are given to illustrate the technical solution of this application in detail.
[0025] Example 1
[0026] Figure 1 This is a flowchart of Example 1 of the method for constructing a metaviromic sequencing library by enzymatic removal of host ribosomal RNA provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:
[0027] S101. The total RNA of the respiratory clinical sample is hybridized with a DNA probe targeting host ribosomal RNA to form a DNA / RNA hybrid strand; before reverse transcription, the RNA in the hybrid strand is digested with RNase H, and the remaining DNA probe in the system is digested with DNase I; after purification, enriched viral RNA is obtained; wherein, the method does not include the step of removing linear RNA.
[0028] It should be noted that this method targets respiratory clinical samples, such as pharyngeal swabs or bronchoalveolar lavage fluid routinely collected in disease surveillance. These samples contain a large number of exfoliated human (host) cells, resulting in extremely high background levels of host nucleic acids, especially ribosomal RNA, while the RNA content of the target pathogen, such as the novel coronavirus or influenza virus, is extremely low. The total RNA in the sample refers to the mixture of all RNA extracted from the aforementioned clinical samples, including both the viral genomic RNA to be detected and the dominant host RNA (such as messenger RNA, transfer RNA, etc.). Preferably, commercial kits such as the QIAampViral RNA Mini Kit or RNeasy Mini Kit are used for nucleic acid extraction. The obtained RNA sample should meet the following requirements: a total volume greater than 11 µL, a Ct / Cq value not higher than 32 as detected by real-time quantitative PCR (Realtime-PCR), or a concentration not lower than 1 ng / µL. Host ribosomal RNA is the most significant interfering component; if it is not removed, it will occupy most of the subsequent sequencing data, overwhelming the viral sequence information and hindering effective analysis.
[0029] Specifically, in order to selectively remove host ribosomal RNA, this application employs a method based on a combination of DNA probe hybridization and specific enzyme digestion. The total RNA of the sample is hybridized with a DNA probe targeting the host ribosomal RNA to form a DNA / RNA hybrid strand. The RNA in the hybrid strand is first digested with RNase H, and then the remaining DNA probe in the system is digested with DNase I. Finally, the enriched viral RNA is obtained through purification.
[0030] It is important to note that the DNA probes are a series of artificially synthesized, short, single-stranded DNA oligonucleotide sequences designed based on specific conserved regions of human host ribosomal RNA, enabling them to achieve base complementarity with these rRNA sequences. Specifically, they target the major rRNAs in human cytoplasm (28S, 18S, 5.8S, 5S) and mitochondrial rRNAs (16S, 12S). The purpose of hybridization is to allow these designed DNA probes to specifically bind to the abundant host ribosomal RNA in the total RNA of the sample, forming a stable DNA / RNA hybrid double strand. This allows subsequent steps to specifically target these hybrid strands, precisely removing host rRNA without affecting the non-complementary viral RNA.
[0031] Specifically, the hybridization reaction is performed in a dedicated probe binding buffer. During the procedure, the following reaction mixture is prepared in a nuclease-free PCR tube, gently pipetted to mix, and briefly centrifuged. The tube is then placed on a PCR instrument and the following program is run: First, denature the RNA and DNA probes at 95°C for 2 minutes to open the double-stranded structure, creating conditions for subsequent complementary binding. Then, the temperature is slowly reduced from 95°C to 22°C at a rate of 0.1°C per second. This slow cooling process allows the DNA probe sufficient time to find and form a precise, stable hybrid double strand with the complementary host rRNA sequence, minimizing non-specific binding. Finally, the tube is held at 22°C for 5 minutes to ensure complete hybridization. After the program is complete, the tube is briefly centrifuged and placed on ice before proceeding to the next step. The resulting hybridization reaction mixture typically contains total RNA from the sample, a DNA probe mixture, and a buffer with optimized ion concentrations. See Table 1 for specific composition.
[0032] Table 1 Examples of hybridization reaction systems
[0033] After hybridization, a large number of DNA / RNA hybrid strands are formed. RNase H is a ribonuclease whose characteristic is that it can specifically recognize and cleave the RNA portion in the DNA / RNA hybrid double strand, while being inactive with single-stranded RNA, single-stranded DNA, or double-stranded DNA. Utilizing this characteristic, we can cleave and degrade the host ribosomal RNA that has hybridized with the DNA probe, thereby achieving the purpose of removing host rRNA. Specifically, prepare the RNase H digestion reaction solution on ice in the reaction tube after hybridization, gently mix it with a pipette, and collect the sample to the bottom of the tube by brief centrifugation. Then place it in a PCR instrument and incubate at 37°C for 20-30 minutes, preferably 30 minutes. This temperature is commonly used for RNase H to exert its optimal activity. During incubation, the enzyme will efficiently cleave the rRNA in all hybrid strands, breaking them into small fragments.
[0034] After RNase H digestion, the host rRNA has been degraded, but there are still a large number of DNA probes that have been bound to the rRNA fragments and free DNA probes in the system. If these DNA probes are not removed, they may become non-template background in the subsequent library construction steps and interfere with library construction. Therefore, they need to be removed.
[0035] Specifically, prepare the DNase I digestion reaction solution in the reaction tube after RNase H digestion on ice, gently mix with a pipette, briefly centrifuge, and then place it in a PCR instrument and incubate at 37°C for 25-35 minutes. Preferably, incubate at 37°C for 30 minutes to fully degrade all DNA probes.
[0036] It should be noted that after the enzymatic digestion described above, the reaction system is filled with degraded rRNA fragments, DNA probe fragments, enzyme proteins, various salt ions, and buffer components. These impurities need to be removed to obtain a clean solution rich in the target viral RNA for the next fragmentation step. This method uses RNA purification magnetic beads for purification. Specifically, a certain volume of magnetic bead suspension (to provide high-solution salt-binding conditions) is added to the digestion product. RNA will reversibly bind to the surface of the magnetic beads (usually silica-coated beads), while impurities such as proteins, salts, and small nucleic acid fragments remain in the solution. By mixing the reaction system with the magnetic beads and allowing it to stand, the magnetic beads are adsorbed under the action of an external magnetic field. The supernatant is discarded, and residual impurities are removed by washing with ethanol. Finally, the RNA bound to the magnetic beads is eluted with nuclease-free water.
[0037] In practice, add RNA purification magnetic bead suspension (e.g., 110 µL, approximately 2.2 times the volume of the digested product) to the DNase I digestion product, mix thoroughly with a pipette, and incubate on ice for 15 minutes to allow RNA binding. Then, place the tube on a magnetic rack and incubate for 5 minutes until the solution is clear. Carefully remove the supernatant. While still magnetically attached, add 200 µL of freshly prepared 80% ethanol along the tube wall, incubate for 30 seconds, and discard the ethanol. Repeat this ethanol washing step once. After brief centrifugation, return the tube to the magnetic rack, aspirate any remaining liquid, and air dry at room temperature for 5-10 minutes until the magnetic beads are dull but not cracked. Add 18 µL of fragmentation buffer to the dried magnetic beads, resuspend thoroughly, and incubate at room temperature for 2 minutes. Incubate again on a magnetic rack for 5 minutes until the solution is clear, then carefully transfer 16 µL of the supernatant to a new nuclease-free PCR tube. Through this purification process, the final product is a solution of enriched viral RNA with most of the host ribosomal RNA removed, laying the foundation for the subsequent construction of a high-quality sequencing library.
[0038] S102. The viral RNA is fragmented.
[0039] It should be noted that this step involves controlled fragmentation of the purified viral RNA to prepare short RNA fragments with suitable length distributions for subsequent high-throughput sequencing library construction. After obtaining complete or longer viral RNA fragments, they need to be randomly fragmented to a specific length range. This is because mainstream sequencing platforms have an optimal range for library insert fragment lengths. Fragmentation ensures efficient and stable cluster generation and sequencing operations. In addition, random fragmentation helps eliminate secondary structure or sequence bias of the template RNA, allowing subsequent sequencing reads to cover the entire viral genome more randomly and evenly, thus laying the foundation for high-precision whole-genome assembly and variant analysis.
[0040] Specifically, the obtained fragmented buffer supernatant containing viral RNA, which already contains the necessary Mg2+, is used. 2+ Ions are placed in a PCR instrument and incubated at 80-90℃ for 4-8 minutes, preferably at 85℃ for 6 minutes. This process utilizes the principle of high-temperature metal ion-catalyzed hydrolysis. Under high temperature and in the presence of divalent magnesium ions, the phosphodiester bonds of the RNA chain undergo random breakage. By strictly controlling the reaction temperature and time, the RNA template can be randomly fragmented to the target length range of 200-500 bp. After the reaction, the sample tubes are immediately centrifuged and placed on ice to quickly terminate the hydrolysis reaction and prevent excessive fragmentation.
[0041] Through this process, the original viral RNA is transformed into a series of short fragments with a concentrated length distribution, the ends of which are functional groups generated by hydrolysis, which can be directly used as templates for the next step of reverse transcription.
[0042] S103. The fragmented RNA is reverse transcribed to synthesize the first-strand cDNA. The second-strand cDNA is synthesized simultaneously in the same reaction system using the first-strand cDNA as a template. End repair and dA tail addition are performed to obtain double-stranded cDNA.
[0043] It should be noted that this step mainly includes two stages: first-strand cDNA synthesis and second-strand cDNA synthesis with simultaneous end modification. First, first-strand cDNA synthesis, i.e., reverse transcription, is performed. Using fragmented viral RNA as a template, random primers are added. These random primers are a mixture of oligonucleotides consisting of 6 to 9 random base sequences, capable of binding to multiple random sites on the RNA template. This achieves comprehensive initiation of highly fragmented, sequence-unknown viral RNA templates, suitable for situations in metavinomics where unknown or diverse viral sequences may exist. The reaction is carried out in an optimized reverse transcriptase system. First, it is incubated at 25°C for 10 minutes; this low-temperature annealing step facilitates full binding of the random primers to the template RNA. Then, it is incubated at 42°C for 15 minutes, the temperature at which reverse transcriptase exhibits optimal activity. The enzyme synthesizes complementary DNA strands along the template, forming an RNA / DNA hybrid double strand. Finally, it is incubated at 70°C for 15 minutes to inactivate the reverse transcriptase and disrupt the secondary structure of the RNA, preparing for the next reaction. In this way, the single-stranded viral RNA information is transformed into a more stable complementary DNA (cDNA) strand.
[0044] The second-strand cDNA is then synthesized and its ends modified. This step is performed continuously in the same reaction system, using the first-strand cDNA synthesized in the previous step as a template. DNA polymerase (usually an enzyme with strand displacement activity) is used to synthesize its complementary strand. Specifically, the reaction is first incubated at 16°C for 30 minutes. This relatively low temperature is beneficial for the continuity and fidelity of DNA polymerase, ensuring the complete synthesis of the second strand. Simultaneously or after the second strand synthesis, a specific enzyme mixture in the reaction system (usually containing T4 DNA polymerase, Klenow fragments, etc.) performs end-repair, smoothing out any uneven ends of the double-stranded cDNA, such as protrusions or depressions caused by RNA degradation or synthesis. Next, using the end-transferase activity of the Klenow fragment (3'→5'exo-), an additional adenine deoxynucleotide (dA) is uniformly added to the 3' end of the smoothed double-stranded cDNA, i.e., a dA tail is added. After the above reaction is completed, the system is incubated at 65°C for 15 minutes to inactivate the relevant enzyme activities. Through this process, a complete, straight double-stranded cDNA product with a dA protrusion at the end was finally obtained.
[0045] S104. The adapter with the dT tail is ligated to the double-stranded cDNA to obtain the ligation product.
[0046] It should be noted that after step S103, a single dA overhang tail has been uniformly added to the 3' end of the double-stranded cDNA, while the adapter used in this step has a dT tail (i.e. a single thymine deoxynucleotide overhang at the 3' end). This dA-dT tail design can achieve specific complementary pairing, which avoids adapter self-ligation or non-specific ligation, and ensures high efficiency of adapter-cDNA ligation.
[0047] Specifically, the ligation reaction is carried out in an optimized ligase premix containing high-fidelity DNA ligase, reaction buffer, and other auxiliary factors. This premix provides the optimal reaction environment for adapter ligation, ensuring ligation efficiency and fidelity. During the procedure, prepare the reaction solution in a PCR tube according to the preset system. Add the double-stranded cDNA product, the adapter with the dT tail, and the ligase premix. Gently pipette 10 times to mix thoroughly, then briefly centrifuge to collect the sample at the bottom of the tube. Incubate at 20°C for 15 minutes. This temperature is optimal for DNA ligase activity; the 15-minute incubation time ensures that the adapter fully binds to the cDNA ends and completes ligation, avoiding incomplete ligation that could lead to reduced library yield or decreased sequencing efficiency.
[0048] After the ligation reaction is complete, in addition to the target ligation product (double-stranded cDNA with adapter), the system still contains unbound free adapters, excess ligase, and buffer components. If these impurities are not removed, they will form adapter dimers and other non-target products during subsequent PCR amplification, consuming sequencing resources and interfering with sequencing result analysis. Therefore, magnetic bead purification is necessary immediately. Specifically, after vortexing the DNA purification magnetic beads, add 60 μL of the DNA purification magnetic beads to the ligation product, cap the tube, and incubate it on a low-speed micro-vortex mixer for 5 minutes to allow the target DNA to fully bind to the magnetic bead surface. Then, after a brief centrifugation, place the PCR tube on a magnetic rack for 5 minutes to allow the liquid to become clear. Carefully aspirate the supernatant. While maintaining the magnetic rack, add 200 μL of freshly prepared 80% anhydrous ethanol along the tube wall, let it stand for 30 seconds, then aspirate the supernatant. Repeat this ethanol washing step. Next, remove residual impurities; then, remove the PCR tube, briefly centrifuge, and place it back on the magnetic rack, aspirate any remaining liquid, and allow it to air dry for about 30 seconds (be careful not to let the aggregated magnetic beads crack); add 22 μL of nuclease-free water (NFW) to the dried magnetic beads to rinse them, gently tap the tube wall to resuspend the beads, and let it stand at room temperature for 2 minutes; finally, after a brief centrifugation, place the PCR tube back on the magnetic rack, wait for the solution to become clear, and carefully aspirate 20 μL of the supernatant to transfer it to a new nuclease-free PCR tube to obtain the purified ligation product.
[0049] S105. The ligation product is amplified by PCR and purified to obtain a metaviid sequencing library.
[0050] It should be noted that after adapter ligation and purification using S104, the concentration of the ligation product is low, and some fragments may not have been successfully ligated. Specific PCR amplification can achieve efficient enrichment of the target product. At the same time, the introduction of a tag sequence (index) to distinguish samples and primer binding sites adapted to the sequencing platform during the amplification process can improve sequencing efficiency and reduce experimental costs.
[0051] Specifically, the PCR amplification reaction is carried out in a high-fidelity amplification enzyme premix, which contains core components such as high-fidelity DNA polymerase, dNTPs, and reaction buffer. This premix ensures amplification efficiency while reducing base mismatches, guaranteeing the accuracy of the viral genome sequence and meeting the sequence fidelity requirements of whole genome assembly. During the procedure, the reaction solution is prepared in a nuclease-free PCR tube according to the preset system. The purified ligation product, tagged PCR primers, and amplification enzyme premix are added. The mixture is then pipetted or gently tapped 10 times to mix. The sample is collected at the bottom of the tube by brief centrifugation and then placed in a PCR instrument for amplification. The amplification program follows the conventional high-throughput sequencing library amplification logic: first, double-stranded DNA is denatured at high temperature to unwind; then, annealing is performed to allow the primers to specifically bind to the template; finally, extension is used to synthesize new strands. This cycle is repeated to enrich the target fragment.
[0052] It should be noted that the number of PCR amplification cycles needs to be precisely adjusted according to the amount of starting RNA: 17-18 cycles for 10-99 ng of starting RNA; 14-15 cycles for 100-499 ng; 12-13 cycles for 500-999 ng; and 10-11 cycles for ≥1 µg. This is because low starting RNA amounts require more cycles to achieve sufficient sequencing concentration, while high starting RNA amounts require fewer cycles to avoid over-amplification leading to sequence bias, adapter dimer accumulation, and increased base error rate.
[0053] After PCR amplification, in addition to the target sequencing library fragment, the system still contains unconsumed primers, dNTPs, polymerase, and non-specific amplification products (such as adapter dimers). These impurities can interfere with sequencing signals and reduce data quality, therefore, magnetic bead purification is necessary immediately. Specifically, after vortexing to purify the DNA beads, add 45 μL to the PCR product, gently tap the tube wall to mix, and incubate on a low-speed micro-vortex mixer for 5 minutes to allow the target DNA fragment to specifically bind to the magnetic bead surface. After brief centrifugation, place the PCR tube on a magnetic rack for adsorption. Once the liquid becomes clear, carefully aspirate the supernatant. While maintaining the magnetic adsorption state, add 200 μL of freshly prepared 80% anhydrous ethanol along the tube wall, let stand for 30 seconds, then aspirate the supernatant. Repeat this ethanol washing step once. Thoroughly remove any residual impurities; remove the PCR tube, centrifuge briefly, then place it back on the magnetic rack, aspirate any remaining liquid, and allow it to air dry for about 30 seconds (be careful not to let the magnetic beads crack); add 22 μL of nuclease-free water (NFW) to the dried magnetic beads to rinse them, gently tap the tube wall to resuspend the beads, and let it stand at room temperature for 2 minutes; after a brief centrifugation, place the PCR tube back on the magnetic rack, and wait for the solution to become clear. Carefully aspirate 20 μL of the supernatant and transfer it to a new nuclease-free PCR tube to obtain the purified metaviidome sequencing library.
[0054] The method provided in this embodiment targets clinical throat swab samples of novel coronavirus and influenza virus. It employs an enzymatic strategy of removing host ribosomal RNA via specific DNA probe hybridization combined with sequential digestion of RNase H and DNase I, and incorporates Mg... 2+A high-quality metaviromic sequencing library was successfully constructed using an optimized library preparation process that incorporates high-temperature controllable fragmentation, random primer reverse transcription, and simultaneous end repair and dA tail addition during second-strand synthesis. This method also employs a strategy of dynamically adjusting the number of PCR cycles based on the initial RNA input. Ultimately, this method enables one-time library preparation for clinical samples of respiratory pathogens such as the novel coronavirus and influenza virus, obtaining complete viral genome sequences with high coverage and low host background. This effectively solves the technical challenges of existing technologies that suffer from significant nucleic acid recovery losses and reliance on over-amplification, making it impossible to effectively obtain full-length viral sequences from low-load samples. Therefore, this provides an efficient and reliable technical means for accurate pathogen typing, mutation tracking, transmission chain analysis, and drug resistance monitoring.
[0055] Example 2
[0056] Corresponding to the aforementioned embodiment of a method for constructing a metaviromic sequencing library by enzymatic removal of host ribosomal RNA, this application also provides an embodiment of a kit for constructing a metaviromic sequencing library by enzymatic removal of host ribosomal RNA. Specifically, the kit includes:
[0057] Remove host ribosomal RNA components, including DNA probes targeting host ribosomal RNA, RNase H, DNase I, and their respective reaction buffers;
[0058] Library preparation components include reagents for RNA fragmentation, cDNA synthesis, adapter ligation, and PCR amplification;
[0059] And, purification components.
[0060] The library preparation components include: RNA fragmentation buffer, reverse transcription reagent, DNA ligase, adapter with dT tail, PCR primers with tagged sequences, and PCR amplification enzyme premix; the purification components include RNA purification magnetic beads and DNA purification magnetic beads.
[0061] Specifically, Table 2 shows the reagent kit components as illustrated in this application:
[0062] Table 2. Reagent Kit Components
[0063] Component names main components Nuclease-free Water Nuclease-free water RNA Clean Beads RNA purification magnetic beads Magnet Beads DNA purification magnetic beads Ribosomal RNA Probe Ribosomal RNA probe Probe Buffer Probe binding buffer RNase H Buffer RNAse digestion buffer RNase H RNA enzyme DNase I Buffer DNA digestion buffer DNase I DNA enzymes Fragment Buffer RNA fragmentation premixed reaction solution 1st Synthetic Master Mix One-chain synthesis premixed reaction solution 2nd Synthetic Master Mix Second chain synthesis premixed reaction liquid Ligase Master Mix Ligase premix DNA Adapter connector Hifi Amplification Mix amplification enzyme premix Index Primers with tagged sequences
[0064] To further verify the effectiveness and technical advantages of the enzymatic method for constructing metavinomic sequencing libraries by removing host ribosomal RNA provided in this application, the following experiments were conducted.
[0065] I. Clinical Sample Testing of Influenza Virus
[0066] Twenty clinical influenza virus pharyngeal swab samples were collected, and RNA was extracted and detected by real-time quantitative PCR (Real-time PCR). The Ct values ranged from 18.08 to 31.30, covering high, medium, and low viral load levels (see Table 3). The RNA from these samples was processed to remove host ribosomal RNA using the method provided in this application, and a metaviromic sequencing library was constructed. High-throughput sequencing was performed using a Zhenmai Biotechnology FASTAseq300 sequencer. Parameterized assembly and coverage analysis were performed on the sequencing data, and the results are shown in Table 4.
[0067] Table 3. RT-PCR detection results of influenza virus clinical samples
[0068] Table 4. Genomic coverage statistics of influenza virus clinical sample sequencing libraries
[0069] Sample number Data output (Gb) Total number of sequences Target sequence number 10× Coverage 50× coverage 100× Coverage Flu1 2.2 14751918 9778 99.89% 92.70% 22.77% Flu2 2.3 15775286 8505 99.57% 77.95% 11.48% Flu3 2.1 14285018 14302 99.39% 92.88% 68.73% Flu4 5.6 37933906 283116 99.97% 99.93% 99.92% Flu5 2.1 14030492 11334 99.82% 86.80% 50.00% Flu6 4.3 29116146 19160 99.93% 98.29% 92.78% Flu7 4.3 29132994 30242 99.93% 99.65% 97.86% Flu8 1.7 11603414 43371 99.93% 99.65% 98.66% Flu9 3.4 25454010 12393 99.79% 95.03% 57.79% Flu10 4.0 29991244 357943 99.93% 99.89% 99.84% Flu11 2.9 21422294 192743 99.95% 99.89% 99.82% Flu12 2.4 19069990 3241895 100.00% 99.96% 99.94% Flu13 2.8 20552240 625936 99.96% 99.93% 99.87% Flu14 5.0 38659872 6314875 100.00% 99.98% 99.97% Flu15 5.6 42897578 180184 99.90% 99.78% 99.63% Flu16 5.4 41530070 416452 99.95% 99.88% 99.81% Flu17 2.4 18570526 173373 99.93% 99.74% 99.28% Flu18 5.7 44353930 18818 99.18% 92.20% 71.70% Flu19 3.3 25261410 27265 99.34% 92.80% 84.40% Flu20 4.1 31359710 29505 99.75% 95.74% 88.87%
[0070] Viral genome sequences were successfully obtained from all 20 samples, with 10× coverage exceeding 99%. High-load samples (e.g., Flu12, Ct=18.08) achieved 99.94% 100× coverage. Even for low-load samples (Flu2) with a Ct value as high as 31.30, the 10× viral genome coverage still reached 99.57%, demonstrating the excellent detection sensitivity and genome coverage capability of this method for low-load clinical samples. Figure 2 For the depth distribution map of the influenza virus whole genome sequence shown in this application, please refer to... Figure 2 A random selection of an influenza virus sample showed that the genome-wide coverage depth distribution map of the sequencing reads uniformly covered the entire viral genome, with no obvious gaps or biases.
[0071] II. Clinical Sample Testing for Novel Coronavirus
[0072] Twenty clinical pharyngeal swab samples of novel coronavirus were collected, and RNA was extracted and detected by real-time quantitative PCR. The Ct values of the ORF1ab gene ranged from 19.57 to 30.34 (see Table 5). The library was constructed and sequenced using the method described in this application, and the assembly coverage statistics are shown in Table 6.
[0073] Table 5. RT-PCR detection results of clinical samples of novel coronavirus
[0074] Sample number Sample types RT-PCR results nCOV1 Throat swab Novel coronavirus ORF1ab: 23.54; Novel coronavirus N: 24.20 nCOV2 Throat swab Novel coronavirus ORF1ab: 25.94; Novel coronavirus N: 25.81 nCOV3 Throat swab Novel coronavirus ORF1ab: 24.37; Novel coronavirus N: 24.81 nCOV4 Throat swab Novel coronavirus ORF1ab: 19.57; Novel coronavirus N: 20.36 nCOV5 Throat swab Novel coronavirus ORF1ab: 22.28; Novel coronavirus N: 21.95 nCOV6 Throat swab Novel coronavirus ORF1ab: 23.14; Novel coronavirus N: 25.59 nCOV7 Throat swab Novel coronavirus ORF1ab: 30.34; Novel coronavirus N: 32.30 nCOV8 Throat swab Novel coronavirus ORF1ab: 26.72; Novel coronavirus N: 28.39 nCOV9 Throat swab Novel coronavirus ORF1ab: 27.76; Novel coronavirus N: 29.67 nCOV10 Throat swab Novel coronavirus ORF1ab: 23.96; Novel coronavirus N: 23.64 nCOV11 Throat swab Novel coronavirus ORF1ab: 24.26; Novel coronavirus N: 24.81 nCOV12 Throat swab Novel coronavirus ORF1ab: 25.46; Novel coronavirus N: 24.46 nCOV13 Throat swab Novel coronavirus ORF1ab: 21.33; Novel coronavirus N: 21.66 nCOV14 Throat swab Novel coronavirus ORF1ab: 22.00, Novel coronavirus N: 21.45 nCOV15 Throat swab Novel coronavirus ORF1ab: 23.00, Novel coronavirus N: 23.57 nCOV16 Throat swab Novel coronavirus ORF1ab: 24.88, novel coronavirus N: 25.54 nCOV17 Throat swab Novel coronavirus ORF1ab: 26.12, novel coronavirus N: 26.24 nCOV18 Throat swab Novel coronavirus ORF1ab: 25.86, novel coronavirus N: 25.71 nCOV19 Throat swab Novel coronavirus ORF1ab: 24.50, novel coronavirus N: 25.20 nCOV20 Throat swab Novel coronavirus ORF1ab: 21.37, novel coronavirus N: 22.02
[0075] Table 6. Genomic coverage statistics of sequencing libraries of clinical samples of novel coronavirus.
[0076] Sample number Data output (Gb) Total number of sequences Target sequence number 10× Coverage 50× coverage 100× Coverage nCOV1 1.7 18843982 473633 99.55% 99.43% 99.40% nCOV2 2.8 30386858 62628 99.42% 98.40% 89.34% nCOV3 1.9 20712830 84491 99.45% 98.51% 93.95% nCOV4 3.9 43129164 2451564 99.65% 99.58% 99.51% nCOV5 3.7 40444392 445767 99.54% 99.45% 99.37% nCOV6 1.8 19945616 612991 99.54% 99.42% 99.37% nCOV7 2.1 22985090 27794 99.04% 79.36% 29.11% nCOV8 1.4 15628004 120142 99.49% 99.17% 98.49% nCOV9 1.9 20134804 231090 99.46% 99.24% 99.08% nCOV10 5.2 56001174 322640 99.51% 99.41% 99.36% nCOV11 4.6 49418780 926049 99.55% 99.49% 99.42% nCOV12 5.1 54779432 139889 99.42% 99.22% 98.71% nCOV13 3.5 37646010 4263779 99.65% 99.56% 99.54% nCOV14 3.2 35176784 375167 99.54% 99.43% 99.34% nCOV15 3.5 38934190 130330 99.49% 99.11% 96.75% nCOV16 4.0 43913928 40422 99.37% 93.50% 54.83% nCOV17 1.7 18079462 50925 99.40% 96.79% 80.29% nCOV18 1.9 20537782 58753 99.42% 98.00% 86.07% nCOV19 3.6 39076338 57199 99.42% 98.33% 87.45% nCOV20 1.7 18755640 633344 99.59% 99.47% 99.41%
[0077] After processing with this method, all novel coronavirus samples achieved a 10× genome coverage of over 99%, with most samples exceeding 99% 100× coverage. For example, the 100× coverage reached 99.40% for samples with a Ct value of approximately 23 (nCOV1); even for low-load samples with a Ct value as high as 32.30 (nCOV7), a 10× coverage of 99.04% was still obtained. This result further confirms that this method can effectively overcome high host background interference and stably obtain nearly complete viral genome sequences from clinical samples with different viral loads. Figure 3 For a depth distribution map of the whole genome sequence of a novel coronavirus shown in this application, please refer to... Figure 3 The coverage depth distribution map of a randomly selected novel coronavirus sample also shows uniform and continuous coverage.
[0078] III. Verification Test of Ribosomal RNA Removal Efficacy
[0079] To directly evaluate the effectiveness of the enzymatic removal of host ribosomal RNA in this method, five influenza virus samples and five novel coronavirus clinical samples were selected for comparative experiments. Two treatments were applied to the same sample: (A) libraries were constructed using the method described in this application (after the host ribosomal RNA removal step); (B) libraries were constructed directly without the host ribosomal RNA removal step. Both sets of libraries were sequenced in equal quantities, and the data were aligned to the human genome. The proportion of host sequences was calculated, and the results are shown in Table 7.
[0080] Table 7 Comparison of the effects of enzymatic removal of host ribosomal RNA steps
[0081] Test methods Sample number Specimen types Data output (Gb) Total number of sequences Number of host sequences Host ratio ribosomal RNA Flu21 Throat swab 1.9 12700000 827471 6.52% non-ribosomal RNA Flu21 Throat swab 1.9 12700000 9220605 72.60% ribosomal RNA Flu22 Throat swab 2.1 14200000 2167929 15.27% non-ribosomal RNA Flu22 Throat swab 2.1 14200000 12243560 86.22% ribosomal RNA Flu23 Throat swab 5.6 37900000 1851190 4.88% non-ribosomal RNA Flu23 Throat swab 5.6 37900000 26660010 70.34% ribosomal RNA Flu24 Throat swab 1.9 13000000 942754 7.25% non-ribosomal RNA Flu24 Throat swab 1.9 13000000 9247997 71.14% ribosomal RNA Flu25 Throat swab 1.7 11600000 416691 3.59% non-ribosomal RNA Flu25 Throat swab 1.7 11600000 8629748 74.39% ribosomal RNA nCOV21 Throat swab 1.7 18800000 2031005 10.80% non-ribosomal RNA nCOV21 Throat swab 1.7 18800000 15364394 81.73% ribosomal RNA nCOV22 Throat swab 2.8 30300000 3488135 11.51% non-ribosomal RNA nCOV22 Throat swab 2.7 30300000 25835886 85.27% ribosomal RNA nCOV23 Throat swab 3.9 43100000 1403672 3.26% non-ribosomal RNA nCOV23 Throat swab 3.8 43100000 33281832 77.22% ribosomal RNA nCOV24 Throat swab 3.7 40400000 3260754 8.07% non-ribosomal RNA nCOV24 Throat swab 3.6 40400000 32483274 80.40% ribosomal RNA nCOV25 Throat swab 1.6 17700000 2834743 16.02% non-ribosomal RNA nCOV25 Throat swab 1.6 17700000 14250276 80.51%
[0082] Experimental results show that after treatment with host ribosomal RNA using this method, the proportion of host sequences in the sequencing data of all test samples was significantly reduced, decreasing from 70.34%–86.22% to 3.26%–16.02%. This demonstrates that the DNA probe hybridization combined with RNase H / DNase I digestion strategy employed in this application can efficiently and specifically remove the vast majority of host ribosomal RNA from samples, thereby increasing the proportion of effective data for virus analysis in the sequencing data several times over, greatly reducing the waste of sequencing data, and improving the economy and efficiency of detection.
[0083] In summary, the above series of experiments fully validated that the library construction method and kit provided in this application can efficiently remove host background (significantly reducing the proportion of host sequences from over 70% to below 20%), and can achieve high-coverage sequencing for clinical samples with different viral loads. For samples with Ct values not exceeding 32, a viral genome coverage of over 99% (10×) can be obtained. Even for low-load samples with high Ct values, high-quality whole-genome sequences can be effectively obtained. In addition, this method is both broad-spectrum and practical. It does not require the pre-design of virus-specific probes or primers, and can simultaneously achieve high-sensitivity detection and whole-genome assembly of multiple respiratory viruses such as novel coronavirus and influenza virus, thus providing a powerful tool for the identification of unknown pathogens and in-depth variation analysis of known pathogens.
[0084] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for constructing a metavinomic sequencing library by enzymatically removing host ribosomal RNA, characterized in that, The method is used to obtain the full-length viral genome sequence from respiratory clinical samples with a Ct value not higher than 32 as detected by real-time quantitative PCR. The method includes: The total RNA of the respiratory clinical sample was hybridized with a DNA probe targeting host ribosomal RNA to form a DNA / RNA hybrid strand; before reverse transcription, the RNA in the hybrid strand was digested with RNase H, and the remaining DNA probe in the system was digested with DNase I; after purification, enriched viral RNA was obtained. The viral RNA was fragmented. The fragmented RNA was reverse transcribed to synthesize the first-strand cDNA. In the same reaction system, the second-strand cDNA was synthesized simultaneously using the first-strand cDNA as a template. End repair and dA tail addition were performed to obtain double-stranded cDNA. The adapter with the dT tail is ligated to the double-stranded cDNA to obtain the ligation product; The ligation product was amplified by PCR and purified to obtain a metaviromic sequencing library.
2. The method according to claim 1, characterized in that, The host is human, and the DNA probe targets at least one of human cytoplasmic 28S, 18S, 5.8S, 5S ribosomal RNA and mitochondrial 16S, 12S ribosomal RNA.
3. The method according to claim 1, characterized in that, The hybridization conditions are as follows: denaturation at 95°C for 2 minutes, followed by cooling to 22°C at a rate of 0.1°C / second, and holding at 22°C for 5 minutes; the RNase H digestion conditions are incubation at 37°C for 20-30 minutes; and the DNase I digestion conditions are incubation at 37°C for 25-35 minutes.
4. The method according to claim 1, characterized in that, The step of fragmenting the viral RNA includes: In Mg 2+ Incubate in buffer at 80-90℃ for 4-8 minutes.
5. The method according to claim 1, characterized in that, The process of reverse transcribing fragmented RNA to synthesize first-strand cDNA includes: The reaction was carried out using random primers at 25°C for 10 minutes, 42°C for 15 minutes, and 70°C for 15 minutes.
6. The method according to claim 1, characterized in that, The number of PCR amplification cycles is determined based on the amount of starting RNA: 17-18 cycles for 10-99 ng of starting RNA; 14-15 cycles for 100-499 ng; 12-13 cycles for 500-999 ng; and 10-11 cycles for ≥1 µg.
7. The method according to claim 1, characterized in that, The conditions for synthesizing the second-strand cDNA were: incubation at 16°C for 30 minutes, followed by incubation at 65°C for 15 minutes.
8. The method according to claim 1, characterized in that, The conditions for connector connection are: incubation at 20°C for 15 minutes; the purification is performed using the magnetic bead method, and the virus includes novel coronavirus and / or influenza virus.
9. A kit for constructing a metavinomic sequencing library by enzymatically removing host ribosomal RNA, characterized in that, The kit is prepared based on the method described in any one of claims 1-8, and the kit comprises: Remove host ribosomal RNA components, including DNA probes targeting host ribosomal RNA, RNase H, DNase I, and their respective reaction buffers; Library preparation components include reagents for RNA fragmentation, cDNA synthesis, adapter ligation, and PCR amplification; And, purification components.
10. The reagent kit according to claim 9, characterized in that, The library preparation components include: RNA fragmentation buffer, reverse transcription reagent, DNA ligase, adapter with dT tail, PCR primers with tagged sequences, and PCR amplification enzyme premix; the purification components include RNA purification magnetic beads and DNA purification magnetic beads.
Citation Information
Patent Citations
Optimized method for constructing RNA high-throughput sequencing library and application of optimized method
CN107385018A
Probe composition for removing rRNA, library building kit and library building method
CN115976163A
Method for improving RNA virus detection rate
CN116622806A
Method for constructing micropterus salmoides virus library by using virus metagenomics
CN117757897A
Cited By
A method for analyzing gene expression of plant-pathogen interaction based on umi tag
CN122168738A