Full-length 16S rRNA gene sequencing method based on UMI
By using UMI tags in 16S rRNA gene sequencing, the problems of primer bias and PCR amplification bias in the prior art are solved, and high accuracy and low cost full-length 16S rRNA gene sequencing are achieved, meeting the needs of high resolution classification and quantitative evaluation.
Patent Information
- Application Number
- CN202410068119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing 16S rRNA gene sequencing methods have primer bias and PCR amplification bias, resulting in accuracy and quantitative errors, which are difficult to meet the needs of high-resolution phylogenetic classification and accurate relative abundance evaluation.
Full-length 16S rRNA genes were labeled during PCR using Unique Molecular Identifier (UMI) tags, sequenced by a long-read sequencer, and data analysis was performed to correct errors and clustering to obtain consistent sequences and relative abundance.
Improve sequencing accuracy to Q30/Q40, eliminates PCR amplification bias, provides high-resolution phylogenetic classification and accurate relative abundance evaluation, reducing costs.
Smart Images

Figure CN120330307A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology. Specifically, the present invention relates to a method for sequencing full-length 16S rRNA genes based on UMI. More specifically, the present invention relates to a sequencing method for full-length 16S rRNA genes using UMI and a long-read sequencer, which has high accuracy, high economy, and eliminates PCR amplification bias. Background Art
[0002] Taxonomy is a scientific discipline that involves the characterization, classification, and naming of biological entities and is crucial for a comprehensive understanding of the diversity of the biosphere. As the most classical marker gene for determining the microbial profile at present, the 16S ribosomal RNA (16S rRNA) gene is the gold standard and a first-line tool for evaluating the taxonomic status of communities. Carl Woese first proposed the universal molecular identifier [1] , the 16S rRNA gene exists in archaeal and bacterial genomes, with highly conserved regions and multiple variable regions containing sufficient taxonomic information. With the development of polymerase chain reaction (PCR) technology and next-generation sequencing (NGS) technology in the early 21st century, the 16S rRNA gene has been incorporated into comprehensive and quality-controlled databases and reliable taxonomic information. So far, the publicly available 16S rRNA sequence databases are growing rapidly, and the current number of entries has exceeded 500,000 [2] .
[0003] High-throughput sequencing (HTS) technology enables the understanding of valuable biodiversity and valuable genetic information from various environments, thus greatly enriching the catalog of microbial species and indeed expanding the tree of life. [3 . Currently, sequencing data of 16S rRNA genes are mainly generated by high-throughput short-read amplicon sequencing. However, this method has several disadvantages. First, amplicons have primer bias, which may lead to the absence of a large amount of known microbial diversity [4 . Second, short-read sequences are of high resolution. Due to the limited information provided by short-read regions, they cannot support phylogenetic classification 5] . Compared with short-read sequencing, long-read sequencing can provide high-resolution phylogenetic classification and detect base modifications, so it is becoming increasingly attractive in microbiology [6-8]With the innovation of long-read sequencing such as Oxford Nanopore Technologies (Nanopore) or Pacific Biosciences (PacBio), circular consensus sequencing has been applied to detect 16S rRNA genes in different ecosystems, such as the human gut microbiota. [7] Pathogens [6,9] Marine microbiota
[10] Wastewater
[11] And digested sludge [8] Currently, long-read sequencers have achieved an accuracy of Q20 for most sequences (>50%), but for various research purposes, error-free sequences with higher accuracy are still required.
[0004] In addition to the accuracy of the sequence, the quantitative error caused by PCR amplification bias is also an important issue affecting the results of scientific problems.
[12] Microbiome research mainly uses the 16S rRNA gene to evaluate the relative abundance of microbial taxa in a community.
[13] However, these measurements do not accurately reflect the absolute taxon concentration because 16S rRNA gene sequencing must undergo PCR amplification to meet the DNA quantity requirements of the sequencer, and PCR amplification can cause quantity bias of DNA templates due to various factors such as annealing temperature, resulting in a large difference between the relative quantification results of PCR products and the original data.
[14]
[0005] Therefore, developing a full-length 16S rRNA gene sequencing method with high accuracy, low cost, and capable of eliminating PCR amplification bias is an urgent problem to be solved currently. Summary of the Invention
[0006] The present invention provides a full-length 16S rRNA gene sequencing method based on UMI tags, and applies this method to a long-read sequencer, thereby verifying the following two points: (1) The method of the present invention improves the accuracy (i.e., accuracy rate) of the final consensus sequence, enabling the wide application of long-read sequencing technology; (2) The method of the present invention obtains the correct relative abundance by eliminating the amplification bias in the PCR process of the 16S rRNA gene.
[0007] A Unique Molecular Identifier (UMI) is a molecular barcode that is unique for each target DNA fragment. In combination with the corresponding downstream bioinformatics analysis, UMI can provide error correction to improve accuracy and eliminate PCR bias to provide correct relative abundances. These two features enable the wide application of long-read sequencing. Therefore, the present invention proposes a UMI-based full-length 16S rRNA gene sequencing method for long-read sequencers. In this method, UMI is added to the target fragment, i.e., the full-length 16S rRNA gene, through a PCR process, and then the UMI-labeled full-length 16S rRNA gene is further amplified to meet the sequencing requirements of the long-read sequencer. Finally, by clustering and correcting sequences with the same UMI, consensus sequences with corrected relative abundances are obtained.
[0008] Specifically, the object of the present invention is achieved by the following technical solutions:
[0009] A UMI-based full-length 16S rRNA gene sequencing method, comprising the following steps:
[0010] Step 1, extracting genomic DNA of a sample to be tested;
[0011] Step 2, ligating UMI to the genomic DNA through a PCR reaction and constructing a DNA library, and then sequencing the DNA library using a long-read sequencer to obtain sequencing data;
[0012] Step 3, performing data analysis on the sequencing data to determine consensus sequences;
[0013] Step 4, determining the classification information of the consensus sequences and / or calculating the relative abundance of each consensus sequence.
[0014] In some embodiments of the method according to the present invention, in Step 1, the sample to be tested is selected from human samples, animal samples, and environmental samples, such as oral samples of humans and animals, intestinal flora samples of humans and animals, soil samples, sludge samples, pathogen samples, marine microbial flora samples, volcanic samples, mine samples, and wastewater samples.
[0015] In some embodiments of the method according to the present invention, in Step 2, the ligating UMI to the genomic DNA through a PCR reaction and constructing a DNA library includes a first-round PCR reaction, a second-round PCR reaction, and an optional third-round PCR reaction performed in sequence.
[0016] In some embodiments of the method according to the present invention, the primers used in the first-round PCR reaction include a UMI forward primer and a UMI reverse primer. The UMI forward primer is CAAGCAGAAGACGGCATACGAGATNNNYRNNNYRNNNYRNNNAGRGTTYGATYMTGGCTCAG (SEQ ID NO: 1), and the UMI reverse primer is AATGATACGGCGACCACCGAGATCNNNYRNNNYRNNNYRNNNRGYTACCTTGTTACGACTT (SEQ ID NO: 2), where Y represents the base C or T, R represents the base A or G, M represents the base A or C, and N represents one of the bases A, C, G, and T.
[0017] In some specific embodiments of the method according to the present invention, the reaction system of the first-round PCR reaction contains 5 ng of genomic DNA, 0.5 U of Platinum Taq DNA high-fidelity polymerase, 1× high-fidelity PCR buffer, 100 μM of each dNTP, 1.5 mM of MgSO4, and 500 nM of each of the UMI forward primer and the UMI reverse primer, and the final reaction volume is 50 μL.
[0018] In some preferred embodiments of the method according to the present invention, the reaction conditions of the first-round PCR reaction are:
[0019]
[0020] In some embodiments of the method according to the present invention, the primers used in the second-round PCR reaction include a synthesized forward primer and a synthesized reverse primer. The synthesized forward primer is CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 3), and the synthesized reverse primer is AATGATACGGCGACCACCGAGATC (SEQ ID NO: 4).
[0021] In some specific embodiments of the method according to the present invention, the reaction system of the second-round PCR reaction contains all the product DNA obtained from the first-round PCR, 0.5 U of Platinum Taq DNA high-fidelity polymerase, 1× high-fidelity PCR buffer, 100 μM of each dNTP, 1.5 mM of MgSO4, 500 nM of each of the synthesized forward primer and the synthesized reverse primer, and the final reaction volume is 50 μL.
[0022] In some preferred embodiments of the method according to the present invention, the reaction conditions of the second-round PCR reaction are:
[0023]
[0024] In some embodiments of the method according to the present invention, the reaction system, reaction conditions, and primers of the third-round PCR are the same as those of the second-round PCR, except that the denaturation, annealing, and extension steps are carried out for 2 or 3 cycles.
[0025] In some embodiments of the method according to the present invention, step 3 includes the following steps:
[0026] Screen the sequencing data to obtain sequences with lengths between 1300 bp and 1700 bp;
[0027] Extract and screen the UMIs of each obtained sequence to retain the target sequences with UMIs of the correct length, pattern, and sequence;
[0028] Perform clustering analysis on the target sequences, and cluster the target sequences with the same UMI group in the same cluster;
[0029] Align, correct and optimize, remove chimeras, and trim the sequences in each cluster to obtain consensus sequences.
[0030] In some embodiments of the method according to the present invention, when aligning the sequences in each cluster, a coverage rate > 30 is used as the cut-off value.
[0031] In some embodiments of the method according to the present invention, in step 4, the relative abundance of each consensus sequence is calculated according to the number of UMIs.
[0032] As can be seen from the above technical solutions, the full-length 16S rRNA gene sequencing method based on UMI provided by the present invention has at least the following beneficial effects:
[0033] Compared with the existing 16S rRNA gene sequencing methods, the full-length 16S rRNA gene sequencing method based on UMI proposed by the present invention adds UMI tags before the traditional PCR process, thereby improving the accuracy of long-read sequencing to Q30 / Q40 (99.9% / 99.99%), and obtaining error-free full-length 16S rRNA genes. And it also eliminates the quantitative bias caused by PCR amplification and corrects the relative abundance of the final consensus sequences of the full-length 16S rRNA genes.
[0034] In addition, the method of the present invention has economy, and its cost is at least 50% lower than other methods on the market. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings, wherein:
[0036] Figure 1Shows the complete flow chart of the UMI-based full-length 16S rRNA gene sequencing method of the present invention, where A and B show the PCR process, and C shows the composition of the UMI-labeled 16S rRNA gene.
[0037] Figure 2 Shows the sequence processing process of the UMI-based Nanopore 16S rRNA gene sequencing of the present invention.
[0038] Figure 3 Shows the principle of correcting sequence errors by UMI of the present invention.
[0039] Figure 4 Shows the principle of correcting PCR amplification bias by UMI of the present invention. The upper process performs counting analysis by the number of sequences, and the lower process uses UMI for quantitative analysis.
[0040] Figure 5 Shows the gel electrophoresis photos of PCR products at different annealing temperatures.
[0041] Figure 6 Shows the similarity and coverage information of all consensus sequences. The X-axis is the coverage information of the cluster where the sequence is located, and the Y-axis is the similarity result between the consensus sequence and the reference genome. The left figure shows the similarity and coverage information of all consensus sequences, and the right figure shows the consensus sequences with the coverage range of the corresponding cluster between 0 - 100.
[0042] Figure 7 Shows the similarity and coverage information of all consensus sequences of each bacterium in the mock community. The X-axis is the coverage information of the cluster where the sequence is located. The Y-axis is the similarity result between the consensus sequence and the reference genome. Detailed implementation mode
[0043] The present invention is further described below in conjunction with embodiments. It should be understood that the embodiments are only used to further illustrate and explain the present invention and are not used to limit the present invention.
[0044] The experimental methods used in the following embodiments are all conventional experimental methods in the art unless otherwise specified. The experimental materials used in the following embodiments are all purchased from biochemical reagent sales companies unless otherwise specified.
[0045] Embodiment 1
[0046] 1 Materials and methods
[0047] 1.1 Samples and DNA extraction
[0048] In this example, three types of samples were used: Escherichia coli (E. coli), mock community, and two enriched samples. The E. coli strain K-12 was purchased from the Global Bioresource Center (ATCC, USA), and its NCBI genome ID is NC_000913.3. The ZymoBIOMICS Microbial Community DNA Standard (D6305) was obtained from ZymoResearch (Irvine, California), which contains 10 species, including 8 bacteria and 2 yeasts (Table 1). Among them, the PCR amplification of the 16S rRNA gene does not target these two yeasts. Finally, two enriched samples taken from the following reactors were used in the following experiments: a partial nitrification-anaerobic ammonium oxidation (PNA) reactor and an anaerobic ammonium oxidation upflow anaerobic sludge bed (Anammox-UASB) reactor that had been operating stably in the inventor's laboratory for more than three years.
[0049] According to the manufacturer's protocol, the FastDNA SPIN Kit for Soil (MP Biomedicals, France) was used to extract the genomic DNA of the above samples. The concentration and purity of the extracted DNA were quantified using a Qubit 2.0 fluorometer (Invitrogen, USA) and verified by a Nanodrop UV / Vis spectrophotometer (ND-1000, Nanodrop Technologies, USA).
[0050] Table 1: ZymoBIOMICS TM Microbial mock community simulating D6305
[0051]
[0052]
[0053] 1.2 PCR using UMI, sequencing library preparation, and long-read sequencing on Nanopore
[0054] UMI is a short sequence that can identify the input DNA molecules in the bioinformatics step and reduce the quantitative bias introduced by PCR amplification 16,17]Based on the inventors' research, conventional UMI quantification is generally used for second-generation RNA sequencing, with a sequence of 4-6 N bases, namely "NNNN", "NNNNN", or "NNNNNN". The method of the present invention applies UMI quantification to correct the amplification bias caused by DNA PCR amplification, and improves the UMI sequence according to the sequencing characteristics of third-generation technology, which has a high error rate and is prone to sequencing deletion of consecutive identical base sequences (such as "AAAAAA"). The UMI sequence is changed to "NNNYRNNNYRNNNYRNNN", thus effectively eliminating the quantification bias caused by PCR amplification.
[0055] Figure 1 Shows the complete flow chart of UMI-based full-length 16S rRNA gene sequencing.
[0056] The PCR process of UMI contains two or three (optional) cycles. In the first cycle of PCR, according to the instructions in Table 2, 5 ng of genomic DNA is used, and the UMI is labeled onto the DNA fragment with the UMI primer. The UMI primer is designed to have three parts: a synthetic primer site, UMI, and a 16S rRNA universal primer. The UMI forward primer is CAAGCAGAAGACGGCATACGAGATNNNYRNNNYRNNNYRNNNAGRGTTYGATYMTGGCTCAG (SEQ ID NO: 1), and the UMI reverse primer is AATGATACGGCGACCACCGAGATCNNNYRNNNYRNNNYRNNNRGYTACCTTGTTACGACTT (SEQ ID NO: 2). Among them, CAAGCAGAAGACGGCATACGAGAT and AATGATACGGCGACCACCGAGATC are synthetic primers, NNNYRNNNYRNNNYRNNN is UMI, and AGRGTTYGATYMTGGCTCAG and RGYTACCTTGTTACGACTT are 16S rRNA universal primers. Together, they can generate 1.2×10 18 possible UMI results (4 12×2 ×2 6×2 = 1.2×10 18)。The first-round PCR contained 5 ng of DNA, 0.5 U of Platinum Taq DNA High-Fidelity Polymerase (error rate: 0.003 - 0.005%, 6 times lower than Taq), 1× High-Fidelity PCR Buffer, 100 μM of each dNTP, 1.5 mM of MgSO4, and 500 nM each of the UMI forward primer and reverse primer, with a final reaction volume of 50 μL. The PCR program started with pre-denaturation (95 °C, 3 minutes), followed by 2 cycles of denaturation (95 °C, 15 seconds), annealing (51 °C, 30 seconds), and extension (72 °C, 2 minutes). Then, the PCR products were purified using AMPure XP magnetic beads (Beckman Coulter, USA) according to the manufacturer's instructions, with a bead / sample ratio of 0.6. The purified PCR products that had been ligated with the UMI sequences were used for the next round of PCR.
[0057] Table 2: Experimental conditions for the first-round PCR reaction of UMI
[0058]
[0059] As Figure 1 As shown in and Table 3, the second-round PCR used synthetic primers to amplify the UMI-tagged PCR products generated in the first-round PCR. The second-round PCR contained all the products amplified in the first-round PCR (i.e., the UMI-tagged PCR products), 0.5 U of Platinum Taq DNA High-Fidelity Polymerase, 1× High-Fidelity PCR Buffer, 100 μM of each dNTP, 1.5 mM of MgSO4, 500 nM each of the synthetic forward primer and reverse primer, with a final reaction volume of 50 μL. The PCR program started with pre-denaturation (95 °C, 3 minutes), followed by 25 cycles of denaturation (95 °C, 15 seconds), annealing (60 °C, 30 seconds), and extension (72 °C, 2 minutes), followed by one cycle of final extension (72 °C, 5 minutes). The PCR products were purified using AMPure XP magnetic beads (Beckman Coulter, USA), with a bead / sample ratio of 0.6.
[0060] Table 3: Experimental conditions for the second-round PCR reaction of UMI
[0061]
[0062]
[0063] To further amplify the DNA yield to meet the requirements of DNA library construction for sequencers, the third round of PCR is optional. The main role of the third round of PCR is to obtain a sufficient amount of PCR products for library construction and sequencing. If the amount of the second-round PCR products is sufficient for library construction and sequencing, the third round of PCR is not required. Otherwise, the third round of PCR is needed. The reaction conditions of the third round of PCR are exactly the same as those of the second round of PCR, except for the annealing of 2 or 3 cycles recommended based on experience. According to the manufacturer's instructions, the purified PCR products are used to prepare a sequencing library by the ligation method for long-read sequencing. For the Nanopore sequencer, after DNA library construction using the SQK-LSK109 kit, approximately 20 fmol of the DNA library is input into the flow cell of the Nanopore sequencer for 48-hour sequencing.
[0064] 1.3 Data Analysis of the UMI-based Full-length 16S rRNA Gene Sequencing Method
[0065] As Figure 2 shown, data analysis mainly includes quality filtration of sequences, UMI extraction and filtration, sequence clustering based on UMI, and generation of UMI consensus sequences. Nanopore sequencing generates raw reads in the fastq format with Q>7. First, the raw reads are screened to retain sequences with lengths between 1300 bp and 1700 bp. Then, the UMI of each sequence is extracted and filtered to retain UMI sequences with the correct length (18 bp) and sequence pattern (NNNYR). USEARCH clusters sequences with the same UMI group together
[18] . And align the sequences in the same cluster with Minimap2
[19] , and correct and optimize (polish) them with Racon
[20] and Medaka to generate the final consensus sequence. Chimeric detection and removal are performed through UCHIME2
[21] with the silva.gold database. Finally, the UMI and primer sequences of the consensus sequence are trimmed using Cutadapt to generate the final error-free sequence (clean sequence).
[0066] 1.4 Classification Annotation and Relative Abundance Calculation of Consensus Sequences
[0067] Using 16S rRNA databases such as the National Center for Biotechnology Information (NCBI)
[23] , the Genome Taxonomy Database (GTDB), EzBioCloud
[24] , Greengene
[25] and Silva
[26] , through BLAST
[22] , the taxonomic information of the final consensus sequence was determined. The relative abundance of each consensus sequence was calculated using the count of UMI numbers rather than the count of sequence numbers, and the specific calculation process is as Figure 4 shown.
[0068] 2 Results and Discussion
[0069] 2.1 Results of the ZymoBIOMICS mock community using the UMI-based full-length 16S rRNA gene sequencing method
[0070] UMIs can reduce the error rate of sequences with sufficient read coverage and provide consensus sequences with the same accuracy as shotgun sequencing
[27] . In the present invention, UMI primers with three parts were designed: a synthetic primer site, a UMI, and a 16S rRNA universal primer. The synthetic primer site and the 16S rRNA universal primer were designed to bind to the DNA template for PCR amplification. The UMI part achieved the functions of sequence error correction (i.e., error correction) and elimination of PCR amplification bias. Considering the randomness of sequencing errors, if there are sufficient sequences with the same UMI, the sequencing errors can be completely corrected to obtain the final consensus sequence.
[0071] As Figure 3 shown, there are three sequences with inconsistent sequencing base results at the first error site: sequence_1 is "G", and sequences_2 and sequence_3 are "A". Therefore, according to the principle of error correction, the correct sequence at this site should be "A". Using the same method, the correct sequencing bases at the other two sites are "T" and "G" respectively. Finally, with the help of the UMI, an error-free consensus sequence was obtained.
[0072] In addition, the UMI-based full-length 16S rRNA gene sequencing method can achieve absolute quantification by eliminating amplification bias during PCR, and the principle is as Figure 4 shown. For example, three sequences A, B, and C with a relative abundance ratio of 1:2:1 in the original community. During PCR, the amplification efficiencies of different sequences are different, which ultimately leads to an abundance bias between the sequencing results and the original community. For example, if the amplification efficiencies of the three sequences are B > A > C, then PCR can obtain an abundance ratio of A:B:C = 4:9:2. Once a sequence is linked to a UMI, the number of sequences with the same UMI should be counted as 1 because they all originate from the same initial template. Based on this principle, if the total amount of the sample is known, then the correct relative abundance ratio (A:B:C = 1:2:1) can be obtained, and even absolute quantification can be calculated.
[0073] 2.2 Determination of Annealing Temperature for the Escherichia coli - Based UMI - Full - Length 16S rRNA Gene Sequencing Method
[0074] The UMI - based full - length 16S rRNA gene sequencing method is a promising solution for accurate sequencing analysis of any system. The annealing temperature is a key parameter in the UMI labeling process, and it should be optimized because it affects the binding of primers to the DNA template and significantly influences the PCR yield. In the first PCR cycle, the gradient annealing temperature was set from 50.0 °C to 63.0 °C (A: 63.0 °C, B: 62.2 °C, C: 60.7 °C, D: 58.3 °C, E: 55.0 °C, F: 52.6 °C, G: 51.0 °C, H: 50.0 °C). In the second PCR cycle, the gradient annealing temperature was set from 50.0 °C to 62.0 °C (A: 62.0 °C, B: 61.3 °C, C: 59.9 °C, D: 57.7 °C, E: 54.6 °C, F: 52.4 °C, G: 50.9 °C, H: 50.0 °C). An electrophoresis gel process was performed to determine the PCR yield based on the concentration of the PCR products. According to Figure 5 , annealing temperatures of 51 °C and 60 °C were selected for the first and second rounds of PCR, respectively. After determining the reaction conditions, after 2 cycles of the first - round PCR and 25 cycles of the second - round PCR, the entire UMI process was carried out with Escherichia coli to obtain UMI - labeled PCR products.
[0075] 2.3 Results of the Consensus Sequences for the UMI - Based Full - Length 16S rRNA Gene Sequencing Method
[0076] After determining the experimental conditions of PCR, the UMI - labeling process was carried out using mock samples. The final PCR products were purified and used in the flow cell of a Nanopore sequencer to generate sequencing data. A total of 9303657 raw sequences were obtained for the UMI - based full - length 16S rRNA gene sequencing, and finally 21401 consensus sequences were obtained, with a data utilization rate of 45.7% (the number of reads was 4266461). Among the 21401 consensus sequences, only 1 sequence was identified as a chimera, with a chimeric rate of 0.004%. Then these 21401 consensus sequences were aligned with the reference genome by BLAST, Figure 6 and Figure 7The similarity results and corresponding coverage information of the total mock community and the corresponding individual bacteria were summarized respectively. Generally speaking, the average similarity of the consensus sequences to the reference genomes was basically above 99% (Q20), and within the coverage range of 0 - 20, the similarity continuously increased with the increase in the coverage rate. When the coverage rate > 20, the consensus sequences based on UMI-based full-length 16S rRNA gene Nanopore sequencing could reach an accuracy rate of 99.9% (Q30) on average. To ensure the quality and accuracy of the sequencing results, a coverage rate > 30 was finally selected as the cut-off value, and 17,830 (with a coverage rate > 30 in their original clusters) were obtained among 21,401 consensus sequences.
[0077] Once the coverage cut-off value of the clusters was determined, the sequencing results of each bacterial species in the mock community were summarized in Table 4, including BIN numbers, corresponding sequence numbers, etc. Two types of sequencing errors were counted in the present invention, namely mismatches and deletions. Mismatches mainly refer to the base inconsistencies between the consensus sequences and the reference genomes, which may be caused by nucleotide mutations during the PCR process or sequencing errors not corrected by UMI. Deletion errors refer to the insertions and deletions of nucleotides during Nanopore sequencing, which are related to the recognition ability of the Nanopore sequencer in homopolymer fragments. Generally speaking, the sequencing accuracy rate of the consensus sequences could reach Q30, but the sequencing accuracy rate of Escherichia coli was only Q20, and the mismatch error rate was 0 - 07%, much higher than that of other bacteria. In addition, the average mismatch and deletion error rates of the remaining 7 bacteria were 0.01% and 0.03% respectively. These results indicate that the UMI-based full-length 16S rRNA gene sequencing method has good performance in correcting mismatch errors, but needs to be optimized for homopolymer sequencing, even though corresponding arithmetic correction methods (such as Medaka) have been applied.
[0078]
[0079] 2.4 Results of UMI-based full-length 16S rRNA gene sequencing of real samples from PNA and Anammox-UASB reactors
[0080] Using the Nanopore Flowcell, two samples from the PNA and Anammox-UASB reactors were sequenced by the UMI-based full-length 16S rRNA gene sequencing method, and the results are summarized in Table 5. Generally speaking, the UMI-based full-length 16S rRNA gene sequencing method performed well, and the sequences of both samples could reach a Q30 sequencing accuracy rate. Among them, there were 139 consensus sequences in PNA and 896 consensus sequences in Anammox-UASB. The 16S rRNA genes of the main functional flora in the PNA reactor were recovered, such as Nitrosomonas, and denitrifying bacteria in Denitratisoma. In the Anammox-UASB system, the 16S rRNA genes of a large number of bacteria with the ability to degrade complex organic matter were captured, such as the genera OLB14 and Bryobacteraceae. In addition to the known taxa, the UMI-based full-length 16S rRNA gene sequencing method also revealed many new taxa in the system that were overlooked by other methods. For example, some species in the PNA system revealed by the UMI-based method were not detected by metagenomics
[28] , probably due to the failure of 16S rRNA gene assembly. In addition, the similarity of the 16S rRNA gene of a bacterium in the Anammox-UASB system to that of the known taxa was only 89.6%, but it matched well with the sequences of uncultured species collected from the Anammox system in the public database (similarity to KC238418.1 was 99% and coverage rate was 100%). All these results demonstrated the ability of the UMI-based full-length 16S rRNA gene sequencing method as a rapid sequencing method in discovering new species.
[0081] Table 5: Results of sequencing samples according to the UMI-based Nanopore Flowcell
[0082] PNA Anammox-UASB Number of sequences with Q > 7 127374 170659 <![CDATA Total number of clusters > <![CDATA 13272 > <![CDATA 10,049 > Number of clusters (coverage > 30) 139 896 Number of sequences in the cluster 7068 34941 Data utilization ratio 5.5% 20.5% Similarity to PacBio sequences 99.9%(Q30) 99.9%(Q30) Error rate of mismatches 0.01% 0.01% Error rate of deletions 0.06% 0.05%
[0083] 2.5 Comparison of different 16S rRNA gene sequencing methods
[0084] Currently, there are several high-throughput 16S rRNA gene sequencing methods based on short-read / long-read sequencing technologies (Table 6), including Illumina, Nanopore, and PacBio sequencers. Among different methods, according to the data of commercial companies, the main advantages of short-read Illumina sequencing are high accuracy (Q40) and low price ($1.07×10 per sequence) -2HKD), but it provides low species resolution and long turnaround time. Long-read sequencing by PacBio can also provide high accuracy, but it takes a long time, and it takes about six months to obtain sequencing results from commercial companies. Therefore, it cannot meet the needs of many studies, such as the real-time monitoring of bacterial communities. In addition, the cost of PacBio sequencing is also high. Nanopore sequencing is fast and can provide the basic composition of the whole community, but it is limited by the accuracy of the original sequence. UMI is a good solution that combines the advantages of high accuracy and short turnaround time, and can obtain sequences with higher accuracy for any sample, providing more convenience for future research. The method for full-length 16S rRNA gene sequencing based on UMI proposed in the present invention can provide error-free full-length 16S rRNA gene sequences with high accuracy (Q30 / Q40) to meet various research purposes. In addition, the method of the present invention is the only method that can eliminate PCR amplification bias to obtain correct relative abundance results, and the economic cost is 2.30×10 per sequence -4 HKD. Of course, in addition to the Nanopore sequencer, the method of the present invention can also be applied to the PacBio sequencer to further increase its accuracy and correct amplification bias.
[0085]
[0086] 3 Conclusions
[0087] The present invention designs a method for full-length 16S rRNA gene based on UMI for long-read sequencers and applies it to the Nanopore sequencing platform. The UMI primers of the method of the present invention are specifically designed to include three parts: full-length 16S rRNA gene primers, UMI primers, and synthesis primers. The present invention also evaluates PCR reaction conditions such as the annealing temperature of the primers, designs the overall experimental process of the method of the present invention, and determines the overall data analysis process of the original sequence to obtain the final consensus sequence. The results show that this method for full-length 16S rRNA gene based on UMI can increase the accuracy of long-read sequencing to Q30 / Q40 (99.9% / 99.99%), thereby obtaining error-free full-length 16S rRNA genes of reactor simulation samples and real samples. In addition, this method can eliminate the quantitative bias caused by PCR amplification, obtain the correct relative abundance of sequences, and is much cheaper than other methods.
[0088] References
[0089] 1. Woese, C.R., O. Kandler, and M.L. Wheelis, Towards a natural system of organisms: proposal for the domains Archaea, Bacteria, and Eucarya. Proceedings of the National Academy of Sciences, 1990. 87(12): p. 4576 - 4579.
[0090] 2. Yarza, P, et al., Uniting the classification of cultured and uncultured bacteria and archaea using 16S rRNA gene sequences. Nature Reviews Microbiology, 2014. 12(9): p. 635 - 645.
[0091] 3. Reuter, Jason A., D.V. Spacek, and Michael P. Snyder, High-Throughput Sequencing Technologies. Molecular Cell, 2015. 58(4): p. 586 - 597.
[0092] 4. Eloe-Fadrosh, E.A., N.N. Ivanova, T. Woyke, and N.C. Kyrpides, Metagenomics uncovers gaps in amplicon-based detection of microbial diversity. Nature Microbiology, 2016. 1(4): p. 15032.
[0093] 5. Callahan, B.J., et al., Ultra-accurate microbial amplicon sequencing with synthetic long reads. Microbiome, 2021. 9(1): p. 130.
[0094] 6. Leggett, R.M., et al., Rapid MinION profiling of preterm microbiota and antimicrobial-resistant pathogens. Nature Microbiology, 2020. 5(3): p. 430-442.
[0095] 7. Matsuo, Y., et al., Full-length 16S rRNA gene amplicon analysis of human gut microbiota using MinION TM nanopore sequencing confers species-level resolution. BMC Microbiology, 2021. 21(1): p. 35.
[0096] 8. Lam, T.Y.C., et al., Superior resolution characterisation of microbial diversity in anaerobic digesters using full-length 16S rRNA gene amplicon sequencing. Water Research, 2020. 178: p. 115815.
[0097] 9. Moon, J., et al., Rapid diagnosis of bacterial meningitis by nanopore 16S amplicon sequencing: A pilot study. International Journal of Medical Microbiology, 2019. 309(6): p. 151338.
[0098] 10. Curren, E., T. Yoshida, V.S. Kuwahara, and S.C.Y. Leong, Rapid profiling of tropical marine cyanobacterial communities. Regional Studies in Marine Science, 2019. 25.
[0099] 11. Numberger, D., et al., Characterization of bacterial communities in wastewater with enhanced taxonomic resolution by full-length 16S rRNA sequencing. Scientific Reports, 2019. 9(1): p. 9673.
[0100] 12. Jiang, S.-Q., et al., High-throughput absolute quantification sequencing reveals the effect of different fertilizer applications on bacterial community in a tomato cultivated coastal saline soil. Science of The Total Environment, 2019. 687: p. 601-609.
[0101] 13. Tettamanti Boshier, F.A., et al., Complementing 16S rRNA Gene Amplicon Sequencing with Total Bacterial Load To Infer Absolute Species Concentrations in the Vaginal Microbiome. mSystems, 2020. 5(2): p. e00777-19.
[0102] 14. Tkacz, A., M. Hortala, and P.S. Poole, Absolute quantitation of microbiota abundance in environmental samples. Microbiome, 2018. 6(1): p. 110.
[0103] 15. Yang, L., et al., Use of an improved high-throughput absolute abundance quantification method to characterize soil bacterial community and dynamics. Science of The Total Environment, 2018. 633: p. 360 - 371.
[0104] 16. Islam, S., et al., Quantitative single-cell RNA-seq with unique molecular identifiers. Nature Methods, 2014. 11(2): p. 163 - 166.
[0105] 17. Kivioja, T., et al., Counting absolute numbers of molecules using unique molecular identifiers. Nature Methods, 2012. 9(1): p. 72 - 74.
[0106] 18. Edgar, R.C., Search and clustering orders of magnitude faster than BLAST. Bioinformatics, 2010. 26(19): p. 2460 - 2461.
[0107] 19. Li, H., Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 2018. 34(18): p. 3094 - 3100.
[0108] 20. Vaser R., I. N. Nagarajan and M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome research, 2017. 27(5): p. 737 - 746.
[0109] 21. Edgar, R.C., UCHIME2: improved chimera prediction for amplicon sequencing. bioRxiv, 2016: p. 074252.
[0110] 22. Camacho, C., et al., BLAST+: architecture and applications. BMC Bioinformatics, 2009.10(1): p. 421.
[0111] 23. Sayers, E.W., et al., Database resources of the national center for biotechnology information. Nucleic Acids Research, 2022.50(D1): p. D20 - D26.
[0112] 24. Yoon, S.-H., et al., Introducing EzBioCloud: a taxonomically united database of 16S rRNA gene sequences and whole - genome assemblies. International Journal of Systematic and Evolutionary Microbiology, 2017.67(5): p. 1613 - 1617.
[0113] 25. DeSantis, T.Z., et al., Greengenes, a Chimera - Checked 16S rRNA Gene Database and Workbench Compatible with ARB. Applied and Environmental Microbiology, 2006.72(7): p. 5069 - 5072.
[0114] 26. Quast, C., et al., The SILVA ribosomal RNA gene database project: improved data processing and web - based tools. Nucleic Acids Research, 2013.41(D1): p. D590 - D596.
[0115] 27. Karst, S.M., et al., High-accuracy long-read amplicon sequences using unique molecular identifiers with Nanopore or PacBio sequencing. Nature Methods, 2021.
[0116] 28. Liu, L., et al., High-quality bacterial genomes of a partial-nitritation / anammox system by an iterative hybrid assembly method. Microbiome, 2020. 8(1): p. 155.
Claims
1. A full-length 16S rRNA gene sequencing method based on UMI, comprising the following steps: Step 1, extracting genomic DNA of a sample to be tested; Step 2, ligating UMI to the genomic DNA through a PCR reaction and constructing a DNA library, and then sequencing the DNA library using a long-read sequencer to obtain sequencing data; Step 3, performing data analysis on the sequencing data to determine a consensus sequence; Step 4, determining the classification information of the consensus sequence and / or calculating the relative abundance of each consensus sequence.
2. The full-length 16S rRNA gene sequencing method based on UMI according to claim 1, wherein In Step 1, the sample to be tested is selected from human samples, animal samples and environmental samples, such as oral samples of humans and animals, intestinal flora samples of humans and animals, soil samples, sludge samples, pathogen samples, marine microbial flora samples, volcanic samples, mine samples and wastewater samples.
3. The full-length 16S rRNA gene sequencing method based on UMI according to claim 1, characterized in that, In Step 2, the ligating UMI to the genomic DNA through a PCR reaction and constructing a DNA library includes a first-round PCR reaction, a second-round PCR reaction and an optional third-round PCR reaction carried out in sequence.
4. The full-length 16S rRNA gene sequencing method based on UMI according to claim 3, wherein The primers used in the first-round PCR reaction include a UMI forward primer and a UMI reverse primer. The UMI forward primer is CAAGCAGAAGACGGCATACGAGATNNNYRNNNYRNNNYRNNNAGRGTTYGATYMTGGCTCAG (SEQ ID NO: 1), and the UMI reverse primer is AATGATACGGCGACCACCGAGATCNNNYRNNNYRNNNYRNNNRGYTACCTTGTTACGACTT (SEQ ID NO: 2), where Y represents base C or T, R represents base A or G, M represents base A or C, and N represents one of bases A, C, G and T; Preferably, the reaction system of the first-round PCR reaction contains 5 ng of genomic DNA, 0.5 U of Platinum Taq DNA high-fidelity polymerase, 1× high-fidelity PCR buffer, 100 μM of each dNTP, 1.5 mM of MgSO4 and 500 nM of each of the UMI forward primer and the UMI reverse primer, and the final reaction volume is 50 μL; Preferably, the reaction conditions of the first-round PCR reaction are:
5. The full-length 16S rRNA gene sequencing method based on UMI according to claim 3, wherein The primers used in the second-round PCR reaction include a synthesized forward primer and a synthesized reverse primer. The synthesized forward primer is CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 3), and the synthesized reverse primer is AATGATACGGCGACCACCGAGATC (SEQ ID NO: 4); Preferably, the reaction system of the second round of PCR reaction contains all the product DNA obtained from the first round of PCR, 0.5 U of Platinum Taq DNA high-fidelity polymerase, 1× high-fidelity PCR buffer, 100 μM of each dNTP, 1.5 mM of MgSO4, 500 nM of each synthesized forward primer and synthesized reverse primer, and the final reaction volume is 50 μL; Preferably, the reaction conditions of the second round of PCR reaction are:
6. The full-length 16S rRNA gene sequencing method based on UMI according to claim 3, wherein, The reaction system, reaction conditions and primers of the third round of PCR are the same as those of the second round of PCR, except that the denaturation, annealing and extension steps are carried out for 2 or 3 cycles.
7. The full-length 16S rRNA gene sequencing method based on UMI according to claim 1, characterized in that, Step 3 includes the following steps: Screen the sequencing data to obtain sequences with lengths between 1300 bp and 1700 bp; Extract and screen the UMI of each obtained sequence to retain the target sequences with UMI of correct length, pattern and sequence; Perform cluster analysis on the target sequences, and cluster the target sequences with the same UMI group in the same cluster; Align, correct and optimize the sequences in each cluster, remove the chimeras therein, and trim to obtain a consensus sequence.
8. The full-length 16S rRNA gene sequencing method based on UMI according to claim 7, wherein, When aligning the sequences in each cluster, a coverage rate > 30 is used as the cut-off value.
9. The full-length 16S rRNA gene sequencing method based on UMI according to claim 1, wherein In step 4, the relative abundance of each consensus sequence is calculated according to the number of UMIs.
Citation Information
Cited By
Circular RNA full-length sequencing method based on unique molecular marker and library amplification quantitative control
CN121320506A