Method for constructing high-quality 16S rDNA sequencing library based on double-end UMI technology and application thereof
By adding paired-end UMI sequences to both ends of the 16S rDNA molecule and combining high-fidelity enzyme and magnetic bead purification technology, a high-quality 16S rDNA sequencing library can be constructed, which solves the problems of high error rate and many chimeras in the existing technology, improves the accuracy and molecular diversity of sequencing data, and is particularly suitable for microbiome research and clinical diagnosis.
Patent Information
- Application Number
- CN202511455888.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-02-24
AI Technical Summary
Existing 16S rDNA sequencing technologies suffer from high error rates, numerous chimeras, and insufficient data accuracy, particularly in the detection of low-abundance species. Current methods struggle to effectively correct sequencing errors and distinguish true mutations.
The paired-end UMI technology was used to add 12 bp random base UMI sequences to both ends of the 16S rDNA molecule. Through paired-end UMI cross-validation, the uniqueness of the original molecule was marked and traced. Combined with high-fidelity enzyme and magnetic bead purification technology, a high-quality sequencing library was constructed.
It significantly improves library quality and data analysis accuracy, reduces the false positive rate of chimeras, enhances molecular diversity and sequencing data accuracy, and is suitable for microbiome research and clinical diagnosis.
Smart Images

Figure CN121555609A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a method for constructing high-quality 16S rDNA sequencing libraries based on paired-end UMI technology, which is particularly suitable for fields such as microbiome research, clinical diagnosis, and environmental monitoring. Background Technology
[0002] 16S rRNA amplicon sequencing technology has become a core tool in environmental, medical, and industrial microbiology research due to its ability to rapidly resolve the composition and diversity of complex microbial communities. However, this technology is susceptible to various sources of error throughout the entire process, from sample processing to data analysis. These include PCR amplification errors (such as primer bias, polymerase mismatch, and chimera formation), inherent errors in sequencing platforms (such as the base substitution error rate of Illumina MiSeq), and limitations of bioinformatics analysis. For example, chimeric products can lead to up to 70% false positive species identification, while random errors and primer bias in PCR amplification can significantly distort the true distribution of microbial abundance. Furthermore, the error rate of sequencing instruments (such as the approximately 0.65% error rate of the PacBio SMRT platform) further reduces the accuracy of sequence data. The combined effect of these errors results in a significant deviation between sequencing results and the actual microbial composition, particularly in the detection of low-abundance species.
[0003] To address the aforementioned issues, existing technologies primarily focus on two aspects: experimental optimization and bioinformatics correction. At the experimental level, optimizing PCR cycle number, annealing temperature, and primer design (e.g., the universal primer 515F / 806R for the V3-V4 region) can partially reduce amplification bias; using mock communities as standards is used to assess systematic errors in the sequencing process. At the bioinformatics level, chimera detection tools (e.g., Chimera Slayer) and algorithms based on consistent sequences (e.g., Amplicon Sequence Variant) are widely used for noise reduction. However, these methods have significant limitations: experimental optimization cannot completely eliminate random errors; standards are only used for quality control rather than directly correcting data; bioinformatics tools have limited error correction capabilities for low-depth sequencing data and rely on threshold settings (e.g., OTU clustering with 97% similarity), which may lead to a loss of species resolution. For example, while ASV can improve classification accuracy, its accuracy is limited by the quality of the original data and cannot distinguish between true mutations and sequencing errors.
[0004] Unique molecular identifier (UMI) technology, by adding random barcodes to each raw DNA molecule, can track repetitive sequences during PCR amplification and sequencing, effectively distinguishing real biological signals from technical noise. In fields such as tumor genomics and single-cell transcriptomics, UMI has been shown to reduce sequencing error rates to below 0.1% and enable the detection of ultra-low frequency mutations. For example, structured UMI design, by optimizing nucleotide arrangement, reduces non-specific amplification and improves detection sensitivity. However, the application of this technology in 16S rDNA sequencing has not been fully explored. Although V3-V4 region primers (such as 338F / 806R) are widely used due to their high versatility, their integration with UMI design and quantitative validation based on standards still lack systematic research. In existing 16S rDNA sequencing pipelines, the absence of UMI makes it difficult to trace amplification and sequencing errors at the molecular level, limiting the accuracy of data correction. Summary of the Invention
[0005] This invention addresses the problems of high error rates and numerous chimeras in existing 16S rDNA sequencing library construction techniques by proposing an innovative solution. Specifically, this invention provides a method for constructing a high-quality 16S rDNA sequencing library with paired-terminal UMIs. These UMIs are composed of random bases, and their structure is specifically designed as NNNNNNNNNNNN, with a total length of 12 base pairs (bp). Furthermore, this invention incorporates 4... 12 A UMI sequence with extremely high randomness. In this invention, the letter "N" in the primer sequence is selected from any one of the bases A, T, C, and G.
[0006] By adding UMI sequences to both ends of the 16S rDNA molecule during the library construction process of this invention, precise labeling of the original molecules is achieved. This step not only greatly enhances the uniqueness and traceability of molecules in the library but also provides strong support for subsequent data analysis. During the data analysis stage, the uniqueness of the paired-end UMIs allows for the accurate identification and removal of chimeras generated during library construction. Simultaneously, since each UMI corresponds to an original 16S rDNA molecule, the molecular proportions in the original sample can be accurately reconstructed, thereby significantly improving the quality of the library and the accuracy of data analysis.
[0007] This invention proposes an innovative method for constructing high-quality 16S rDNA sequencing libraries containing paired-end unique molecular identifiers (UMIs) by amplifying the V3V4 region of 16S rDNA using modified universal primers. The modification involves adding unique UMI sequences and sequencing primer Reads to the 5' ends of a pair of conventional amplification primers (forward and reverse primers) during the first round of PCR amplification. Specifically, the forward primer is supplemented with the UMI1 sequence and sequencing primer Read1, and the reverse primer is supplemented with the UMI2 sequence and sequencing primer Read2. Then, in the second round of PCR amplification, using the amplified product obtained after the first round of PCR amplification and magnetic bead purification as a template, a pair of library construction primers containing an Index tag (specifically, in this pair of primers, the forward primer contains an Index tag, an Illumina sequencing adapter P5 sequence, and sequencing primer Read1, and the reverse primer contains an Index tag, an Illumina sequencing adapter P7 sequence, and sequencing primer Read2) is used for PCR amplification. This step embeds the P5 / P7 adapters and index tags required for sequencing into both ends of the library fragment, constructing a complete library that can be used for subsequent sequencing on the Illumina platform (see [link to Illumina library]). Figure 1 The specific steps are as follows: (a) Design a paired molecular marker system by adding UMI: UMI1 and UMI2, composed of random bases, to both ends of a 16S rDNA molecule. Both UMI1 and UMI2 are 12bp random sequences (NNNNNNNNNNNN). (b) DNA extraction: Total DNA was extracted from the sample using a combination of mechanical disruption and enzymatic digestion as the starting material for subsequent reactions. Alternatively, microbial double-stranded DNA standards could be used as the starting material, such as ZymoBIOMICS® Microbial Community DNA Standard (D6305) and ZymoBIOMICS® GutMicrobiome Standard (D6331) used in the construction method. (c) First round of PCR amplification, the first round of amplification uses a pair of amplification primers including sequencing primer Read, UMI sequence and 16S V3-V4 universal primers; (d) Magnetic bead purification to remove non-specific amplification products; (e) The second round of PCR amplification introduces the index tag and sequencing adapter, and a pair of library construction primers are used in the second round of amplification; (f) Obtain the target fragment library by purification with magnetic beads; (g) Quality testing of the final product, i.e. library quality control.
[0008] Furthermore, for both UMI1 and UMI2, each N is independently selected from four bases: A, T, C, and G. Therefore, the complexity of each UMI sequence is achieved through 4¹² random combinations, and each UMI corresponds to an original 16S rDNA molecule.
[0009] Further, in step (c), the forward primer structure of the pair of amplification primers is 5'-sequencing primer Read1+UMI1+16S V3-V4 universal primer-3'; the reverse primer structure of the pair of amplification primers is 5'-sequencing primer Read2+UMI2+16S V3-V4 universal primer-3'. Preferably, the first round of PCR amplification uses an 8-cycle pre-amplification program, and a high-fidelity DNA polymerase, such as KAPA HiFi Hot Start enzyme, is preferred during amplification, with an annealing temperature preferably of 58°C. Preferably, the base sequences of the pair of amplification primers are SEQ ID No. 1 and SEQ ID No. 2.
[0010] Further, in step (d), non-specific products smaller than 300 bp are removed using AMPure XP magnetic beads with a volume of 0.8 × 10⁸.
[0011] Further, in step (e), the second round of PCR amplification is performed for 20 cycles, with the amplification conditions being: denaturation at 98°C for 20 seconds, annealing at 64°C for 15 seconds, and extension at 72°C for 15 seconds. Preferably, the forward primer of the pair of library construction primers contains an Index tag, an Illumina sequencing adapter P5 sequence, and sequencing primer Read1, and the reverse primer of the pair of library construction primers contains an Index tag, an Illumina sequencing adapter P7 sequence, and sequencing primer Read2. Preferably, the base sequences of the pair of library construction primers are SEQ ID No. 3 and SEQ ID No. 4.
[0012] Furthermore, in step (e), non-specific products smaller than 300 bp are removed using AMPure XP magnetic beads with a volume of 0.8 × 10⁸.
[0013] Further, in step (g), the Agilent 4200 detects the target fragment peak at 600±50 bp, and the final product concentration is ensured to be >20 nM by quantification using a nucleic acid quantification instrument. The nucleic acid quantification instrument used can be any commercially available instrument such as the Qubit fluorometer (Invitrogen).
[0014] In recent years, the standardized application of standards (such as ZymoBIOMICS) has provided a benchmark for evaluating 16S rDNA sequencing errors, while the read length enhancement of high-throughput sequencing platforms (such as Illumina NovaSeq) has made UMI integration feasible. This invention directly quantifies the error rate reduction effect of UMI by embedding UMI sequences into universal primers V3-V4, combined with amplification of standard templates and alignment with consistent sequences. For example, by comparing the matching degree between the UMI and standard sequences before and after adding UMI, the decrease in the mismatch rate can be accurately calculated, thus avoiding the limitations of traditional methods that rely on indirect statistical models. This integration strategy not only fills the technological gap in 16S sequencing using UMI but also provides a new paradigm for the precision and standardization of microbiome research.
[0015] Compared with the prior art, the technical features and beneficial effects of the present invention are as follows: (1) Construction of a dual-end UMI molecular marker system: By introducing unique molecular identifiers (UMIs) at both ends of the 16S rDNA molecule, a dual-end marker system is formed. This design achieves unique marking of each original molecule through molecular tracing algorithms (such as dual-end UMI cross-validation), significantly improving the traceability of molecules in the library; (2) High-complexity UMI sequence design: The UMI is composed of a 12 bp random base sequence (NNNNNNNNNNNN), which is constructed by 4 12 (i.e., 16,777,216) random combinations generate an ultra-high complexity tag system, effectively avoiding tag interference and improving molecular diversity and sequencing data accuracy; (3) Staged amplification process optimization: adopting a two-stage PCR amplification strategy (first round low cycle number pre-amplification + second round exponential amplification), combined with high-fidelity enzyme (KAPA HiFi HotStart) and magnetic bead purification technology, reducing the proportion of non-specific amplification products and increasing library concentration, thereby meeting the needs of high-throughput sequencing; (4) Chimera suppression and data error correction: through the molecular tracing function of dual-end UMI, accurately identify and remove chimeras generated by PCR amplification in the data analysis stage, and at the same time use UMI molecular counting to correct amplification bias, so as to reduce the false positive rate of rare species; (5) Original sample information restoration: each UMI uniquely corresponds to an original 16S rDNA molecules support absolute molecular quantification analysis. Experiments of this invention show that this method can restore the original molecular abundance in complex microbial communities, significantly improve the accuracy of microbial community structure analysis, and is especially suitable for the detection of low-abundance pathogens in clinical samples; (6) Improved efficiency of the whole process: The library construction process is simplified, the single sample operation time is shortened to 4 hours, and it is compatible with the Illumina platform for parallel processing of 96 samples. In addition, combined with the UMI-assisted automated data analysis algorithm, the research cycle is compressed and the labor cost is reduced. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram illustrating the construction principle of the 16S rDNA sequencing library with paired-end UMIs constructed according to the present invention.
[0018] Figure 2 This is a flowchart illustrating the construction process of the 16S rDNA sequencing library with paired-end UMIs as described in this invention.
[0019] Figure 3 This is an electrophoretic detection image of the 16S rDNA sequencing library template with paired-end UMIs constructed using this invention. This method can be applied even with extremely low template amounts; when the template amount in the system is only 0.06 ng, a clear band can be observed.
[0020] Figure 4 This is an electrophoretic image of a 16S rDNA sequencing library with paired-end UMIs constructed using this invention. The final library bands were obtained by comparing different standards and different cycle numbers used in library construction.
[0021] Figure 5 This is a data analysis diagram of a 16S rDNA sequencing library with paired-end UMIs constructed using this invention. The sequencing data obtained using UMI screening and all data were compared with the microbial sequences in the standards after sequencing the library constructed using this invention. The results show that after UMI correction, the proportion of sequences identical to those in the standard microorganisms significantly increased, thus improving the accuracy of the obtained data. Detailed Implementation
[0022] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the contents of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by this invention.
[0023] The construction process of the 16S rDNA sequencing library with paired-end UMIs constructed in this invention is as follows: Figure 2 As shown.
[0024] To achieve the technical objective of this invention, the following uses standard DNA (D6305 and D6331, ZymoBIOMICS) as an example to illustrate the implementation process and operating procedures of the technical solution in detail with reference to embodiments: 1. Dilute the standard DNA (D6305 and D6331, ZymoBIOMICS) by diluting the original concentration of 10 ng / μl to 1 ng / μl.
[0025] 2. First round of PCR amplification (UMI introduction) (1) Reaction system configuration Prepare a 50 μl reaction system according to the following composition:
[0026] DNA templates extracted from microorganisms can also be used instead of standards. Primer sequences are shown in SEQ ID No. 1 and SEQ ID No. 2.
[0027] (2) The PCR instrument settings are as follows:
[0028] The conventional method (as a control) requires 20-25 PCR cycles to amplify the target region of a 16S rDNA sequencing library without UMI. Here, we will use 20 cycles as an example. The method of the present invention optimizes the number of PCR cycles required to amplify the target region of a 16S rDNA sequencing library containing UMI to 8.
[0029] (3) Product purification 1) Add AMPure XP magnetic beads (Beckman Coulter) at a volume ratio of 0.8, vortex mix well, and incubate at room temperature for 5 minutes; 2) Separate the magnetic beads using a magnetic rack, and wash them twice with 200 μl of 70% ethanol. 3) After air drying, the sample was washed with 24 μl of enzyme-free water, and the final library concentration was ≥15 ng / μl (detected by Qubit fluorometer (Invitrogen)).
[0030] 3. Second round of PCR amplification (sequencing adapter ligation) (1) Reaction system configuration Prepare a 50 μl reaction system according to the following composition:
[0031] Primer sequences are shown in SEQ ID No. 3 and SEQ ID No. 4.
[0032] (2) Set the PCR instrument to the following conditions:
[0033] In conventional methods, the number of PCR cycles required for adding sequencing adapters without UMI is 6-10. Here, we will use 8 cycles as an example. In the method of this invention, the number of PCR cycles required for adding sequencing adapters with UMI is optimized to 20.
[0034] (3) Document Quality Inspection 1) Target bands were screened using 1.5% agarose gel electrophoresis, and purified and recovered by gel cutting using the NucleoSpin Gel and PCR Clean-up Kit (MN); Using standard DNA D6305 as a template, sequencing libraries were constructed with different amounts of DNA template. The final PCR products were then subjected to agarose gel electrophoresis. The results showed that the electrophoretic bands became brighter with increasing amounts of standard DNA template. Bands were observed when the template amount was only 0.03 ng, and distinct bands were observed when the template amount reached 0.5 ng. Figure 3 As shown in the figure. These results suggest that this method has high sensitivity and can still construct high-quality 16S rDNA sequencing libraries even with trace amounts of template DNA.
[0035] To evaluate the advantages of the method of this invention in improving sequencing accuracy, this embodiment selected two standard DNA samples as templates and constructed 16S rDNA sequencing libraries using both the conventional method (20 cycles of PCR in the first round, 8 cycles of PCR in the second round) and the UMI method of this invention (8 cycles of PCR in the first round, 20 cycles of PCR in the second round). Four independent replicate experiments were performed, and the final PCR libraries were subjected to gel electrophoresis. The results are shown in Figure 4: regardless of the standard DNA used, both the UMI method of this invention and the conventional method yielded bright and clear electrophoretic bands. This indicates that the UMI method of this invention can stably construct high-quality libraries.
[0036] (4) Data Analysis To verify the effectiveness of this invention in improving the accuracy of sequencing data, a comparative data analysis was conducted: the homogeneous sequence data obtained using the UMI method was compared with the original data without UMI error correction. The results showed that for standard DNA D6305, the proportion of mismatched sequences in the data without UMI error correction was as high as 30.69% compared to the known reference sequence. After UMI error correction, the proportion of mismatched sequences significantly decreased to only 21.50%. Furthermore, the optimized number of PCR cycles (8 cycles in the first round, 20 cycles in the second round) also significantly reduced the mismatch rate, decreasing it from 25% to 15.85%. For standard DNA D6331, the trend was similar to D6305, but the effect of UMI error correction was more pronounced: it reduced the mismatch rate from 14.96% to 4.51%, and after optimizing the number of PCR cycles, the mismatch rate further decreased from 14.79% to 3.95%, as shown in the results. Figure 5 As shown.
[0037] In summary, the UMI method of this invention can significantly reduce the mismatch rate of sequencing data, thereby effectively improving sequencing accuracy.
[0038] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a high-quality 16S rDNA sequencing library based on paired-end UMI technology, characterized in that... Includes the following steps: (a) Design a paired molecular marker system by adding UMI: UMI1 and UMI2, composed of random bases, to both ends of a 16S rDNA molecule. Both UMI1 and UMI2 are 12bp random sequences (NNNNNNNNNNNN). (b) Extract total DNA from the sample as starting material; (c) First round of PCR amplification, the first round of amplification uses a pair of amplification primers including sequencing primer Read, UMI sequence and 16SV3-V4 universal primer; (d) Magnetic bead purification to remove non-specific amplification products; (e) Sequencing adapters are introduced in the second round of PCR amplification, using a pair of library preparation primers; (f) Obtain the target fragment library by purification with magnetic beads; (g) Perform quality testing on the final product.
2. The method according to claim 1, characterized in that: The forward primer structure of the pair of amplification primers in step (c) is 5'-sequencing primer Read1+UMI1+16S V3-V4 forward primer-3', and the reverse primer structure of the pair of amplification primers is 5'-sequencing primer Read2+UMI2+16S V3-V4 reverse primer-3'.
3. The method according to claim 1, characterized in that: The first round of PCR amplification in step (c) uses an 8-cycle pre-amplification program, high-fidelity DNA polymerase, and an annealing temperature of 58°C.
4. The method according to claim 1, characterized in that: The second round of PCR amplification in step (e) is performed in 20 cycles, with the following amplification conditions: denaturation at 98°C for 20 seconds, annealing at 64°C for 15 seconds, and extension at 72°C for 15 seconds.
5. The method according to claim 1, characterized in that: In steps (d) and (f), magnetic bead purification uses 0.8 × volume of AMPure XP magnetic beads to specifically remove non-target fragments smaller than 300 bp.
6. The method according to claim 1, characterized in that: The quality control in step (g) includes using an Agilent 4200 to detect the target fragment peak at 600±50bp and using a nucleic acid quantification instrument to ensure that the final product concentration is >20nM.
7. The method according to claim 2, characterized in that: The base sequences of the amplification primer pair are SEQ ID No 1 and SEQ ID No 2.
8. The method according to claim 1, characterized in that: The method is adapted to the Illumina sequencing platform. In the second round of amplification in step (e), the forward primer of the pair of library preparation primers contains an Index tag, an Illumina sequencing adapter P5 sequence, and a sequencing primer Read1, and the reverse primer of the pair of library preparation primers contains an Index tag, an Illumina sequencing adapter P7 sequence, and a sequencing primer Read2.
9. The method according to claim 8, characterized in that: The base sequences of the library construction primer pair are SEQ ID No 3 and SEQ ID No 4.
10. The application of the 16S rDNA sequencing library constructed by the method of any one of claims 1-9 in microbiome research, clinical pathogen detection, or environmental microbial monitoring.