A method for high-throughput target sequence screening of DNA library based on single molecule tag UMI
Patent Information
- Application Number
- CN202611230642.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
本发明旨在解决合成DNA文库筛选领域的核心技术瓶颈:一是传统克隆筛选法操作繁琐、周期长、成本高、通量有限;二是采用固定序列UMI时易发生标签碰撞,无法保证模板分子与标签的一一对应;三是单纯稀释模板虽可缓解标签碰撞,但会导致PCR扩增效率急剧下降甚至扩增失败,即“标签一一对应”与“PCR扩增稳定性”不可兼得的技术矛盾
1、突破核心技术矛盾,从根本上解决固定UMI的标签碰撞难题
Smart Images

Figure CN122811344A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, specifically to the field of high-throughput nucleic acid screening and synthetic biology, and particularly to a method for high-throughput target sequence screening of DNA libraries based on single-molecule tags (UMI). This invention can be widely applied to scenarios such as correct sequence enrichment of chemically synthesized or enzymatically synthesized DNA libraries, quality control of gene synthesis products, and sequence sorting for DNA data storage, providing a technical solution for high-throughput, low-cost acquisition of high-purity target DNA fragments. Background Technology
[0002] With the rapid development of synthetic biology, antibody library screening, and DNA information storage, the industry's demand for rapid delivery of long fragments, low error rates, and customized genes is growing exponentially. Complex oligonucleotide libraries (DNA libraries) are key materials for diverse applications such as synthetic biology, recombinant protein production, and DNA storage. Currently, DNA acquisition mainly relies on two technical routes: chemical synthesis (such as the solid-phase phosphoramide method) and enzymatic synthesis (such as synthesis based on terminal deoxynucleotidyl transferase TdT). However, due to synthetic errors, both methods generate DNA pools containing a large number of erroneous sequences during the preparation process, rather than a single pure target sequence. How to efficiently and accurately screen out the correct target fragments from these pools has become a key technical bottleneck restricting downstream applications.
[0003] Currently, the mainstream method for screening target sequences from DNA pools is still molecular cloning combined with Sanger sequencing. This requires multiple steps, including E. coli transformation, overnight culture, single clone selection, colony PCR identification, plasmid extraction, and Sanger sequencing verification. The overall process is cumbersome, costly, and time-consuming. Moreover, the longer the gene length, the more difficult it is to screen for the correct clone, which is difficult to match the production capacity requirements of high-throughput synthesis.
[0004] In recent years, Unique Molecular Identifier (UMI) technology has been widely used in scenarios such as low-frequency mutation detection in ctDNA and RNA-Seq transcriptome analysis due to its ability to label original DNA molecules. The conventional technical route involves labeling a template molecule with a random, degenerate UMI sequence via ligation or PCR. After amplification and high-throughput sequencing, the original template is traced using the UMI. Consistency analysis of multiple reads under the same UMI filters out errors introduced by PCR amplification and sequencing, thereby improving the detection sensitivity of low-frequency mutations. In this type of technology, the core role of the UMI is to correct sequencing data errors, with the technical goal of improving detection accuracy and ultimately outputting bioinformatics analysis results. It typically uses a degenerate UMI of 6-12 random bases, relying on a very high theoretical number of combinations for labeling. Furthermore, multiple copies of the same template corresponding to the same UMI are more conducive to error correction analysis, and a strict one-to-one correspondence between the UMI and the template molecule is not required. This type of technical solution is designed for clinical gene testing scenarios and differs fundamentally from the technical requirements of target sequence sorting and physical enrichment in synthetic DNA libraries, making it unsuitable for direct application.
[0005] If UMI technology is extended to the screening of correct sequences and physical recovery of synthetic DNA libraries—that is, after locating the correct sequence through UMI, using UMI as primers for reverse amplification and recovery of the target nucleic acid—a new technical challenge will be encountered that cannot be solved by existing technologies. If random degenerate UMI is used, the subsequent UMI-specific amplification stage suffers from high primer synthesis costs, poor batch-to-batch consistency, and difficulty in achieving standardized mass production and reuse. If a fixed-sequence UMI primer pool with lower cost and better stability is used, a contradiction will arise where the total number of UMI combinations is far lower than the total copy number of template molecules, resulting in a "tag collision" problem.
[0006] Tag collisions fall into two categories: first, the same UMI combination simultaneously labels both correct and incorrect sequences; second, the same UMI combination simultaneously labels multiple different target synthetic sequences. Both types of collisions prevent the accurate enrichment of a single correct target sequence during subsequent UMI-specific amplification, directly causing screening failure.
[0007] Specifically, taking chip synthesis as an example, chip-based synthesis can achieve a loading capacity of fmol, which, based on Avogadro's constant, translates to a total DNA molecule copy number in the billions. Considering the difficulty of engineering implementation and cost control, using 384 fixed UMI ligation primers each for forward and reverse directions only yields approximately 140,000 specific primer combinations (384 × 384). When using conventional homogeneous UMI ligation technology, the same UMI combination will correspond to tens of thousands of DNA molecules, inevitably leading to tag collisions. This ultimately results in the inability to establish a unique correspondence between UMIs and a single target sequence in the sequencing data, making it impossible to obtain a pure target sequence through UMI-specific amplification.
[0008] To address the tag collision problem with fixed UMIs, the conventional approach in this field is to dilute the initial template concentration, thereby reducing the probability of the same UMI combination labeling multiple templates by decreasing the total template amount. However, in this application scenario, to achieve a one-to-one correspondence between UMIs and templates, template molecules need to be diluted tens of thousands or even hundreds of thousands of times, resulting in a template concentration that is only one-ten-thousandth to one-hundred-thousandth of that in conventional PCR; excessive template dilution will lead to PCR amplification failure. This creates a technical contradiction between "ensuring a one-to-one correspondence between UMIs and templates" and "maintaining a stable PCR amplification success rate," for which there is currently no effective solution.
[0009] In summary, there is an urgent need in this field to develop a DNA library target sequence screening method based on fixed sequence UMI, which can guarantee a one-to-one correspondence between template and tag, and has both high throughput and low cost, in order to break through the bottleneck of existing synthetic DNA screening technology. Summary of the Invention
[0010] The purpose of this invention is to address the shortcomings of existing technologies by providing a high-throughput target sequence screening method for DNA libraries based on single-molecule tag (UMI) and a matching primer combination. This invention aims to solve the core technical bottlenecks in the field of synthetic DNA library screening: first, traditional cloning screening methods are cumbersome, time-consuming, costly, and have limited throughput; second, when using fixed-sequence UMI, tag collisions are prone to occur, making it impossible to guarantee a one-to-one correspondence between template molecules and tags; third, while simply diluting the template can alleviate tag collisions, it leads to a sharp decrease in PCR amplification efficiency or even amplification failure, i.e., the technical contradiction between "one-to-one tag correspondence" and "PCR amplification stability" is mutually exclusive.
[0011] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: In a first aspect, the present invention claims a method for high-throughput target sequence screening of DNA libraries based on single-molecule tag UMI.
[0012] The method for high-throughput target sequence screening of DNA libraries based on single-molecule tag UMI, as claimed in this invention, may include the following steps: (A1) All DNA samples to be screened are grouped to obtain at least one group, and each group is defined as a DNA library; the sequences of all DNA samples to be screened are composed of a template 5' end universal sequence, a synthetic sequence, and a template 3' end universal sequence from the 5' end to the 3' end; in all DNA samples to be screened, the template 5' end universal sequence is the same, the template 3' end universal sequence is also the same, and the template 5' end universal sequence and the template 3' end universal sequence are neither the same nor reverse complementary; (A2) Each DNA library corresponds to a PCR reaction system. Using the DNA library as a template, all primers in the UMI ligation primer combination are mixed and added to the corresponding PCR reaction system. The UMI ligation primer set is a fixed sequence primer pool, consisting of several forward UMI ligation primers and several reverse UMI ligation primers; wherein, the forward UMI ligation primers, from the 5' end to the 3' end, consist of the 3' end portion of the second-generation sequencing 5' end adapter, the left-end UMI sequence, and the universal sequence A; the reverse UMI ligation primers, from the 5' end to the 3' end, consist of the 3' end portion of the second-generation sequencing 3' end adapter, the right-end UMI sequence, and the universal sequence B; The 3' end portion of the second-generation sequencing 5' adapter in all the forward UMI ligation primers is the same, the universal sequence A in all the forward UMI ligation primers is also the same, and the left-end UMI sequence in all the forward UMI ligation primers is different. The 3' end portion of the second-generation sequencing 3' adapter in all the reverse UMI ligation primers is the same, the universal sequence B in all the reverse UMI ligation primers is also the same, and the right-end UMI sequence in all the reverse UMI ligation primers is different. The universal sequence A is completely identical to the universal sequence at the 5' end of the template; the universal sequence B is the reverse complementary sequence to the universal sequence at the 3' end of the template. The reaction system is divided into multiple independent microdroplet reaction units using water-in-oil microdroplet generation technology. The template DNA molecule and the UMI ligation primer pair follow a Poisson distribution in the microdroplet, such that each microdroplet contains no more than one template DNA molecule and no more than one pair of UMI ligation primers. The physical separation of the microdroplets achieves a one-to-one correspondence between the template DNA molecule and the paired-end UMI combination formed by the left-end UMI sequence carried by the forward UMI ligation primer and the right-end UMI sequence carried by the reverse UMI ligation primer. The first step of PCR amplification is performed in the microdroplet reaction unit, and the first step PCR product is obtained after demulsification and collection. The number of the DNA libraries in step (A1) corresponds to the number of the first step PCR products obtained in this step.
[0013] (A3) Each of the first step PCR products corresponds to a PCR reaction system. Using the first step PCR product as a template, the second step PCR amplification is performed using the UMI amplification primer pair to obtain the second step PCR product. The UMI amplification primer pair consists of a universal forward primer and a reverse primer; wherein, the universal forward primer is a complete second-generation sequencing 5' adapter, and its 3' end portion is completely identical to the 3' end portion sequence of the second-generation sequencing 5' adapter in the forward UMI ligation primer; the reverse primer is a complete second-generation sequencing 3' adapter, which consists of the 5' end portion of the second-generation sequencing 3' adapter, a barcode sequence, and a 3' end portion that is completely identical to the 3' end portion sequence of the second-generation sequencing 3' adapter in the reverse UMI ligation primer; The universal forward primer sequence is the same in all PCR reaction systems, while the reverse primers in different PCR reaction systems carry different barcode sequences (i.e., different DNA libraries can be distinguished by the barcode sequence). (A4) Perform high-throughput sequencing on the PCR product of the second step, and associate the left-end UMI sequence, the right-end UMI sequence and the corresponding synthetic sequence based on the sequencing data, and screen out the unique set of the left-end UMI sequence and the right-end UMI sequence corresponding to the synthetic sequence that completely matches the preset target sequence; (A5) Using the PCR product of the second step as a template, PCR amplification is performed using the UMI-specific PCR primer pair composed of the left-end UMI sequence and the right-end UMI sequence obtained by screening in (A4) to obtain the target DNA sequence.
[0014] Further, in (A1), the number of DNA libraries can be 1 to 8. Correspondingly, in (A3), the number of UMI amplification primer pairs can be 1 to 8.
[0015] Further, in (A1), each of the DNA libraries contains N DNA samples to be screened, where N is a positive integer greater than 1. In some embodiments of the present invention, the number of DNA libraries is 1, and N is 96.
[0016] Further, in (A2), the number of each of the forward and reverse UMI ligation primers constituting the UMI ligation primer combination is independently selected from 200 to 1000. In some embodiments of the present invention, the number of each of the forward and reverse UMI ligation primers constituting the UMI ligation primer combination is 384.
[0017] Further, in (A2), the total number of the paired-end UMI combinations is less than 1 / 200 of the total copy number of the template DNA molecules in the DNA library. In some embodiments of the invention, the total number of the paired-end UMI combinations is approximately 1 / 271 of the total copy number of the template DNA molecules in the DNA library.
[0018] In some embodiments of the present invention, the UMI ligation primer combination consists of 384 forward UMI ligation primers shown in SEQ ID NO:1 to SEQ ID NO:384 and 384 reverse UMI ligation primers shown in SEQ ID NO:385 to SEQ ID NO:768.
[0019] In some embodiments of the present invention, the nucleotide sequence of the universal forward primer in the UMI amplification primer pair is SEQ ID NO:769, and the nucleotide sequence of the reverse primer in the UMI amplification primer pair is selected from any one of SEQ ID NO:770 to SEQ ID NO:777.
[0020] In some embodiments of the present invention, the forward primer in the UMI-specific PCR primer pair is selected from any one of SEQ ID NO:778 to SEQ ID NO:1161, and the reverse primer is selected from any one of SEQ ID NO:1162 to SEQ ID NO:1545.
[0021] Further, in the PCR reaction system of (A2), the final concentration of the UMI ligation primer combination is 0.004 μM-0.4 μM; in the PCR reaction system of (A2), the final concentration of the template is 0.1 ng-1 ng / μL. In some embodiments of the present invention, in the PCR reaction system of (A2), the final concentration of the UMI ligation primer combination is 0.04 μM; in the PCR reaction system of (A2), the final concentration of the template is 0.2 ng / μL. Specifically, in some embodiments of the present invention, the PCR reaction system of (A2) is shown in Table 1.
[0022] Further, in the reaction procedure for the first step of PCR amplification performed in (A2), the denaturation-annealing-extension cycle is repeated twice. In some embodiments of the present invention, the denaturation-annealing-extension cycle is repeated twice in the reaction procedure for the first step of PCR amplification. Specifically, in some embodiments of the present invention, the reaction procedure for the first step of PCR amplification performed in (A2) is shown in Table 2.
[0023] Furthermore, in (A2), all the left-end UMI sequences and all the right-end UMI sequences are preferably also different.
[0024] In some embodiments of the present invention, in the PCR reaction system of (A3), the final concentrations of the universal forward primer and the reverse primer constituting the UMI amplification primer pair are both 0.2 μM. In some embodiments of the present invention, in the reaction program for the second step of PCR amplification in (A3), the number of denaturation-annealing-extension cycles is 30.
[0025] In some embodiments of the present invention, the PCR reaction system of (A3) is shown in Table 3; the reaction procedure for the second step of PCR amplification in (A3) is shown in Table 4.
[0026] In some embodiments of the present invention, in the PCR amplification reaction system performed in (A5), the final concentrations of both the left-end UMI sequence and the right-end UMI sequence constituting the UMI-specific PCR primer pair are 0.2 μM. In some embodiments of the present invention, in the PCR amplification reaction program performed in (A5), the number of denaturation-annealing-extension cycles is 30.
[0027] In some embodiments of the present invention, the reaction system for the PCR amplification performed in (A5) is shown in Table 6; the reaction procedure for the PCR amplification performed in (A5) is shown in Table 7.
[0028] Furthermore, in (A4), after high-throughput sequencing of the PCR product from the second step, a step may also be included to distinguish different DNA libraries in the sequencing data based on specific barcode sequences that correspond one-to-one with each DNA library.
[0029] Furthermore, in (A2), after the first step of PCR amplification is completed, a step of purifying the first step PCR product is also included.
[0030] Furthermore, in (A3), after the second step of PCR amplification, a step of purifying the PCR product from the second step is also included. The purified product is the final library, used for the high-throughput sequencing described in (A4).
[0031] Furthermore, in step (A4), the high-throughput sequencing includes, but is not limited to, second-generation sequencing platforms; when the target synthesized sequence is longer, third-generation sequencing platforms (such as single-molecule real-time sequencing, nanopore sequencing) can be used to cover the complete long fragment sequence with long reads, further expanding the applicable length range of the method.
[0032] Secondly, this invention claims protection for a primer combination.
[0033] The primer combinations claimed in this invention include (or consist of): (B1) UMI ligation primer set; the UMI ligation primer set consists of 384 forward primers shown in SEQ ID NO:1 to SEQ ID NO:384 and 384 reverse primers shown in SEQ ID NO:385 to SEQ ID NO:768; (B2) Eight UMI amplification primer pairs; the nucleotide sequence of the forward primer in each UMI amplification primer pair is SEQ ID NO:769; the nucleotide sequences of the reverse primers in the eight UMI amplification primer pairs are SEQ ID NO:770 to SEQ ID NO:777, respectively; (B3) UMI-specific PCR primer set; the UMI-specific PCR primer set consists of 384 forward primers shown in SEQ ID NO:778 to SEQ ID NO:1161 and 384 reverse primers shown in SEQ ID NO:1162 to SEQ ID NO:1545.
[0034] Thirdly, the present invention claims the use of the primer combinations described in the second aspect above in high-throughput target sequence screening of DNA libraries based on single-molecule tag UMI.
[0035] The beneficial effects of this invention are: 1. Overcome the core technological contradictions and fundamentally solve the tag collision problem of fixed UMI. This invention innovatively employs a water-in-oil microdroplet-separated PCR system. Based on Poisson distribution, it precisely controls the concentration of template and primers, and through physical separation, it forces a one-to-one correspondence of "single droplet = 1 template + 1 pair of UMI primers" without the need for excessive template dilution. This completely avoids the tag collision problem that inevitably occurs when fixing UMIs in a homogeneous system, and ensures stable PCR reaction efficiency within each droplet. It breaks through the technical bottleneck of the incompatibility between "one-to-one tag correspondence" and "PCR amplification success rate" in existing technologies.
[0036] 2. The efficiency of correct sequence enrichment is significantly improved. This invention uses UMI single-molecule tags for sequence tracing and precise sorting, achieving a target fragment correct sequence yield of up to 90.2%, which is more than 3 times more efficient than traditional cloning screening methods. It effectively solves the industry pain points of high error rate and low correct sequence ratio in synthetic DNA libraries.
[0037] 3. The screening cycle is significantly shortened, and the throughput is significantly increased. This invention is a whole cell-free system that eliminates the need for cumbersome operations such as cloning transformation, bacterial culture, and bacterial selection verification. It can screen and capture nearly 800 DNA fragments within 2 days, which is more than 50% shorter than the traditional cloning screening route that takes more than 4 days. The throughput of a single screening session can reach 768 fragments, which is 8 times higher than the traditional method and can fully meet the production capacity requirements of high-throughput DNA synthesis.
[0038] 4. Screening costs are significantly reduced, and the technology is highly adaptable to engineering applications. Using a fixed-sequence UMI primer pool allows for batch synthesis and reuse, with primer costs far lower than those of random degenerate UMI schemes. Taking the screening of 768 fragments as an example, this scheme only requires one high-throughput sequencing and conventional PCR reagents, reducing the overall cost by more than 80% compared to the traditional cloning + Sanger sequencing route, demonstrating significant economic advantages and engineering application value.
[0039] 5. Widely applicable scenarios, compatible with multi-length segment filtering The entire screening process is completed in vitro, without the need for culture media, antibiotics, or other biological reagents, and generates no biological waste liquid, meeting the requirements of green synthesis development. It is compatible with DNA libraries from various sources, including chemical synthesis and enzymatic synthesis, and is also compatible with second-generation and third-generation sequencing platforms. It can be used for screening short fragments at the hundred-base level, and can also be extended to the correct sequence sorting of long fragments at the thousand-base level. It is widely used in many core synthetic biology scenarios such as gene synthesis quality control, antibody library screening, and DNA data storage sequence sorting. Attached Figure Description
[0040] Figure 1 This invention presents the complete workflow of a high-throughput target sequence screening method for DNA libraries based on water-in-oil microdroplet generation technology and single-molecule tag (UMI). A PCR reaction system is prepared by using a DNA library and a primer pool for UMI ligation. The reaction system is divided into multiple independent microdroplet reaction units using water-in-oil microdroplet generation technology to obtain single-molecule microdroplets. This ensures that the template DNA molecule and the UMI ligation primer pair follow a Poisson distribution within the droplets. PCR is then performed within the droplets to ligate the UMI to both ends of the sequence. After demulsification and purification, the product is collected and subjected to a second step of PCR amplification and enrichment, simultaneously completing library construction. The library is then subjected to high-throughput sequencing to analyze the correspondence between the target sequence and the UMI, screening for uniquely matching UMI combinations. Finally, UMI-specific PCR is used to specifically capture the target sequence from the DNA library.
[0041] Figure 2 The library construction process for adding UMI tags to this invention is as follows: The first step, PCR, involves two cycles to ligate the UMI to both ends of the synthesized fragment, while simultaneously adding partial sequencing adapter sequences. The second step, PCR, amplifies the target sequence, and simultaneously adds the complete sequencing adapter sequences to both ends to complete library construction.
[0042] Figure 3 The image shows the agarose gel electrophoresis results after UMI amplification in Example 1. After ligating UMI via microdroplet PCR, UMI amplification and library construction were performed, with each of the three lanes corresponding to three replicate experiments.
[0043] Figure 4 This diagram illustrates UMI sequence analysis and target fragment acquisition. The data is compared with the target sequence, and reads that correctly match the target fragment are selected. The corresponding UMI primer numbers are output, and these UMI primers are then used for amplification to obtain the correct target sequence. Detailed Implementation
[0044] To address the shortcomings of existing DNA synthesis technologies, such as high error rates, insufficient enrichment methods, and low screening throughput, this invention proposes introducing fixed UMI sequences at both ends of the DNA library to construct a "UMI-synthesized sequence-UMI" single-molecule library. It innovatively employs a water-in-oil microdroplet-separated PCR system for UMI labeling, fundamentally solving the tag collision problem of fixed UMIs through physical separation, ensuring a one-to-one correspondence between the template molecule and the UMI combination. Library construction is completed through two PCR steps, and the correspondence between sequence information and UMIs can be obtained simultaneously in a single high-throughput sequencing run. Bioinformatics analysis can accurately identify all UMI combinations that perfectly match the target sequence, and finally, high-purity target fragments are obtained through UMI-specific PCR amplification.
[0045] This process requires no cloning or cell culture and can screen and capture nearly 800 DNA fragments within two days, providing a new, cell-free, cloning-free, and high-throughput technical route for low-cost, customized, long-fragment gene synthesis.
[0046] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0047] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0048] The process flow of this invention includes 6 steps, as detailed below: (1) UMI sequence design: 384 sets of UMI sequences with a length of 18bp were designed at both ends, while ensuring that the GC content was 40-60% and the Tm value was around 60℃; (2) DNA synthesis: A DNA fragment with a length of 220nt / 220bp is synthesized, wherein the left and right sides of the fragment are the universal sequence at the 5' end of the template (which is completely identical to the universal sequence A below) and the universal sequence at the 3' end of the template (which is inversely complementary to the universal sequence B below). (3) Primer design: including UMI ligation primers (designing 384 pairs of UMI ligation primers, the forward primers including the 3' end portion of the second-generation sequencing 5' adapter + the left-end UMI sequence + universal sequence A, the reverse primers including the 3' end portion of the second-generation sequencing 3' adapter + the right-end UMI sequence + universal sequence B. The synthesized 384 forward and reverse primers are mixed into one tube. Among them, the sequences of the 384 forward primers are SEQ ID NO:1 to SEQ ID NO:384, and the sequences of the 384 reverse primers are SEQ ID NO:385 to SEQ ID NO:768), UMI amplification primers (designing the forward universal 5' phosphorylation modified primer PCR-F (complete second-generation sequencing 5' adapter) and the reverse 8 barcode primers PCR-BCX (complete second-generation sequencing 3' adapter, containing barcode). Among them, the sequence of the PCR-F primer is SEQ ID NO:769, and the sequences of the 8 PCR-BCX primers are SEQ ID NO:770 to SEQ ID NO:769. ID NO: 777), UMI-specific PCR primers (such as the 384 forward UMI primers UMI-FX shown in SEQ ID NO: 778 to SEQ ID NO: 1161 and the 384 reverse UMI primers UMI-RX shown in SEQ ID NO: 1162 to SEQ ID NO: 1545).
[0049] (4) Two-step PCR library construction: The first step is to perform UMI-labeled PCR in the water-in-oil microdroplet system, and to achieve single droplet, single template, and single primer pair by Poisson distribution regulation to ensure one-to-one correspondence; after demulsification and purification, the second step is to perform conventional PCR to complete the sequencing adapter and sample barcode.
[0050] (5) High-throughput sequencing and analysis: After circularization, digestion, and DNB making, the library was subjected to PE150 high-throughput sequencing using the BGI Genomics G99 platform. After sequencing, bioinformatics analysis was performed to screen out UMIs that completely matched the target sequence.
[0051] (6) UMI-specific PCR amplification: Based on the screened UMI, select the corresponding UMI-specific PCR primers and perform PCR amplification using the final library as a template to obtain the target DNA sequence.
[0052] When the target synthesized sequence is longer (e.g., kilobase length), the second-step PCR product can be adapted to a third-generation sequencing platform (e.g., single-molecule real-time sequencing, nanopore sequencing) for long-read sequencing. Based on UMI-associated sequence information and screening, the applicable length range of this method can be further expanded.
[0053] Figure 1 This invention presents the complete workflow of a high-throughput target sequence screening method for DNA libraries based on water-in-oil microdroplet generation technology and single-molecule tag (UMI). A PCR reaction system is prepared by using a DNA library and a primer pool for UMI ligation. The reaction system is divided into multiple independent microdroplet reaction units using water-in-oil microdroplet generation technology to obtain single-molecule microdroplets. This ensures that the template DNA molecule and the UMI ligation primer pair follow a Poisson distribution within the droplets. PCR is then performed within the droplets to ligate the UMI to both ends of the sequence. After demulsification and purification, the product is collected and subjected to a second step of PCR amplification and enrichment, simultaneously completing library construction. The library is then subjected to high-throughput sequencing to analyze the correspondence between the target sequence and the UMI, screening for uniquely matching UMI combinations. Finally, UMI-specific PCR is used to specifically capture the target sequence from the DNA library.
[0054] Figure 2 The library construction process for adding UMI tags to this invention is as follows: The first step, PCR, involves two cycles to ligate the UMI to both ends of the synthesized fragment, while simultaneously adding partial sequencing adapter sequences. The second step, PCR, amplifies the target sequence, and simultaneously adds the complete sequencing adapter sequences to both ends to complete library construction.
[0055] Figure 3 The image shows the agarose gel electrophoresis results after UMI amplification in Example 1. After ligating UMI via microdroplet PCR, UMI amplification and library construction were performed, with each of the three lanes corresponding to three replicate experiments.
[0056] Figure 4 This diagram illustrates UMI sequence analysis and target fragment acquisition. The data is compared with the target sequence, and reads that correctly match the target fragment are selected. The corresponding UMI primer numbers are output, and these UMI primers are then used for amplification to obtain the correct target sequence.
[0057] Example 1: UMI screening and acquisition of target fragments for enzymatic DNA synthesis 1. Processing and purification of enzymatically synthesized DNA fragments (1) Select a DNA library containing 96 enzyme-catalyzed synthetic products (synthesized using a synthesizer developed by Zhonghe Gene, the sequences of the 96 enzyme-catalyzed synthetic products are double-stranded DNA as shown in SEQ ID NO:1546 to SEQ ID NO:1641), and vortex to mix. (2) Take 200 μL of DNA library, add 300 μL of magnetic beads (Vaht DNA CleanBeads, catalog number N411) equilibrated to room temperature, mix by pipetting, and let stand at room temperature for 10 min. (3) Place the sample on the magnetic rack and wait for the solution to clarify (about 5 min), then carefully remove the supernatant; (4) Keep the sample on the magnetic rack at all times, add 500 μL of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 s, and carefully remove the supernatant; (5) Repeat step (4) once, for a total of two rinses; (6) Keep the sample on the magnetic rack at all times, and open the lid to dry the magnetic beads for about 2-5 minutes at room temperature; (7) Remove the sample from the magnetic rack, add 80 μL of ddH2O, vortex or pipette to mix thoroughly, and let stand at room temperature for 2 min. After the solution has clarified, let it stand on the magnetic rack for 5 min, and carefully aspirate the supernatant into a new PCR tube; (8) The concentration of the purified product was determined using a Qubit fluorometer and diluted to 10 pg / μL.
[0058] 2. Preparation of UMI ligation primers Order 384 pairs of UMI ligation products (SEQ ID NO:1 to SEQ ID NO:768), each with a single primer concentration of 10 μM. Mix all 768 primers in equal volumes into one tube, then dilute it 10 times to obtain a primer mixture with a total concentration of 1 μM, for later use.
[0059] 3. Preparation of the microdroplet UMI ligation PCR reaction system This step utilizes water-in-oil microdroplet generation technology to divide the reaction system into millions of independent microdroplet reaction units. Based on the Poisson distribution, the template concentration is controlled to ensure that the vast majority of droplets contain only one template molecule and one pair of UMI connecting primers. Through physical separation, a one-to-one correspondence between the template and the paired-end UMI combination is achieved, fundamentally avoiding tag collisions.
[0060] Prepare the aqueous PCR reaction system according to Table 1; Table 1. Microdroplet UMI ligation PCR reaction system
[0061] In this PCR reaction system, the total number of paired-end UMI combinations is 384 × 384 = 147456, and the total copy number of the template DNA molecule is approximately 4 × 10⁻⁶. 7 .
[0062] 4. Microfluidic droplet generation Using the Xinyi Bio Droplet Generator, place the PCR tube into the corresponding position on the instrument, insert the droplet generation chip into the matching generation chip mechanical slot, press down the slot cover to fix the chip, add 30μL of PCR reaction solution to the water well, add 180μL of microdroplet generation oil to the oil phase, then cover the water well and oil well of the chip with the microdroplet generation chip sealing gasket, and place it in the instrument to generate droplets.
[0063] 5. Microdroplet PCR The microdroplets generated above were amplified according to the PCR procedure shown in Table 2; Table 2. Microdroplet UMI ligation PCR reaction procedure
[0064] 6. Demulsification and purification After microdroplet amplification, carefully remove the lower layer of oil. Add 30 μL of demulsifier B and briefly centrifuge; demulsifier B can be observed in the lower layer. Add 1 μL of demulsifier A, gently vortex the mixture, and briefly centrifuge. Incubate the centrifuge tube at 60°C for 10 min, then at 95°C for 10 min, and finally cool to room temperature. Remove the centrifuge tube; the water and oil will separate into clear layers. Carefully transfer the upper aqueous phase to obtain the product and dispose of the lower layer of discarded demulsifier. Purify the aqueous product using 1.5× magnetic beads, reconstitute with 48 μL of water, and obtain the purified product from the first step of PCR.
[0065] 7. UMI amplification (1) Take 46 μL of the above purified product as a template and prepare the PCR reaction system according to Table 3; Table 3. UMI amplification PCR reaction system
[0066] Note: Primer 5P-PCR-F is a universal forward primer (5' phosphorylated) for UMI amplification, and its nucleotide sequence is SEQ ID NO:769. Primer PCR-BCX is a barcode-enabled reverse primer for UMI amplification. Since there is only one DNA library, only one PCR reaction system is needed. Accordingly, only one reverse primer is used in this embodiment, specifically the primer shown in SEQ ID NO:770.
[0067] (2) After the PCR system is prepared, amplification is performed using the PCR reaction procedure shown in Table 4; Table 4. UMI Amplification PCR Reaction Procedure
[0068] (3) After the PCR reaction, the PCR product was purified using 150 μL (1.5×) magnetic beads (VAHTS DNA CleanBeads, catalog number N411) and reconstituted with 52 μL of water. The purified product can be used as a library for high-throughput sequencing.
[0069] 8. Library sequencing (1) Use the MGIEasy circularization module V2.0 to circularize and digest the library obtained in the above steps; (2) DNB preparation and PE150 sequencing were performed on the G99 sequencing platform using the DNBSEQ-G99RS high-throughput sequencing reagent kit (G99FCL PE150).
[0070] 9. Data analysis after machine operation For each enzymatically synthesized DNA fragment, UMI reads that matched the expected sequence and corresponded to only one unique sequence were screened, and their corresponding UMI numbers were output (purpose: to effectively exclude invalid fragments generated during enzymatic DNA synthesis due to mismatches and non-specific amplification, while avoiding sequencing data bias caused by repeated UMI labeling, and significantly improving the efficiency of sequence accuracy identification of subsequent enzymatically synthesized fragments). The results showed a 100% success rate in screening 96 synthesized fragments. Details of the UMI analysis screening results are shown in Table 5.
[0071] Table 5. UMI Analysis Screening Results
[0072]
[0073] Note: "-" indicates not detected. Fragments 1 to 96 are 96 enzymatic DNA synthesis fragments, whose nucleotide sequences are shown in SEQ ID NO:1546 to SEQ ID NO:1641. The "Left-end UMI" and "Right-end UMI" columns represent the unique set of left-end UMI and right-end UMI sequences that perfectly match the preset target sequence, selected based on sequencing data by associating the left-end UMI sequence, right-end UMI sequence, and corresponding synthetic sequence. These sequences are also the forward and reverse primers used in step 7 below for UMI-specific PCR amplification of the corresponding target fragments. In this invention, there are a total of 384 forward primers available for UMI-specific PCR amplification, namely UMI-F1, UMI-F2, ..., UMI-F384, with nucleotide sequences from SEQ ID NO:778 to SEQ ID NO:1161; and a total of 384 reverse primers available for UMI-specific PCR amplification, namely UMI-R1, UMI-R2, ..., UMI-R384, with nucleotide sequences from SEQ ID NO:1162 to SEQ ID NO:1545.
[0074] 10. PCR acquisition of the target fragment Using the purified UMI amplification product obtained in step 7 as a template, UMI-specific PCR amplification was performed using the primer pairs in Table 5 (a total of 96 reaction systems were conducted, with different UMI-specific PCR amplification primer pairs used in each reaction system. For example, for fragment 1 in Table 5, the UMI-specific PCR amplification primer pairs used were UMI-F245 and UMI-R365). The PCR reaction systems and reaction procedures are shown in Tables 6 and 7, respectively.
[0075] Table 6. UMI-specific PCR reaction system
[0076] Note: For each enzymatic DNA synthesis fragment, the primers UMI-FX and UMI-RX used in the PCR reaction system can be found in Table 5.
[0077] Table 7. UMI-specific PCR reaction procedure
[0078] Example 2: Verification of the target fragment obtained after UMI screening 1. Pooling and purification of UMI-specific PCR products (1) After the 96 UMI fragments successfully screened in Example 1 were amplified by UMI-specific PCR, 5 μL of PCR products were taken and mixed into one tube, and then shaken to mix. (2) Take 200 μL of the mixed product and add 300 μL of magnetic beads (Vaht DNA Clean Beads, catalog number N411) for purification, and then reconstitute with 50 μL of water.
[0079] 2. Library construction and PE150 sequencing validation (1) Use MGI’s universal DNA library preparation kit to construct the library according to the operating requirements; (2) DNB preparation and PE150 sequencing were performed on the G99 sequencing platform using the DNBSEQ-G99RS high-throughput sequencing reagent kit (G99FCL PE150).
[0080] 3. Data analysis after machine operation The reads obtained from the PCR reaction were split using the UMI primer sequences used in 96 PCR reactions. At the same time, they were compared with the target fragment sequences. The number of reads split by the UMI primers, the number of reads that were aligned to the target fragment, and the number of reads that were consistent with the expected sequence were counted. The yield of the target fragment obtained after UMI screening was calculated (yield = number of correct sequence reads / number of target gene reads).
[0081] The results showed that among the 96 UMI-specific PCRs, 3 PCRs failed with fewer than 100 reads, while the average yield of the other 93 PCR products was 90.2%, as shown in Table 8.
[0082] Table 8. Validation results of target fragments obtained after UMI filtering
[0083]
[0084]
[0085] Note: "-" indicates no data.
[0086] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. A method for high-throughput target sequence screening of DNA libraries based on single-molecule tag UMI, characterized in that: The method includes: (A1) The DNA to be screened is grouped, and each group is defined as a DNA library; the DNA to be screened consists of a universal sequence at the 5' end of the template, a synthetic sequence, and a universal sequence at the 3' end of the template; (A2) Each DNA library corresponds to a PCR reaction system. Using the DNA library as a template, all primers in the UMI ligation primer set are added to the corresponding system. The UMI ligation primer set consists of several forward UMI ligation primers and several reverse UMI ligation primers; the forward UMI ligation primers consist of the 3' end portion of the next-generation sequencing 5' adapter, the left-end UMI sequence, and the universal sequence A; the reverse UMI ligation primers consist of the 3' end portion of the next-generation sequencing 3' adapter, the right-end UMI sequence, and the universal sequence B; all the left-end UMI sequences and all the right-end UMI sequences are different; The general sequence A is identical to the general sequence at the 5' end of the template; the general sequence B is the reverse complementary to the general sequence at the 3' end of the template. The reaction system is divided into multiple independent microdroplet reaction units using water-in-oil microdroplet generation technology. The template DNA molecule and the UMI ligation primer pair follow a Poisson distribution in the microdroplet, so that each microdroplet contains no more than one template DNA molecule and no more than one pair of UMI ligation primers. The physical separation of the microdroplets achieves a one-to-one correspondence between the template DNA molecule and the paired-end UMI combination composed of the left-end UMI sequence and the right-end UMI sequence. The first step of PCR amplification is performed in the microdroplet reaction unit, and the first step PCR product is obtained after demulsification and collection. (A3) Using the first step PCR product as a template, perform the second step PCR amplification using the UMI amplification primer pair to obtain the second step PCR product; The UMI amplification primer pair consists of a universal forward primer and a reverse primer; the universal forward primer is a 5' end adapter for complete next-generation sequencing; the reverse primer is a 3' end adapter for complete next-generation sequencing and contains a barcode sequence. The universal forward primer sequence is the same in all systems, while the reverse primers in different systems carry different barcode sequences. (A4) Perform high-throughput sequencing on the PCR product of the second step, and associate the paired-end UMI combination with the corresponding synthetic sequence based on the sequencing data, and screen out the unique pair of paired-end UMI combinations that match the target sequence. (A5) Using the PCR product of the second step as a template, PCR amplification is performed using the UMI-specific PCR primer pair composed of the left-end UMI sequence and the right-end UMI sequence in the screened paired-end UMI combination to obtain the target DNA sequence.
2. The method according to claim 1, characterized in that: In (A1), the number of DNA libraries is 1 to 8.
3. The method according to claim 1, characterized in that: In (A1), each of the DNA libraries contains N DNAs to be screened, where N is a positive integer greater than 1.
4. The method according to claim 1, characterized in that: In (A2), the number of each of the forward and reverse UMI ligation primers constituting the UMI ligation primer combination is independently selected from 200 to 1000; and / or, In (A2), the total number of the paired-end UMI combinations is less than 1 / 200 of the total copy number of the template DNA molecules in the DNA library.
5. The method according to claim 1, characterized in that: The UMI ligation primer combination consists of 384 forward UMI ligation primers shown in SEQ ID NO:1 to SEQ ID NO:384 and 384 reverse UMI ligation primers shown in SEQ ID NO:385 to SEQ ID NO:768; and / or, The nucleotide sequence of the universal forward primer in the UMI amplification primer pair is SEQ ID NO:769, and the nucleotide sequence of the reverse primer in the UMI amplification primer pair is selected from any one of SEQ ID NO:770 to SEQ ID NO:777; and / or, The forward primer in the UMI-specific PCR primer pair is selected from any one of SEQ ID NO:778 to SEQ ID NO:1161, and the reverse primer is selected from any one of SEQ ID NO:1162 to SEQ ID NO:1545.
6. The method according to any one of claims 1-5, characterized in that: In the PCR reaction system of (A2), the final concentration of the UMI ligation primer combination is 0.004 μM-0.4 μM; and / or, in the PCR reaction system of (A2), the final concentration of the template is 0.1 ng / μL-1 ng / μL.
7. The method according to any one of claims 1-5, characterized in that: In the reaction procedure for the first step of PCR amplification in (A2), the denaturation-annealing-extension cycle is repeated twice.
8. The method according to any one of claims 1-5, characterized in that: In (A4), after high-throughput sequencing of the PCR product from the second step, the method further includes a step of distinguishing different DNA libraries in the sequencing data based on specific barcode sequences that correspond one-to-one with each DNA library.
9. A primer combination, characterized in that: The primer combination includes: (B1) UMI ligation primer set; the UMI ligation primer set consists of 384 forward primers shown in SEQ ID NO:1 to SEQ ID NO:384 and 384 reverse primers shown in SEQ ID NO:385 to SEQ ID NO:768; (B2) Eight UMI amplification primer pairs; the nucleotide sequence of the forward primer in each UMI amplification primer pair is SEQ ID NO:769; the nucleotide sequences of the reverse primers in the eight UMI amplification primer pairs are SEQ ID NO:770 to SEQ ID NO:777, respectively; (B3) UMI-specific PCR primer set; the UMI-specific PCR primer set consists of 384 forward primers shown in SEQ ID NO:778 to SEQ ID NO:1161 and 384 reverse primers shown in SEQ ID NO:1162 to SEQ ID NO:1545.
10. The application of the primer combination of claim 9 in high-throughput target sequence screening of DNA libraries based on single-molecule tag UMI.