Method for batch detection of gene editing events and related sequencing library construction method thereof
By combining droplet digital PCR with barcode labeling in a one-step amplification technique, the problems of high throughput, low cost, and accuracy in gene editing detection of massive samples have been solved. This technique achieves integrated processing from sample handling to data analysis, reducing costs and the risk of cross-contamination, and improving detection efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies are insufficient for efficiently and cost-effectively performing gene editing detection on massive amounts of samples, and also suffer from cumbersome sample processing steps, risks of cross-contamination, and high barriers to data analysis.
A one-step amplification technique combining droplet digital PCR and barcode labeling was employed. This technique utilizes a droplet-based reaction system that mixes multiple samples in the same reaction tube and leverages the binding properties of LNA-modified primers at different temperatures to amplify the target sequence and perform sample-specific barcode labeling. This is then combined with a bioinformatics workflow for efficient separation and analysis.
It enables high-throughput, low-cost gene editing detection, significantly reducing the cost per sample, minimizing experimental time and cross-contamination risks, improving detection accuracy and throughput, and simplifying the process.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a method for batch detection of gene editing events and a method for constructing related sequencing libraries. Background Technology
[0002] In recent years, gene editing technologies, represented by the CRISPR / Cas9 system, have been developing at an unprecedented pace and have been widely applied in fields such as functional genomics, disease model construction, and precision medicine. As the scale of gene editing experiments continues to expand, how to efficiently, accurately, and cost-effectively detect editing events in massive amounts of samples has become one of the bottlenecks restricting the large-scale application of this technology.
[0003] Traditional gene editing detection methods, such as the T7E1 restriction enzyme digestion method and the Surveyor method, have inherent limitations, including low sensitivity (typically >1-5%), inability to provide specific mutation sequence information, and susceptibility to interference from background polymorphism. While Sanger sequencing can accurately read sequences, it has extremely low throughput, high cost, and difficulty in resolving complex chimeric mutations resulting from gene editing. These methods are insufficient to meet the need for rapid screening of tens of thousands of samples.
[0004] Next-generation sequencing (NGS) technology, with its advantages of high throughput, high sensitivity, and low cost, has become a powerful tool for analyzing gene editing results. To analyze NGS data, various bioinformatics tools have been developed, such as: BATCH-GE (Sci. Rep., 2016) and CRISPResso (Nat. Biotechnol., 2016): These tools enable batch analysis of NGS data, quantitatively assessing insertion / deletion (Indel) efficiency and homology-directed repair (HDR) efficiency. However, they require independent PCR amplification and library construction for each sample before analysis, resulting in cumbersome sample processing steps, high reagent consumption, long processing times, and the risk of cross-contamination between samples.
[0005] Hi-TOM (Sci. China Life Sci., 2019, 2024): This platform achieves high-throughput sequencing of multiple samples and multiple targets through two rounds of PCR (target amplification and barcode labeling) and provides convenient online analysis tools. However, its experimental workflow still relies on stepwise PCR reactions, failing to achieve ultimate compression of the reaction system; while its bioinformatics workflow is easy to use, its customization and flexibility are limited.
[0006] Cas-analyzer (Bioinformatics, 2017) and CRISPR-GA (Bioinformatics, 2014): These online or local tools simplify data analysis, but typically only support the analysis of single or small numbers of samples, making them unsuitable for large-scale screening projects. Furthermore, they only address data analysis issues at the "dry" experimental stage and do not revolutionize sample pretreatment processes at the "wet" experimental stage.
[0007] In summary, existing technologies share a common core problem: sample pretreatment in "wet experiments" and data analysis in "dry experiments" are separate processes. Current high-throughput detection strategies either involve numerous independent reactions at the experimental end, leading to bottlenecks in throughput, cost, and efficiency; or lack intelligent and flexible solutions at the data analysis end that match high-throughput experimental designs. In particular, how to achieve extreme compression and mixing of reaction systems from thousands of samples without introducing high levels of contamination, and to develop a corresponding automated bioinformatics workflow capable of accurately separating and analyzing this mixed data, remains an unsolved technical challenge.
[0008] Therefore, there is an urgent need in this field to develop an integrated solution that combines a revolutionary wet testing process with an intelligent dry analysis process. This solution can fundamentally simplify the operation steps, greatly increase the detection throughput, significantly reduce the unit sample cost, and ensure the accuracy of the detection results. Summary of the Invention
[0009] The technical problem to be solved by this invention is how to accurately detect high-throughput gene editing and / or how to accurately detect batch gene editing or mutations at high throughput and low cost.
[0010] To address the aforementioned technical problems, this invention first provides a method for constructing a sequencing library for (batch) detection of gene editing events, the method comprising the following steps: A1) Prepare n independent PCR reaction systems, wherein the PCR reaction system includes a nucleic acid sequence template containing the target sequence, a specific binding primer pair 1 of the target sequence with bridging sequence 1, and a primer pair 2 with m locked nucleic acid modifications and barcodes; the primer pair 2 contains bridging sequence 2, and the bridging sequence 1 and the bridging sequence 2 are complementary; the n independent PCR reaction system has n different barcodes; A2) The n independent PCR reaction systems are subjected to PCR amplification under the same reaction program to obtain n PCR products; the n PCR products are mixed to obtain the sequencing library; the reaction program includes two temperature-controlled reaction programs, program 1 and program 2, in chronological order, wherein the annealing temperature of program 1 can be 50-62℃ and the annealing temperature of program 2 can be 65-75℃. The n is a natural number greater than or equal to 2, and the m is a natural number greater than or equal to 1.
[0011] The batch can be the number of gene editing events greater than or equal to 2.
[0012] In the above method, the independent PCR reaction system can be microdroplets generated by oil encapsulation using microdroplets, and the PCR reaction system also includes a PCR buffer that can maintain the stability of the microdroplets; A1) further includes the step of mixing n of the microdroplets to obtain a mixed microdroplet library.
[0013] In the above method, n can be a natural number greater than or equal to 8 and less than or equal to 56.
[0014] In the above method, m can be a natural number greater than or equal to 1 and less than or equal to 5.
[0015] In one specific embodiment of the present invention, n is 24 and m is 4.
[0016] In the above method, the PCR buffer is preferably a commercially available reagent that can maintain droplet stability during the PCR process. In one specific embodiment of the present invention, the PCR buffer is a product from Bio-Rad with catalog number 1863010.
[0017] In the above method, the annealing temperature of procedure 1 can be 58°C, and the annealing temperature of procedure 2 can be 70°C.
[0018] In the above method, the number of amplification cycles in program 1 can be 15-25, and the number of amplification cycles in program 2 can be 20-25.
[0019] To address the aforementioned technical problems, the present invention also provides a method for detecting gene editing events and / or determining gene editing efficiency. The method may include constructing a sequencing library of nucleic acid sequences to be detected using the method described above, sequencing the sequencing library using a sequencing platform to obtain sequencing data, and performing batch detection of gene editing based on the analysis results of the sequencing data.
[0020] To address the aforementioned technical problems, the present invention also provides a composition for constructing sequencing libraries for batch detection of gene editing events, the composition comprising the primer pair 2 described above, the droplet-generating oil described above, and a PCR buffer capable of maintaining the stability of the droplets.
[0021] The above composition may also include PCR reaction components.
[0022] To address the aforementioned technical problems, the present invention also provides a gene editing event detection system and / or a gene editing efficiency monitoring system, the system including the composition described above, a microfluidic device, and a PCR instrument.
[0023] The system may also include sequencers, etc.
[0024] This invention belongs to the field of biotechnology and genetic engineering analysis and detection, specifically relating to a high-throughput method for detecting gene editing events. More specifically, this invention relates to an integrated technical solution that combines droplet digital PCR (ddPCR) technology, temperature-controlled one-step amplification and barcode labeling using locked nucleic acid (LNA) modified primers, and a bioinformatics analysis workflow developed based on a large language model, thereby achieving efficient and accurate gene editing analysis of thousands of samples simultaneously.
[0025] The rapid development of gene editing technology has posed significant challenges to the throughput, cost, and accuracy of subsequent mutation detection. Traditional detection methods, such as Sanger sequencing, suffer from low throughput and high cost, making them unsuitable for screening large-scale samples. Existing next-generation sequencing (NGS)-based detection methods typically require individual PCR amplification and barcode ligation for each sample before library construction, which is cumbersome, time-consuming, and consumes a large amount of reagents, while also posing a risk of cross-contamination between samples. Furthermore, the analysis of massive amounts of NGS data requires specialized bioinformatics knowledge, making the construction of analytical workflows highly complex. Therefore, there is an urgent need in this field for a highly efficient gene editing detection solution that can achieve a fully integrated workflow from sample processing to data analysis, significantly improve detection throughput, substantially reduce unit sample cost, and effectively prevent cross-contamination.
[0026] To address the aforementioned technical problems, this invention provides a high-throughput gene editing detection method based on droplet digital PCR and barcode labeling. The overall flowchart of the method can be found in the appendix to the specification. Figure 1 .
[0027] The core concept of this invention is to pre-mix the microdroplet reaction systems of different samples and utilize the binding characteristics of LNA-modified primers at different temperatures to complete the amplification of the target sequence and the embedding of sample-specific barcodes in one step within a single reaction tube. Finally, a dedicated bioinformatics workflow is used to efficiently split and analyze the mixed sequencing data, thereby achieving high-throughput integration of the entire process from "wet experiment" to "dry experiment".
[0028] The specific technical solution includes the following steps: Microdroplet preparation and premixing: For multiple samples to be tested, the genomic DNA, basic PCR reaction system, specific primers with bridging sequences, and barcode ligation primers modified with LNA for each sample were mixed separately.
[0029] Microfluidic devices (such as Bio-Rad QX200) are used to encapsulate each mixture into an individual droplet, forming a droplet library of a single sample.
[0030] The droplet libraries from different samples are premixed for the first time and then combined into a single PCR reaction tube. A preferred number of samples per mix is 24 to achieve an optimal balance between throughput and cross-contamination rate. The preferred buffer used for droplet generation is a commercially available reagent (such as Bio-Rad buffer) that maintains droplet stability during PCR.
[0031] One-step droplet PCR amplification and barcode labeling: The above-mentioned mixed droplet library was placed in a reaction tube for a one-step PCR reaction. This PCR employed a two-step temperature-controlled cycling program: Low-temperature amplification stage: Cycling is performed at a low annealing temperature (e.g., 50-62℃, preferably 58℃). At this time, the specific primer with bridging sequence binds to the target DNA template to achieve specific amplification of the target fragment (e.g., gene editing site), and the product ends with a bridging sequence.
[0032] High-temperature barcode labeling stage: Cycling is performed at a high annealing temperature (e.g., 65-75℃, preferably 70℃). At this time, the LNA-modified barcode ligation primers, due to their higher melting temperature (Tm), specifically bind to the bridging sequence at the end of the amplicon, integrating the unique barcode sequence into the amplicon, thus completing sample labeling.
[0033] The LNA-modified barcode ligation primers preferably contain 4 or 5 LNA modification sites to create a sufficiently large difference in Tm values between them and the unmodified primers, ensuring the specificity of the two-step reaction.
[0034] Library preparation and sequencing: After PCR, the droplets are broken to release all the amplification products carrying barcodes.
[0035] The amplification products from different mixed droplet libraries (i.e. different PCR reaction tubes) were ultramixed a second time to directly construct a library for next-generation sequencing and perform high-throughput sequencing.
[0036] Bioinformatics analysis: The sequencing data was processed using a bioinformatics analysis workflow called "DropCode". This workflow was developed iteratively through interaction with large language models (such as DeepSeek) and features automation and ease of use.
[0037] The process specifically includes: a. Quality control: Use tools such as FASTP to filter and trim the raw sequencing data.
[0038] b. Data splitting: Using a Python-based script, the sequence data is split into datasets corresponding to each sample based on a pre-stored list of barcodes.
[0039] c. Sequence alignment: Use alignment tools such as BWA-MEM to align the split sequences to the reference genome or target sequence.
[0040] d. Variation identification and statistics: SAMtools was used to process the alignment results, and a custom script was used to count the type of gene editing (such as base substitution, insertion, deletion) and its frequency at the target site for each sample.
[0041] e. Results visualization and output: Generate detailed reports containing allele sequences and frequencies, and support visualization verification using tools such as IGV.
[0042] This invention also protects a dedicated primer pair system for implementing the above-described method, and a detection kit comprising the primer pair system and a droplet generation reagent. Furthermore, this invention protects a computer-readable storage medium for performing the DropCode analysis procedure, and an integrated detection system comprising a microfluidic device, a PCR instrument, a sequencer, and a data processing device.
[0043] The beneficial effects of this invention are: Compared with the prior art, the present invention has the following significant advantages: Boosting throughput to the extreme: By combining "intradroplet labeling" with "inter-droplet premixing," physical separation and molecular indexing are perfectly integrated, achieving maximum multiplexing at the sample level. A single sequencing run can easily detect thousands of samples, increasing throughput by tens to hundreds of times compared to traditional Sanger sequencing or NGS methods requiring separate library preparation.
[0044] Significantly reduced costs: The PCR reaction system for multiple samples is compressed into a single tube, and the construction of sequencing libraries is made through hypermixing, greatly reducing the cost of enzymes, primers, kits, and sequencing.
[0045] Simplified and efficient process: The innovative "one-step" PCR avoids the cumbersome multi-round amplification and purification steps, integrating amplification and barcode labeling into a single closed system, which significantly shortens the experimental cycle and reduces the risk of human error and cross-contamination between samples.
[0046] Precise and reliable detection: Utilizing LNA-modified primers enables precise temperature control, ensuring smooth specific amplification at low temperatures and efficient barcoding at high temperatures, thus improving detection specificity and accuracy. The physical separation of microdroplets effectively reduces PCR amplification bias.
[0047] Intelligent and convenient data analysis: The DropCode workflow, developed with the assistance of a large language model, lowers the threshold for bioinformatics analysis and realizes fully automated analysis from raw data to the final mutation report, ensuring the reproducibility and efficiency of the results.
[0048] Wide range of applications: This invention has been successfully applied to high-throughput detection of gene knockout and base editing events in crops such as maize, verifying the universality and reliability of the method. It can be widely used for mutation screening in multiple fields such as plant and animal breeding, gene function research, and clinical diagnosis. Attached Figure Description
[0049] Figure 1 This is a flowchart of the detection process of the method of the present invention.
[0050] Figure 2 This study evaluates the performance of PCR-compatible buffers in terms of droplet stability, size retention, and cross-contamination control. A represents the stability assessment of droplets generated before and after PCR using four different buffers; B shows the statistical distribution of droplet diameters generated before and after PCR using the four different buffers, with the vertical axis representing droplet diameter distribution; C represents the cross-contamination rate between samples from different repeat levels, with the vertical axis representing the cross-contamination rate and the horizontal axis representing the repeat level.
[0051] Figure 3 This study evaluates the performance of LNA-modified primers in temperature-controlled barcode labeling and amplification specificity. A represents different LNA modification sites; B represents the melting curves of primers with different LNA modification sites; C represents the corresponding Tm values of primers with different LNA modification sites; and D represents temperature-controlled one-step amplification using LNA-modified primers.
[0052] Figure 4This document outlines the development and performance evaluation of DropCode for next-generation sequencing (NGS) data analysis in high-throughput gene editing research. A represents the data quality control and analysis workflow; B presents the analysis results.
[0053] Figure 5 This section showcases the scalability and accuracy of the DropEdit platform in gene editing detection. A represents the RWX-31 biallelic mutation; B represents the visualization of the RWX-31 mutation IGV; and C represents the Sanger sequencing results of the RWX-31 mutation. Detailed Implementation
[0054] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0055] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0056] The instruments and reagents used in the embodiments of this invention are from the following sources: Microfluidic device: Bio-Rad QX200 Droplet Generator PCR buffers: Buffer 1 (Bio-Rad, catalog number 1863025), Buffer 2 (Vazyme, catalog number P222-01), Buffer 3 (TSINGKE, catalog number TSE005), Buffer 4 (mei5bio, catalog number MF002). Optical microscope: Zeiss DM1000 LED.
[0057] The following examples use statistical software to process the data. The experimental results are expressed as mean ± standard deviation. One-way ANOVA test was used. P < 0.05 (*) indicates a significant difference, P < 0.01 (**) indicates a highly significant difference, and P < 0.001 (***) indicates a highly significant difference.
[0058] Example 1. Flowchart of a high-throughput gene editing detection method based on droplet digital PCR and barcode labeling 1. LNA primer design and synthesis Design a ligation primer pair with a barcode (5'-NNNNNNNNNN-3') containing four locked LNA modifications (identified by "L") in the bridging sequence region: BC-LP-4F: 5'-NNNNNNNNNN ACTGG L CCG L TC L GT L TTTAC -3' (where G) L Represents a G base modified with LNA, C L Represents a C base modified with LNA, T L (This represents a T base modified with LNA; the underlined sequence is the bridging sequence). BC-LP-4R: 5'-NNNNNNNNNN CAGG L AA L AC L AG L CTATGAC -3' (where G) L Represents a G base modified with LNA, C L Represents a C base modified with LNA, A L (This represents an A base modified with LNA; the underlined sequence is the bridging sequence). This embodiment designed 50 pairs of barcode sequences, the specific sequences of which are shown in Table 1 (the forward primer barcode sequences CL1F1-CL1F50 in Table 1 correspond to the sequence information of the barcode NNNNNNNNNN in the forward primer BC-LP-4F of the primer pair with locked nucleic acid and barcode; the reverse primer barcode sequences CL1R1-CL1R50 in Table 1 correspond to the sequence information of the barcode NNNNNNNNNN in the forward primer BC-LP-4R of the primer pair with locked nucleic acid and barcode).
[0059] Table 1. Names and sequences of barcode primers
[0060] 2. Preparation of microdroplets Fifty-six different maize (ZC01) DNA samples (partial sample sequences are shown in sequences 3-26 of the sequence listing) were selected from the maize genome, corresponding to three different introduced SNP positions of the ALS gene (its nucleotide sequence is sequence 1 in the sequence listing). PCR reaction mixtures were prepared for each sample. Each reaction system (20 μL) contained: 10 μL of 2× Taq PCR Mix, 50 ng of sample DNA template, target-specific primer pairs with bridging sequences (BS-SP-F and BS-SP-R, 0.2 μM), and barcode ligation primer pairs with four LNA modifications and barcodes designed and synthesized in step 1 (BC-LP-4F and BC-LP-4R, 0.2 μM), and nuclease-free water was added to a final volume of 20 μL.
[0061] The three primer pairs that introduce SNP position specificity are as follows: BS-SP-F1: 5'- ACTGGCCGTCGTTTTAC gaaggtgatggtgttgaac-3' (The underlined sequence is the bridging sequence, which can complement the BC-LP-4F bridging sequence in a barcode-linked primer pair containing four locked nucleic acid modifications). BS-SP-R1: 5'- CAGGAAACAGCTATGAC ccatccaggatcatgtcc-3' (The underlined sequence is the bridging sequence, which can be complementary to the bridging sequence BC-LP-4R in the link primer pair containing four locked nucleic acid modifications with barcodes). Subsequently, using a microfluidic device (Bio-Ray QX200 droplet generator), the independent PCR reaction mixtures of 56 samples were encapsulated in oil to generate independent water-in-oil droplets, forming a droplet library for each sample.
[0062] 3. Microdroplet premixing The droplet libraries from different samples were premixed for the first time, with 24 droplets from different samples grouped together and mixed in equal volumes in the same 0.2 mL PCR tube to form a 24-fold mixed droplet library. The droplets from different samples contained different barcode sequences to distinguish between them. The number of samples in the mix was 24-fold to achieve an optimal balance between throughput and cross-contamination rate. The buffer used for droplet generation was a commercially available reagent (Buffer 1) that maintains droplet stability during PCR.
[0063] 4. One-step droplet PCR amplification and barcode labeling The above-mentioned mixed droplet library, placed in a PCR reaction tube, was used for a one-step PCR reaction. This PCR employed a two-step temperature-controlled cycling program: Low-temperature amplification stage (Program 1): 25 cycles are performed at a low annealing temperature (94℃ 30s, 58℃ 30s, 72℃ 1min). At this time, the target-specific primer pair with bridging sequence binds to the sample DNA template to achieve specific amplification of the target fragment (gene editing target). The product ends with a bridging sequence.
[0064] High-temperature barcode labeling stage (Program 2): 25 cycles are performed at a higher annealing temperature (94℃ 30s, 70℃ 30s, 72℃ 1min). During this stage, the LNA-modified barcode ligation primer pair, due to its higher melting temperature (Tm), specifically binds to the bridging sequence at the end of the amplification product from the low-temperature amplification stage, integrating the unique barcode sequence into the amplicon, thus completing sample labeling. The LNA-modified barcode ligation primer contains four LNA modification sites to create a sufficiently large Tm difference between it and the target-specific primer pair, ensuring the specificity of the two-step reaction.
[0065] 5. Second hypermixing to obtain sequencing libraries After PCR, the droplets are broken to release all the amplification products carrying barcodes.
[0066] Amplification products from different mixed droplet libraries (i.e., different PCR reaction tubes) were ultramixed a second time to directly obtain a sequencing library for high-throughput detection of gene editing events. This sequencing library was then subjected to high-throughput sequencing using a next-generation sequencing platform to obtain raw sequencing data.
[0067] 6. DropCode Bioinformatics Process Analysis By interacting with the large language model DeepSeek, code is iteratively generated and optimized, ultimately integrated into an automated workflow called "DropCode". This workflow can run on Windows 10 and later systems with WSL installed, or on native Linux systems.
[0068] The "DropCode" workflow was used to process and analyze sequencing data.
[0069] The DropCode process specifically includes: Place the following files in the input folder: sequencing_data.fastq (raw sequencing data), barcodes.xlsx (barcode list), target.fasta (target sequence), and reference.fasta (reference genome).
[0070] After the process is run, perform the following steps in sequence (see process details). Figure 4 (A) Quality control: Using the FASTP software with parameters set to -Q 30 -L 70, low-quality read lengths and connectors are removed to obtain the quality-controlled data.
[0071] Data splitting: Using the Python-based script Split.py, the quality control data was split into datasets corresponding to each sample based on different barcode sequences according to the barcode list (barcodes.xlsx, some sequences are shown in Table 1).
[0072] Sequence alignment: The BWA-MEM algorithm is used to align the split read lengths with target.fasta and reference.fasta to generate a SAM file.
[0073] File processing: Use SAMtools to convert SAM files to BAM files, and then sort and index them.
[0074] Result generation: Execute the Python script Allele.py generated by LLM to count the various allele sequences and their frequencies at the target locus for each sample, generating the results as shown in the attached figure. Figure 4 The tabular report shown in B.
[0075] Loading the sorted BAM file into the Integrated Genomics Viewer (IGV) allows for a visual view of sequence alignment and verification of edit events.
[0076] Example 2. Comparison of Process Optimization of the Method of the Invention 1. Comparison of different PCR buffers The independent PCR reaction system mixtures of the 56 samples in step 2 of Example 1 were used with a microfluidic device to generate water-in-oil droplets using droplet-generating oil and four PCR buffers (Buffer 1, Buffer 2, Buffer 3 and Buffer 4) that can maintain droplet stability, to obtain a droplet library for each sample.
[0077] Then, the droplet mixing in step 3 of Example 1 and the one-step droplet PCR amplification and barcode labeling reaction in step 4 were performed to evaluate the stability of four different PCR buffers. The results are as follows: Morphological observation: Before and after the PCR reaction, small amounts of microdroplets were aspirated and their morphology was observed and photographed under a Zeiss DM1000 LED optical microscope (see [reference]). Figure 2(A). The results showed that only Buffer 1 maintained the integrity and uniformity of droplet structure before and after PCR, with no fusion or rupture observed. Buffer 2 showed a small number of fusion events ( Figure 2 (Red arrow in A), Buffers 3 and 4 became completely unstable. Figure 2 (Blue arrow in the middle A)
[0078] Particle size statistical analysis: The diameter of hundreds of droplets before and after PCR was measured using image analysis software. For example... Figure 2 As shown in Figure B, the droplet diameter of Buffer 1 was 102.77 ± 2.81 μm before PCR and 101.61 ± 2.81 μm after PCR, with no significant change (p>0.05). However, the droplet diameters of Buffer 3 and Buffer 4 changed significantly before and after PCR (p<0.0001). Therefore, Buffer 1 is the optimal choice.
[0079] 2. Comparison of fluxes for different droplet premixing methods To evaluate the cross-contamination rate under different mixed throughputs, the independent PCR reaction system mixtures of 56 samples in step 2 of Example 1 were gradient-mixed using microfluidic droplet libraries: droplets of 8, 16, 24, 32, 40, 48 and 56 different samples (all generated using Buffer 1) were taken and mixed in equal amounts in 7 0.2 mL PCR tubes to form 7 mixed droplet libraries with different multiplicity.
[0080] Seven mixed droplet libraries and an unmixed single-sample droplet library were subjected to one-step droplet PCR amplification and barcode labeling reaction as described in step 4 of Example 1.
[0081] After PCR, the droplets were lysed, the products were recovered, mixed, and then used to construct an NGS library for sequencing, yielding a total of 4.17 × 10⁴ samples. 8 Sequencing read lengths. Following the bioinformatics analysis workflow in step 6 of Example 1, sequencing reads were split back into the original samples based on barcodes, and the proportion of reads incorrectly assigned to non-target samples was calculated, defined as the cross-contamination rate. Results ( Figure 2 The results (C) show that the contamination rate is low when mixing 24 layers or less (24 layers: 4.03 ± 1.81%), while the contamination rate rises to 25.57% when mixing 56 layers. Therefore, considering both throughput and accuracy, 24 layers are preferred as the standard premixing scheme.
[0082] 3. Comparison of the effects of primer pairs with different amounts of LNA modification 3.1 Design a series of primer pairs containing different amounts of LNA modification in the bridging sequence region. CK-F / R: Unmodified control: CK-F: 5'-ACTGGCCGTCGTTTTAC-3'; CK-R: 5'-CAGGAAACAGCTATGAC-3' LP-1F / R to LP-5F / R: Each contains 1 to 5 LNA modifications (identified by "L", G). L Represents a G base modified with LNA, C L Represents a C base modified with LNA, A L Represents an A base modified with LNA, T L (Represents a T base with LNA modification) LP-1F: 5'-ACTGGCCG L TCGTTTTAC-3'; LP-1R: 5'-CAGGAAAC L AGCTATGAC-3'; LP-2F: 5'-ACTGGCCG L TC L GTTTTAC-3'; LP-2R: 5'-CAGGAAAC L AG L CTATGAC-3'; LP-3F: 5'-ACTGG L CCG L TC L GTTTTAC-3'; LP-3R: 5'-CAGGAA L AC L AG L CTATGAC-3'; LP-4F: 5'-ACTGG L CCG L TC L GT L TTTAC-3'; LP-4R: 5'-CAGG L AA L AC L AG L CTATGAC-3'; LP-5F: 5'-AC L TGG L CCG L TC L GT L TTTAC-3'; LP-5R: 5'-CAGG L AA L ACL AG L CTATG L AC-3'; 3.2 Analysis of melting temperature (Tm) Using a Bio-Rad CFX96 real-time quantitative PCR instrument, the reaction system was as follows: 10 μL of 2× Universal Green qPCR SuperMix (Takegold, catalog number AQ631), 0.2 μM template (DNA template from Example 1), 0.2 μM each of forward and reverse primers, and water to a final volume of 20 μL. A melting curve program was run from 50℃ to 80℃. Results ( Figure 3 Figures B and C show that the Tm values of the primers significantly increased with the increase of LNA modification. The Tm values of the unmodified CK-F / R were 62.08±0.20℃ and 59.08±0.20℃, respectively, while the Tm values of LP-4F / R and LP-5F / R increased by about 10-15℃, enabling them to bind efficiently at 70℃ annealing, while the unmodified primers were almost ineffective at this temperature.
[0083] 4. Verification of specificity and efficiency of one-step PCR: In step 2 of Example 1, different groups were set up to test the efficiency of the one-step PCR verification of the present invention: Control Group 1 (Two-Round Method): Two PCR reaction systems were prepared for each sample, and PCR reactions were performed in two rounds. In the first round, only the target-specific primer pair (BS-SP group BS-SP-F1 and BS-SP-R1) was used to amplify the target sequence, and the first-round PCR product was obtained. In the second round, the first-round PCR product was used as a template, and the barcode ligation primer pair (BC1-PF and BC1-PR) without LAN modification and with barcode sequence and Cy3 fluorescent label at the 5' end was used to perform PCR reaction to add barcode. BC1-PF: 5'-GGACTGCCAT ACTGGCCGTCGTTTTAC -3' (where the bold sequence is the barcode sequence, the underlined sequence is the bridging sequence, which can be complementary to the bridging sequence on BS-SP-F1, and the 5' end G band is labeled with Cy3 fluorescence). BC1-PR: 5'-ACTTCAGGAA CAGGAAACAGCTATGAC -3' (where the bold sequence is the barcode sequence, the underlined sequence is the bridging sequence, which can be complementary to the bridging sequence on BS-SP-R1, and the 5' end A band is labeled with Cy3 fluorescence). Experimental group (one-step method of this invention): Target-specific primer pairs (BS-SP-F1 and BS-SP-R1 of the BS-SP group) and barcode ligation primer pairs (BC-LP-4F2 / BC-LP-4R2) with 4 LNA modifications and barcodes and Cy3 fluorescent labeling at the 5' end were simultaneously added to independent PCR reaction systems of the same sample. A two-step temperature-controlled cycling method was used: first, a 25-cycle low-temperature amplification program 1 (94℃ 30s, 58℃ 30s, 72℃ 1min) was performed, followed by a 25-cycle high-temperature labeled amplification program 2 (94℃ 30s, 70℃ 30s, 72℃ 1min); the PCR reaction products obtained after the two-step temperature-controlled cycling PCR reaction were analyzed.
[0084] BC-LP-4F2: 5'-GGACTGCCAT ACTGG L CCG L TC L GT L TTTAC -3' (where G) L Represents a G base modified with LNA, C L Represents a C base modified with LNA, T L This represents a T base modified with LNA. The bolded sequence is the barcode sequence, and the underlined sequence is the bridging sequence, which can pair complementaryly with the bridging sequence on BS-SP-F1. The 5' G end is labeled with Cy3 fluorescence. BC-LP-4R2: 5'-ACTTCAGGAA CAGG L AA L AC L AG L CTATGAC -3' (where G) L Represents a G base modified with LNA, C L Represents a C base modified with LNA, A L The bolded sequence represents the barcode sequence, and the underlined sequence represents the bridging sequence, which can pair complementaryly with the bridging sequence on BS-SP-R1. The 5' end A is labeled with Cy3 fluorescence.
[0085] Control group 2 (non-LNA mixed method): Target-specific primer pairs (BS-SP-F1 and BS-SP-R1 in the BS-SP group) and barcode ligation primer pairs (BC1-PF and BC1-PR) without LAN modification but with added barcode sequences and Cy3 fluorescent labeling at the 5' end were added to the independent PCR reaction system of the same sample. The PCR reaction was carried out in 35 cycles at 58℃. All reaction products were analyzed by 1% agarose gel electrophoresis and Cy3 fluorescence imaging. Results ( Figure 3 (D) shows that the experimental group (one-step method of the present invention, Figure 3 The Single-reaction (LNA) of D (represented by the control group) produced the same result as the control group 1 at all 10 target sites. Figure 3 The presence of highly specific bands of comparable intensity in the two-reaction (non-LNA) bands of the control group (D) demonstrates the effectiveness of the one-step method. Figure 3 The single-reaction (non-LNA) of D can only weakly amplify some targets and has many non-specific bands, proving that without LNA modification, it is impossible to coordinate the competition between the two primer pairs (target-specific primer pairs and barcode-linked primer pairs) in a single tube.
[0086] Example 3. Practical Applications of High-Throughput Gene Editing Detection 1. Detection of the maize Waxy gene knockout library A sequencing library (RWX719, sgRNA sequence: 5'-GAGGTTCAGCTCCGGGTAGT-3') containing a gene-edited line of 636 maize Waxy genes (sequence 2 in the sequence listing) was constructed using the method of the present invention in Example 1 to detect the gene editing status of the gene-edited line. The sequencing library was sequenced using the NovaSeq platform.
[0087] The target-specific primer pairs Waxy-BS-SP-F and Waxy-BS-SP-R sequences are as follows, and the barcode-linked primer pairs with 4 LNA modifications and barcodes are the same as in Example 1 (BC-LP-4F and BC-LP-4R).
[0088] Waxy-BS-SP-F:5'- ACTGGCCGTCGTTTTAC ctgaactgaacaacgccgtc-3'; Waxy-BS-SP-R:5'- CAGGAAACAGCTATGAC cacatgcacgcaggaaaac-3'.
[0089] Using the DropCode workflow analysis from Example 1, all 636 samples were successfully detected, with an average read depth of 7141.8 ( ). Figure 5 (A)
[0090] IGV visualization showed, for example, that the target site in sample RWX-31 contained a T base deletion (-T) and an A base insertion (+A), which is consistent with the results of Sanger sequencing. Figure 5 The results in C) are completely consistent, proving the accuracy of this method.
[0091] In summary, this invention establishes a complete, efficient, and reliable high-throughput gene editing detection platform, which demonstrates excellent performance in high-throughput gene editing detection and gene editing efficiency detection from sample preprocessing to final data analysis.
[0092] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. A method for constructing a sequencing library for detecting gene editing events, characterized in that: The method includes the following steps: A1) Prepare n independent PCR reaction systems, wherein the PCR reaction system includes a nucleic acid sequence template containing the target sequence, a specific binding primer pair 1 of the target sequence with bridging sequence 1, and a primer pair 2 with m locked nucleic acid modifications and barcodes; the primer pair 2 contains bridging sequence 2, and the bridging sequence 1 and the bridging sequence 2 are complementary; the n independent PCR reaction system has n different barcodes; A2) The n independent PCR reaction systems are subjected to PCR amplification under the same reaction program to obtain n PCR products; the n PCR products are mixed to obtain the sequencing library; the reaction program includes two temperature-controlled reaction programs, Program 1 and Program 2, in chronological order, wherein the annealing temperature of Program 1 is 50-62℃ and the annealing temperature of Program 2 is 65-75℃. The n is a natural number greater than or equal to 2, and the m is a natural number greater than or equal to 1.
2. The method according to claim 1, characterized in that: The independent PCR reaction system is a microdroplet generated by oil encapsulation using microdroplets. The PCR reaction system also includes a PCR buffer that can maintain the stability of the microdroplets. A1) further includes the step of mixing n of the microdroplets to obtain a mixed microdroplet library.
3. The method according to claim 1 or 2, characterized in that: The n is a natural number that is greater than or equal to 8 and less than or equal to 56.
4. The method according to any one of claims 1-3, characterized in that: The m is a natural number that is greater than or equal to 1 and less than or equal to 5.
5. The method according to any one of claims 1-4, characterized in that: The annealing temperature of procedure 1 is 58°C, and the annealing temperature of procedure 2 is 70°C.
6. The method according to any one of claims 1-5, characterized in that: The amplification cycle number of Program 1 is 15-25 times, and the amplification cycle number of Program 2 is 20-25 times.
7. A method for detecting gene editing events and / or determining gene editing efficiency, characterized in that: The method includes constructing a sequencing library of the nucleic acid sequence to be detected using the method described in any one of claims 1-6, sequencing the sequencing library using a sequencing platform to obtain sequencing data, and performing batch detection of gene editing based on the analysis results of the sequencing data.
8. A composition for constructing sequencing libraries for batch detection of gene editing events, characterized in that: The composition comprises primer pair 2 as described in claim 1, droplet-generating oil as described in claim 2, and PCR buffer capable of maintaining the stability of the droplets.
9. The composition according to claim 8, characterized in that: The composition also includes PCR reaction components.
10. A gene editing event detection system and / or a gene editing efficiency monitoring system, characterized in that: The system includes the composition of claim 8 or 9, a microfluidic device, and a PCR instrument.