Construction method of high-throughput single-cell ATAC-seq sequencing library

By assembling the Tn5 transposase complex using a single nucleic acid sequence and employing microfluidic chip technology, the problems of DNA fragment loss and low throughput in single-cell ATAC-seq were solved, enabling high-throughput and high-sensitivity detection of open chromatin regions in single cells and reducing costs.

CN120905360APending Publication Date: 2025-11-07ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511159485.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing single-cell ATAC-seq technology based on droplet platforms suffers from problems such as reduced Tn5 transposase activity, low amount of single starting DNA leading to loss of nucleic acid fragments, and low number of captured fragments. Furthermore, traditional nucleic acid sequence assembly methods result in the loss of 50% of DNA fragments.

Method used

The Tn5 transposase complex is assembled using a single nucleic acid sequence and designed with uracil deoxyribonucleotides to avoid stem-loop structure formation. High-efficiency DNA fragmentation and coding labeling are achieved through microfluidic chips. Uracil DNA glycosylase and endonuclease are used to cut symmetrical double-stranded structures to form asymmetric adapter structures. High-throughput operations are performed using microfluidic chips or micropore array systems.

Benefits of technology

It significantly increased the number of DNA fragments detected, improved the throughput and sequencing data quality of single-cell ATAC-seq, reduced experimental costs, and enabled efficient and economical single-cell epigenetics research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120905360A_ABST
    Figure CN120905360A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a high-throughput single-cell ATAC-seq sequencing library, and belongs to the technical field of single-cell chromatin accessibility sequencing. The method comprises the steps that a single nucleic acid specific sequence with uracil deoxyribonucleotide and Tn5 transposase are assembled into a Tn5 transposase complex, and the transposase complex can be randomly combined with a target DNA sequence, cut and inserted into a DNA fragment carried by the target DNA sequence. Carrying out transposition reaction on the permeabilized cell nucleus by using a Tn5 transposase complex; the method comprises the following steps: preparing single-cell liquid drops by adopting a micro-fluidic chip, supplementing the tail ends of DNA fragments generated by enzyme digestion through PCR (Polymerase Chain Reaction), and carrying out single-cell coding marking; and finally carrying out PCR amplification and library construction. A single nucleic acid sequence joint and Tn5 transposase assembly strategy is adopted, so that the problem of PCR amplification failure caused by insertion of same joints of double-end nucleic acid sequences when nucleic acid fragments are inserted in traditional Tn5 transposase assembled by two or more nucleic acid sequences is fundamentally solved, and the recovery rate and the detection number of effective DNA fragments are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of single-cell chromatin accessibility sequencing, and particularly relates to a method for constructing a droplet single-cell ATAC-seq sequencing library of a DNA sequence binding, cleavage and insertion reaction of a Tn5 transposase complex assembled by a single nucleic acid sequence linker. BACKGROUND

[0002] Single-cell assay for transposase-accessible chromatin with high throughput sequencing (scATAC-seq) is a key technology for analyzing cell heterogeneity and epigenetic regulation, and plays an important role in discovering cell type-specific cis-regulatory elements, transcription factors and non-coding gene variation mechanisms related to diseases. The labeling of open chromatin regions depends on the binding of Tn5 transposase to the open DNA region and the insertion of Tn5 transposase itself carrying nucleic acid sequences at the cleavage site. The scATAC-seq technology platform includes microwell plates (such as Fluidigm C1) and microfluidic droplet systems (such as 10x Genomics Chromium). The method based on the well plate needs a large amount of customized modified Tn5 transposase, which is too expensive, and has limitations in sequencing throughput, while the droplet system reduces the experimental cost due to the Tn5 enzyme transposition reaction in the mixed cell population, and has the characteristics of high throughput (tens of thousands of cells can be captured in a single experiment), and is widely used.

[0003] At present, although the single-cell ATAC-seq based on the droplet platform has high throughput, most of the technologies have the problem of insufficient sensitivity. One of the reasons is that the activity of Tn5 transposase will be reduced when the reaction is carried out in the droplet, and the amount of single starting DNA is very small, which may cause the loss of nucleic acid fragments during amplification. Secondly, the systematic problems in experimental design will also lead to low number of captured fragments: the Tn5 transposase widely used at present is assembled using two kinds of nucleic acid sequences; but since the nucleic acid sequence fragments are randomly combined with the enzyme during the assembly of the Tn5 transposase complex, the use of two different nucleic acid sequences will cause the DNA fragments inserted by the Tn5 transposase complex to have a 50% probability of being symmetrical nucleic acid sequence adapters at both ends after the end sequence is filled, and such DNA will be unable to be effectively amplified due to the formation of stem-loop structure, thereby losing 50% of the DNA fragments and reducing the number of DNA fragments. At present, there is still no scATAC-seq sequencing technology that can effectively solve the bottleneck problem of limited sensitivity caused by the loss of nucleic acid fragments.

[0004] In summary, in view of the problems of insert loss of single-cell ATAC-seq sequencing technology and low effective fragments of high-throughput single-cell ATAC-seq sequencing technology, a new Tn5 transposase complex assembly strategy and amplification strategy are urgently needed, which can reduce DNA fragment loss, improve sensitivity, and be highly compatible with microfluidic chips or microwell array systems to realize the unification of high-throughput and high data quality. SUMMARY

[0005] The purpose of the present application is to provide a method for constructing a high-throughput single-cell ATAC-seq sequencing library, which realizes high-throughput capture of single cells and high-sensitivity detection of chromatin open regions.

[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solutions: In a first aspect, a method for constructing a high-throughput single-cell ATAC-seq sequencing library is provided, comprising the following steps: (1) Tn5 transposase complex assembly composed of a single nucleic acid sequence. The present application first forms a transposon by annealing a single nucleic acid sequence adapter with a uracil deoxyribonucleotide to the ME reverse sequence, and then combines the transposon with Tn5 transposase to form a Tn5 transposase complex with a single nucleic acid insertion sequence adapter that can cut DNA.

[0007] (2) Cell lysis and nuclear permeation. The single-cell suspension is directly incubated with a complex lysis and permeation solution containing a non-ionic detergent to simultaneously achieve cell membrane lysis and nuclear membrane permeation, so that the Tn5 transposase complex can efficiently bind to DNA and insert into the labeled chromatin open region.

[0008] (3) Transposition reaction of the nucleus. The permeated nucleus is mixed with the assembled Tn5 transposase complex and incubated at 37°C under constant temperature and shaking conditions to complete the chromatin cutting and adapter insertion.

[0009] (4) Single-cell droplet encapsulation and fragmentized DNA end gap filling reaction. The transposed nucleus suspension and DNA end gap filling reaction reagents (including dNTP, DNA polymerase, buffer, etc.) are encapsulated into water droplets in the microfluidic chip to form oil-in-water monodisperse droplets, which are incubated at the working temperature of the DNA polymerase to complete the end gap filling reaction after Tn5 transposase-mediated DNA fragmentation; or the single cells are distributed into each microwell of the microfluidic microwell array chip and incubated with the DNA end gap filling reaction reagents.

[0010] (5) Synchronous implementation of DNA fragment desymmetrization and single-cell coding addition. Droplets containing single nuclei are co-encapsulated with droplets carrying coding microspheres through a microfluidic chip, and the desymmetrization of DNA fragments and the release of single-cell coding sequences are completed synchronously through enzyme digestion reaction to complete the coding labeling of DNA fragments at the single-cell level; or single cells are distributed into microwells of a microfluidic microwell array chip for enzyme digestion reaction and coding labeling.

[0011] (6) Construction of coded genomic DNA (gDNA) library. After breaking the droplets with a surfactant, the product is purified using magnetic beads, PCR amplified, and a complete single-cell ATAC-seq sequencing library is constructed, and high-throughput sequencing is performed.

[0012] In step (1), the single nucleic acid linker sequence is a fixed sequence, and a uracil deoxyribonucleotide is introduced, so that specific enzyme digestion can be performed by uracil DNA glycosylase and endonuclease VIII in the future, and an asymmetric linker sequence structure is formed with the complete DNA sequence without uracil deoxyribonucleotide after end repair. This design avoids the formation of stem-loop structures by fragmented DNA carrying the same nucleic acid sequence at both ends, which cannot be amplified, thereby significantly improving the library amplification efficiency and sequencing data quality.

[0013] In step (2), the lysis and permeabilization solution can be composed of IGEPAL CA-630, Digitonin, Tween-20. The reagent rapidly lyses the cell membrane through brief incubation and forms a reversible open nuclear pore, avoiding chromatin leakage.

[0014] In step (3), the oscillation incubation (300 - 700 rpm) promotes the sufficient binding of the Tn5 transposase complex to the open regions of the chromatin, avoiding local concentration unevenness. The reaction time is controlled at 20 - 40 minutes.

[0015] In step (4), the DNA polymerase can be DeepVent, Bst, Taq or Q5 super-fidelity DNA polymerase.

[0016] In step (5), the uracil DNA glycosylase and endonuclease work together to specifically cut the symmetric double-stranded structure formed after Tn5 transposase complex fragmentation and end repair, selectively cut the original insertion sequence (containing uracil deoxyribonucleotide), and retain the complementary strand (not containing uracil deoxyribonucleotide), thereby generating an asymmetric linker structure to provide a binding site for single-cell coding sequence addition.

[0017] In step (5), the coding sequence carried by the coding microsphere is used for cell coding addition of gDNA, the single-cell coding sequence is connected with the microsphere by uracil deoxyribonucleotide, and the single-cell coding sequence can be dissociated under the joint action of uracil DNA glycosylase and endonuclease. The coding nucleic acid sequence is composed of a fixed sequence A, a cell coding sequence and a fixed sequence B; the fixed sequence A is located at the 5' end and is a universal sequencing adapter (such as Illumina Read1 sequence), which is used for subsequent sequencing primer binding and does not have non-specific binding with the Tn5 transposase adapter (ME sequence) to avoid sequencing interference; the cell coding sequence is located downstream of the fixed sequence A and is an 8-16 bp unique barcode, which is different for different coding microspheres and is used for marking cells; and the fixed sequence B is consistent with the forward direction of the single nucleic acid adapter sequence used when the Tn5 transposase is assembled, which is used for capturing gDNA and performing coding addition reaction.

[0018] In step (6), a Read2 extension primer is used for extension reaction to form a product capable of being sequenced. The DNA polymerase used in the extension reaction is Taq and its mutants, Phusion, Bst, Bst2.0, Bst2.0 WarmStart, Bst3.0, DeepVent or Phi29, etc.

[0019] In the second aspect, the application provides an ATAC-seq sequencing library constructed by the method of the first aspect, which can be used for cell line, animal tissue, plant tissue or microbial application research.

[0020] Compared with the prior art method, the method of the application has the following beneficial effects: In terms of Tn5 transposase complex assembly, the application innovatively adopts a single nucleic acid sequence assembly strategy. Compared with the traditional two nucleic acid sequence assembly method, this design avoids the problem of PCR amplification failure caused by the insertion of double-end identical adapters from the source. Therefore, the method of the application has higher DNA fragment detection number.

[0021] The method of the application uses microfluidic chip technology to realize droplet fusion coding, can realize the generation and fusion of thousands of droplets per second, and can complete the parallel processing of tens of thousands to millions of single cells in a single experiment. Compared with the single-cell ATAC-seq based on the hole plate, the method of the application improves the throughput by 2-3 orders of magnitude. At the same time, the fully automated droplet generation and fusion process avoids the cumbersome manual operation steps of the hole plate method. The droplet-based microfluidic operation system reduces the reaction system to nanoliter level, significantly reduces the reagent consumption, and greatly reduces the economic cost. The method provides an efficient, economical and reliable technical solution for single-cell epigenetic group research. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1Schematic diagram of single nucleic acid sequence Tn5 transposase complex assembly of the present application; Figure 2 Schematic diagram of key process of single cell ATAC-seq library construction of the present application; Figure 3 Nucleic acid gel electrophoresis chart of HEK293T single cell ATAC-seq data obtained by an embodiment of the method of the present application after library construction is completed; Figure 4 Histogram of insert size and its proportion of HEK293T single cell ATAC-seq data obtained by an embodiment of the method of the present application; Figure 5 Histogram of enrichment of HEK293T single cell ATAC-seq data obtained by an embodiment of the method of the present application in the transcription start site (TSS) region; Figure 6 Quality control index result chart of HEK293T single cell ATAC-seq library obtained by an embodiment of the method of the present application, the quality control index including the number of cells passing the quality control, the median of fragment number, the median of TSS enrichment score and cell density distribution.

[0023] In the figure: 1, Tn5 transposase; 2, single nucleic acid linker sequence; 21, uracil deoxyribonucleotide; 3, ME reverse sequence; 4, gDNA fragment; 5, hydrogel microsphere; 6, cell coding sequence; 61, fixed nucleic acid sequence; 62, cell coding sequence; 63, linker sequence; 7, P5 sequence; 8, T1ME sequence; 9, Read2 sequence; 10, P7 sequence. DETAILED DESCRIPTION

[0024] In order to make the personnel in the art better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application.

[0025] Embodiment 1 According to an embodiment of the present application, HEK293T cell sample single cell ATAC-seq library construction is provided, the key process is as shown in Figure 2 , including the following steps: Cell culture: HEK293T cells are cultured in DMEM / high glucose medium containing 10% v / v fetal bovine serum and 1% v / v penicillin-streptomycin, and passaged every 2-3 days.

[0026] Single nucleic acid sequence Tn5 transposase complex assembly: the assembly schematic diagram is as shown in Figure 1The single nucleic acid sequence linker 50 μΜ, the ME reverse sequence 50 μΜ, 2 x buffer (80 mM Tris-HCl PH 8.0) were mixed uniformly, and were cooled from 95 °C to 65 °C at a rate of 0.1 °C / s, and were kept for 5 minutes, and were then cooled to 4 °C at a rate of 0.1 °C / s. The transposition group solution was mixed with an equal volume of Tn5 transposase, and was incubated at 23 °C for 30 minutes. After the reaction was completed, an equal volume of 50% v / v glycerol was mixed, and was stored at -20 °C.

[0027] Single cell collection: The cells were digested with trypsin to single dispersion, washed with PBS, and filtered through a 40 μm cell strainer to remove aggregated cells.

[0028] Nuclei extraction and permeabilization: 100 μL of cell suspension was mixed uniformly with an equal volume of lysis solution (2x, containing 10 mM Tris-HCl PH 7.6, 10 mM NaCl, 3 mM MgCl2, 0.1% v / v IGE-PAL, 0.1% v / v Tween-20, 0.01% v / v Digitonin), and was left to stand on ice for 3-5 minutes. 1 mL of washing solution (containing 10 mM Tris-HCl PH 7.6, 10 mM NaCl, 3 mM MgCl2, 0.1% v / v Tween-20) was added, and was centrifuged at 500 g at 4 °C for 8 minutes, and the supernatant was removed, and 100 μL of washing solution was used to resuspend the cell pellet; finally, the cells were counted.

[0029] Cell transposition reaction: a reaction solution containing 5 μL of 2x TD buffer (PH 8.0), 0.4 μL of assembled Tn5 transposase complex, 0.1 μL of 1% v / v Digitonin (digitonin), 0.1 μL of 10% v / v Tween-20, and 1.1 μL of DEPC water was prepared. 5000 cells were taken and added to the transposition reaction solution, and the total volume of the solution was made up to 10 μL with PBS. The test tube was placed on a shaking type metal bath, and was incubated at 37 °C at a speed of 500 rpm for 30 minutes. After the reaction was completed, EDTA solution was added, and was incubated at 37 °C for 10 minutes.

[0030] Single-cell droplet preparation and DNA end-repair: The repair-lysis reagent was prepared as follows: 5 μL PCR buffer (10 x), 1.25 μL 10 mM dNTP, 1 μL DNA polymerase, 0.25 μL 20% v / v Tween-20, 3.2 μL 20% v / v Triton X-100, 2 μL 50% w / v PEG, 11.3 μL DEPC water, total 25 μL. After transposition, MgCl2 solution and EDTA were added to neutralize the single-cell suspension, and the cell suspension, repair-lysis reagent, and oil phase (7500 containing 2% surfactant) were added to the syringes, respectively, connected to the corresponding inlets of the microfluidic chip through the hose, and the water-in-oil droplets containing single cells, repair-lysis reagent were obtained at an appropriate flow rate and collected in a 0.2 mL PCR tube. The collected droplets were incubated in a thermal cycler at 74°C for 5 minutes to complete the repair of the DNA fragments.

[0031] Simultaneous implementation of DNA fragment desymmetrization and single-cell code addition: The reaction reagent was prepared containing 2 x PCR buffer, 0.5 mM dNTP, 2% v / v Tween-20, 5 U USER enzyme, and 5 U DNA polymerase, total 50 μL. Subsequently, the single-cell droplets were fused with the code extension reaction droplets using a microfluidic chip at an optimized flow rate to form water-in-oil droplets containing cell gDNA, reaction reagents, and code microspheres. The resulting droplets were placed in a thermal cycler for enzyme digestion at 37°C, extension at 72°C for 20 seconds, and final extension reaction at 70°C for 1 minute, respectively. After the reaction was completed, 40 μL of droplet mixture containing about 1500 cells was taken, the droplets were broken by adding a demulsifier (7500 containing 20% v / v PFO), and the code completed gDNA product was obtained after magnetic bead purification and elution using 20 μL DEPC water.

[0032] Extension reaction: 20 μL of purified product was mixed with 2.5 μL of 10 x PCR buffer, 0.5 μL of 10 mM dNTP mixture, 1.25 μL of 10 μM Read2 extension primer, and 1 unit of DNA polymerase to construct the reaction system, and the reaction was performed in a thermal cycler at 95°C for 1 minute, 62°C for 30 seconds, and 72°C for 2 minutes, respectively. Through this step, the effective addition of sequencing adapters was completed, and the extension product was recovered using magnetic bead purification technology after the reaction was completed. Finally, 20 μL of nuclease-free DEPC was used for elution.

[0033] PCR amplification: prepare the extension reaction reagent containing 20 μL of the reaction product of the previous step, PCR buffer (10x), dNTP, T1ME primer, TruSeq Read2 primer, DNA polymerase. Place it in a thermal cycler to perform PCR reaction. Use magnetic beads to purify the pre-amplification product, and use 20 μL of DEPC water to elute.

[0034] Addition of sequencing library tag and sequencing adapter: prepare the library construction reaction reagent containing 10 ng of the reaction product of the previous step, 25 μL of NEBNext Ultra II Q5 premix (2x), 2 μL of 10 μM P5 universal primer, and 2 μL of 10 μM P7 tag primer. Place it in a thermal cycler to incubate at 98℃ for 30 seconds, 6 PCR cycles (98℃ for 15 seconds, 70℃ for 30 seconds, and 72℃ for 60 seconds), and then incubate at 72℃ for 5 minutes. Use magnetic beads to purify the amplified gDNA, and use 40 μL of DEPC water to elute. The nucleic acid gel electrophoresis diagram of the product is shown in Figure 3 , which shows that the open chromatin is successfully captured, and the PCR amplification specificity and DNA integrity are good. After library quantification and quality control, second-generation sequencing can be performed.

[0035] The insert size (100 bp-600 bp) of the HEK293T single-cell ATAC-seq data obtained by the method embodiment of the present application is shown in Figure 4 , the HEK293T single-cell ATAC-seq data obtained by the method of the present application is significantly enriched in the transcription start site (TSS) region, and the median of the TSS enrichment score is 6.832 ( Figure 5 ); the quality control index result diagram of the HEK293T single-cell ATAC-seq library obtained by the method embodiment of the present application is shown in Figure 6 , for the 1378 single cells passing the quality control, the cell density distribution is concentrated, the median of the fragment number is 1708, and the detected single-cell fragment number is significantly higher than that detected by the existing high-throughput cell ATAC technology.

[0036] It should be noted that the present application is intended to cover any variations, uses, or adaptive changes of the present application, and the specification and examples are only considered as exemplary.

Claims

1. A method for constructing high-throughput single-cell ATAC-seq sequencing library, characterized in that, The method comprises the following steps: a single nucleic acid sequence adapter specifically recognized by Tn5 transposase is annealed with the ME reverse sequence to form a transposon, the transposon is combined with Tn5 transposase to assemble a Tn5 transposase complex with a single nucleic acid insertion sequence adapter; wherein the single nucleic acid sequence comprises a special modified base; nuclei are extracted by cell lysis and subjected to permeabilization treatment; the assembled Tn5 transposase complex with a single nucleic acid insertion sequence adapter is mixed with the permeabilized nuclei to perform a transposition reaction; the transposed single nuclei are packaged into droplets or distributed into independent microwells by a microfluidic chip, so that at most one cell is contained in each droplet or microwell, and DNA fragments generated by Tn5 transposase cleavage and insertion are subjected to end filling in the droplets or microwells; an enzyme combination capable of recognizing the special modified base is used for processing to cleave the 5' end of a nucleic acid chain containing the special modified base, break the symmetry of the two ends of the Tn5 insertion sequence, and form an asymmetric structure with a complementary chain not containing the special modified base; and single-cell coding labeling is performed. The coded genomic DNA is subjected to library construction.

2. The construction method of claim 1, wherein, The cell nuclei are extracted by cell lysis and subjected to permeabilization treatment, which comprises co-incubating a single-cell suspension with a lysis and permeabilization solution containing a non-ionic detergent to simultaneously achieve cell membrane lysis and nuclear membrane permeabilization.

3. The construction method of claim 1, wherein, The transposition reaction is mixing the assembled Tn5 transposase complex and the permeabilized nuclei, incubating under constant temperature shaking at 37°C, and performing Tn5 transposase complex binding, cleavage and sequence adapter insertion on DNA sequences.

4. The construction method of claim 1, wherein, The transposed single nuclei are packaged into droplets or distributed into independent microwells by a microfluidic chip, and DNA fragments generated by Tn5 transposase cleavage and insertion are subjected to end filling in the droplets or microwells, which comprises co-encapsulating a transposed cell nucleus suspension and a DNA end filling reaction reagent to form water-in-oil monodisperse droplets, or distributing single cells into each microwell of a microfluidic microwell array chip and reacting with the DNA end filling reaction reagent, and incubating at a DNA polymerase working temperature; wherein the DNA end filling reaction reagent comprises dNTP, DNA polymerase and buffer.

5. The construction method of claim 1 wherein, The single-cell coding labeling comprises co-encapsulating droplets containing single nuclei and droplets carrying coding microspheres, or distributing single nuclei and coding microspheres into microwells of a microfluidic microwell array chip, so that at most one nucleus and one coding microsphere are contained in each microwell; the coding microsphere comprises a single-cell coding sequence connected to a microsphere through the special modified base, and under the action of the enzyme combination, the single-cell coding primer is released by dissociation, and the symmetry of the DNA fragment is broken and the single-cell level DNA fragment coding labeling is completed by enzyme digestion reaction.

6. The construction method of claim 5, wherein, The single-cell coding sequence is sequentially connected by a universal sequencing adapter, a unique barcode and a fixed sequence, and the fixed sequence is consistent with the single nucleic acid adapter sequence in a forward direction.

7. The construction method of claim 1 or 5, wherein, The special modified base adopts deoxyuracil ribonucleotide, one or more thymine in the base of the single nucleic acid linker sequence is replaced by deoxyuracil; the enzyme combination corresponds to adopt uracil DNA glycosylase and endonuclease VIII.

8. The construction method of claim 1 wherein, The library construction of the coded genomic DNA comprises: using magnetic bead to purify the DNA product after the droplet demulsification, performing PCR amplification on the purified product, and constructing a complete single-cell ATAC-seq sequencing library.

9. An ATAC-seq sequencing library constructed by the method of any one of claims 1-8.

10. The ATAC-seq sequencing library of claim 9 is applied in the research of cell lines, animal tissues, plant tissues and microorganisms.