Indexed TN5 tagmentation-based single cell assay for transposase-accessible chromatin sequencing
Patent Information
- Application Number
- PCT/CN2024/106967
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-01-29
AI Technical Summary
Existing single-cell ATAC-seq technology struggles to balance sensitivity, accuracy, precision, throughput, and cost, limiting its application. In particular, commercial methods require specialized equipment and are costly, while traditional methods have low throughput and are complex to operate.
By employing multiple pairs of coded tagging aptamers, and through parallel large-scale tagging and molecular indexing, combined with flow cytometry sorting and PCR amplification, efficient, multiplexed, and high-throughput ATAC-seq analysis of single cells is achieved.
It achieves efficient and low-cost single-cell ATAC-seq, significantly improving throughput and accuracy, reducing the twin-cell rate, and reducing operational complexity and equipment requirements, making it suitable for single-cell analysis of complex biological systems.
Smart Images

Figure PCTCN2024106967-FTAPPB-I100001 
Figure PCTCN2024106967-FTAPPB-I100002 
Figure PCTCN2024106967-FTAPPB-I100003
Abstract
Description
INDEXED TN5 TAGMENTATION-BASED SINGLE CELL ASSAY FOR TRANSPOSASE-ACCESSIBLE CHROMATIN SEQUENCINGFIELD OF THE INVENTION
[0001] The disclosed invention is generally in the field of genetic analysis and specifically in the area of chromatin sequencing.BACKGROUND OF THE INVENTION
[0002] Studying gene regulation at the single-cell level is becoming increasingly important in biomedical research, particularly for elucidating cellular heterogeneity in intricate biological systems1-3. While single-cell transcriptomics offers insights into gene expression dynamics, single-cell epigenomics delves deeper into the regulatory mechanisms orchestrating these expression profiles1, 2, 4. The Assay for Transposase-Accessible Chromatin sequencing (ATAC-seq) has emerged as a tool to elucidate the cellular epigenetic landscape by unveiling the active regulatory, or 'open chromatin' , regions that control gene expression5, 6. This technology, independent of pre-existing knowledge about epigenetic markers or transcription factors, is advancing our understanding of the interplay between genetic and environmental factors in shaping cellular identity in mixed cell populations7, 8. Recent advancements in scATAC-seq methods, such as in-house prepared single-cell combinatorial indexing (sci) -ATAC-seq5, plate-based scATAC-seq9 and scATAC-seq in small volumes (μATAC-seq) 10, alone or integrated with other single-cell omics11-16, have provided insights into cell state transitions, functional variations, disease initiation, and progression at an ever-growing scale, expanding the depth and breadth of investigations in transcriptional regulation.
[0003] Library quality forms the foundation for data interpretation and the conclusions drawn from scATAC-seq analysis. Key technical determinants for scATAC-seq libraries include sensitivity (the ability to detect all accessible chromatin regions, as evidenced by library complexity) , accuracy (the extent of correspondence between sequenced fragments and authentic ATAC-seq signals from a single cell, but not multiple cells or background noises) , and precision (the ability discern ATAC signals specific to distinct cellular types and states) . Another important consideration is the number of cells the method is capable of analyzing and the time, manual work, and monetary cost. The existing methods struggle to strike a balance between sensitivity, accuracy, precision, throughput, and cost, impeding the widespread applications and development of scATAC-seq technology9, 10, 17. Commercial approaches utilizing droplet-based microfluidic systems or nanowell-based platforms, such as Bio-Rad ddSEQ18, Fluidigm C16, and Takara ICELL810, are costly and require specialized equipment or devices, posing significant financial and accessibility barriers, therefore restricting their use primarily to well-resourced laboratories. Although the plate-based scATAC-seq is relatively simple9 , it suffers from low throughput, accommodating only hundreds to thousands of cells. Scaling up the plate-based approach is associated with a disproportionate surge in laborious operations and costs related to PCR reactions. On the other hand, sci-ATAC-seq does augment cellular throughput to organ scale through multiple rounds of splitting and pooling but at the expense of large-scale assembly of indexed Tn5 and decreased library quality5, 19, 20. These challenges highlight the need for solutions that improve one or more of cost-efficiency, sensitivity, and throughput in the application of scATAC-seq technology.
[0004] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each claim of this application.
[0005] Throughout this specification the word “comprise, ” or variations such as “comprises” or “comprising, ” will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.
[0006] BRIEF SUMMARY OF THE INVENTION
[0007] Disclosed are compositions, systems, and methods for efficient, high throughput, and highly multiplexed forms of single-cell (sc) assay for transposase accessible chromatin (ATAC) sequencing (scATAC-seq) , referred to as indexed transposome scATAC-seq (IT-scATAC-seq) . The disclosed systems and methods make use of a set of N pairs of left and right barcoded tagmentation adapters. The barcoded tagmentation adapters each comprise a transfer strand and a non-transfer strand hybridized together, the hybridized transfer and non-transfer strands form a double stranded transposase recognition sequence and a single stranded barcoded adapter sequence, with the barcoded adapter sequence extending 5’ from the transfer strand.
[0008] The barcoded adapter sequence comprises an adapter barcode and an adapter sequence. In some forms, the adapter barcodes of each of the left and right barcoded tagmentation adapters in the set of pairs of left and right barcoded tagmentation adapters are distinct from each other. In some forms, all of the adapter sequences of all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence and all of the adapter sequences of all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.
[0009] In some forms, the set of N pairs of barcoded tagmentation adapters are comprised in a set of N sets of indexed transposomes. In some forms, for each of the N sets of indexed transposomes, the transposomes in the set comprises a tagmentation transposase and the same pair of left barcoded tagmentation adapter and right barcoded tagmentation adapter. In some forms, each of the N sets of indexed transposomes comprises a distinct pair of pair of left barcoded tagmentation adapters and right barcoded tagmentation adapters. In some forms, each of the N sets of indexed transposomes is distinguished by a pair of adapter barcodes that differs in each of the N sets of indexed transposomes.
[0010] In some forms, each of the N sets of indexed transposomes is independently formed by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes. In some forms, each of the sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters. This results in separate sets of indexed transposomes. In some forms, i *j = N, where each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.
[0011] In some forms, the set of N sets of indexed transposomes are comprised in a set of N distinct sets of tagmentated cells or nuclei. In some forms, the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by the pair of adapter barcodes that differs in each of the N sets of indexed transposomes.
[0012] In some forms, the set of N distinct sets of tagmentated cells or nuclei are disposed in a set of K reaction chambers. In some forms, each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei. In some forms, the K reaction chambers are comprised in L reaction groups, where each reaction group comprises M of the K reaction chambers.
[0013] In some forms, the set of K reaction chambers further comprise a set of P distinct pairs of forward and reverse barcoded first PCR primers. In some forms, a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is disposed in each of the M reaction chambers in a given reaction group of the L reaction groups.
[0014] In some forms, all of the forward barcoded first PCR primers comprise the same forward first primer sequence. In some forms, the forward first primer sequence matches the forward first primer matching sequence. In some forms, all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence. In some forms, the reverse first primer sequence matches the reverse first primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.
[0015] In some forms, the L reaction groups further comprises a different pair of forward and reverse barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers disposed in the contents of the M reaction chambers in each of the L reaction groups. In some forms, all of the forward barcoded second PCR primers comprise the same forward second primer sequence, wherein the forward second primer sequence matches the forward second primer matching sequence. In some forms, all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence, wherein the reverse second primer sequence matches the reverse second primer matching sequence.
[0016] In some forms, each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.
[0017] In some forms, the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination of the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise next-generation sequencing sequences.
[0018] In some forms, the disclosed methods involve producing N distinct sets of tagmentated cells or nuclei. In some forms, the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei, which results in production of N distinct sets of tagmentated cells or nuclei. In some forms, each of the N sets of tagmentated cells or nuclei is distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei.
[0019] In some forms, the disclosed method further involve sorting a single cell or nucleus of each of the N distinct sets of tagmentated cells or nuclei into K reaction chambers (e.g., wells) such that each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei. In some forms, the K reaction chambers are comprised in L reaction groups (e.g., microtiter plates) , where each reaction group comprises M of the K reaction chambers (e.g., M = K / L) .
[0020] In some forms, the disclosed methods further involve performing separate PCR amplifications of the contents of each of the K reaction chambers using a set of P distinct pairs of barcoded first PCR primers. In some forms, for each of the M reaction chambers in a given reaction group, a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is used.
[0021] In some forms, the disclosed methods further involve performing PCR amplification of the contents of the M reaction chambers in a given reaction group using a pair of barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers. In some forms, a single PCR amplification is performed in the pooled contents of all of the M reaction chambers in a given reaction group. In some forms, for each of the L reaction groups, a different pair of barcoded second PCR primers from the set of Q distinct pairs of barcoded second PCR primers is used.
[0022] In some forms, the disclosed methods further involve performing sequencing of the amplified products of the PCR amplification of the contents of the M reaction chambers.
[0023] In some forms, the disclosed methods further comprise independently forming each of the N sets of indexed transposomes by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes. In some forms, each of the N sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters, which results in formation of N separate sets of indexed transposomes. In some forms, i *j = N, where each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.
[0024] In some forms, each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.
[0025] In some forms, all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence and wherein all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.
[0026] In some forms, all of the forward barcoded first PCR primers comprise the same forward first primer sequence. In some forms, the forward first primer sequence matches the forward first primer matching sequence. In some forms, all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence In some forms, the reverse first primer sequence matches the reverse first primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.
[0027] In some forms, all of the forward barcoded second PCR primers comprise the same forward second primer sequence. In some forms, the forward second primer sequence matches the forward second primer matching sequence. In some forms, all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence. In some forms, the reverse second primer sequence matches the reverse second primer matching sequence.
[0028] In some forms, the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination of the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise next-generation sequencing sequences.
[0029] Additional advantages of the disclosed method and compositions will be set forth in part in the description which follows, and in part will be understood from the description, or can be learned by practice of the disclosed method and compositions. The advantages of the disclosed method and compositions will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings illustrate several embodiments of the disclosed method and compositions and together with the description, serve to explain the principles of the disclosed method and compositions.
[0031] Figure 1. Schematic workflow diagram of IT-scATAC-seq. The nuclei are isolated and subjected to parallel bulk transposition reactions with indexed Tn5 complexes. The transposed nuclei from each distinct tagmentation reaction are assigned individually into 384-well plates via fluorescence-activated cell sorting (FACS) . After lysis, the first round of barcoded PCR is performed to distinguish different wells. Then, the 1st PCR products are pooled for a second round of barcoded PCR to cover more plates and incorporate the standard Illumina Truseq adapter, ready for next-generation sequencing (NGS) .
[0032] Figures 2A-2F. Benchmarking of IT-scATAC-seq. Figure 2A, Species mixing experiments demonstrate FACS accuracy using equal numbers of human HEK293T and mouse NIH3T3 cells. The doublet rate without filtering is 1.19%; with low-quality cell filtering (reads < 3000) , the rate is 0.17%. Figure 2B, ATAC-seq signal enrichment around transcription start site (TSS) plotted against unique fragments for three IT-scATACseq libraries (index #1, index #2, index #3, each n = 384) . Figure 2C, ATAC-seq insert fragments frequencies distribution showing nucleosome periodicity of libraries from aggregated single cell profiles (orange line) and randomly sampled individual single cells (n=50, black lines) . Figure 2D, ATAC-seq reads distributions across TSS. Figure 2E, Scatterplots show pairwise correlation in read coverage between three indexed IT-scATAC-seq libraries across all accessible loci. Figure 2F, Genome tracks around CHD4 gene locus, showing aggregate and representative single-cell (n = 200) ATAC-seq signals.
[0033] Figures 3A-3H. Comparison of IT-scATAC-seq with other scATAC-seq methods. Figures 3A-3E, Comparison of duplication rate (sequencing depth) , unique fragments, mapping rates, fraction of reads in peaks (FRiP) , and fraction of mitochondrial (MT) DNA from IT, plate-based and C1 scATAC-seq methods. Quality control metrics obtained from plate-based scATAC-seq (9) . Figures 3F-3H, comparison of median unique fragments, median percentage of genomic fragments and median FRiP from IT and other scATAC-seq libraries. Quality control metrics obtained from benchmarking of current scATAC-seq technologies using PUMATAC (17) .
[0034] Figures 4A-4L. IT-scATAC-seq revealed a high degree of cell fate plasticity during the exit of mouse embryonic stem cells from naive pluripotency. Figure 4A, Schematic images showing mouse embryonic stem cells after undergoing a 48-hour differentiation and being subjected to IT-scATAC-seq profiling. Figure 4B, Log 10 of number of unique fragment plot against the TSS score, and cells within the upper right quadrant (n = 4, 167) passed QC and were subjected to downstream analysis. Figures 4C-4D, Fragment size distribution and enrichment of ATAC-seq signal up and downstream 2Kb to the TSS region. Figures 4E-4J, Violin plots of read alignment rate, duplication rate, log 10 (number of unique fragments) , TSS enrichment score, mitochondrial fraction of total uniquely mapped reads, and FRiP for individual samples. Figure 4K, Heatmap displaying 10 groups of 4, 167 single cells that were clustered based on ESC and mesoderm (meso) , endoderm (endo) and ectoderm (ecto) lineage scores, which were calculated from gene activity scores of a panel of marker genes for each of ESC and germ layers (see Method) Figure 4L, Scatter plots of ESC lineage scores plot against meso, endo, ecto lineage scores for each single cells.
[0035] Figures 5A-5D. High-throughput IT-scATAC-seq dissects cellular heterogeneity in human PBMCs. Figure 5A, Uniform manifold approximation and projection (UMAP) plots showing IT-scATAC-seq profiles of two human PBMC samples colored by sample origin and by unannotated cluster after dimension reduction and batch correction. Figure 5B, UMAP visualization colored by cell type identity, including B, NK and T cells, and monocytes. Figure 5C, Lineage-specific markers overlaid on UMAP embedding, including PAX5, MS4A1 and EBF1 (B cells) ; CD3G, IL7R and CD44 (T cells) ; FCGR3A (CD16) , NKG7 and IL2RB (NK cells) ; CD14, CEPB and CCR2 (monocytes) . Visualization is colored by normalized gene scores. The MAGIC algorithm was used for smoothing the drop-out noises. Figure 5D, Heatmap displaying the chromatin accessibility and gene expression of 9, 934 significant (Pearson correlation r > 0.45 and adjusted p-value < 0.1) peak-gene links.
[0036] Figures 6A-6D. Identification of epigenomic features involved in immune lineages specification. Figure 6A, Heatmap showing Z-score of normalized chromatin accessibility of 7, 208 cell identity specific marker peaks (FDR ≤ 0.1, log2 fold-change ≥1) identified over all scATAC-seq clusters. Figure 6B, MA plots displaying log2 accessibility signal foldchange (y-axis) versus mean signal fold-change (x-axis) . Indicated cell population was compared against the rest cell types. Significant differential peaks (FDR ≤ 0.1, log2 fold-change ≥1) was colored and number was indicated. Figure 6C, Heatmap showing enriched transcription factor binding motifs in marker peaks. Representative logo plots visualizing sequences of motifs associated with highest variability in chromatin accessibility among different cell groups. Figure 6D, Lineage-enriched motifs drawn on the UMAP embedding and colored based on motif deviation scores that have been imputed with MAGIC.
[0037] Figures 7A-7C. Purification, assembly and quality control of indexed Tn5 transposome complex. Figure 7A, SDS-PAGE and Commassie blue staining of purified Tn5 (53.3 kDa) , and a gradient of BSA were loaded for the standard curve plotting. Figure 7B, Standard curve was fitted based on loaded BSA and calculation of the equation and R-squared value. Figure 7C, Quality control of Tn5 activities by DNA electrophoresis of Tn5 transposome complex cleaved genomic DNA samples. Fourteen paired indexed Tn5 transposome were randomly chosen for quality control, the undigested genome DNA as negative control and conventional A / B assembled Tn5 (Tn5-A / B) as positive control; the up and down bluelines correspond to 1kb and 100bp size.
[0038] Figures 8A-8D. Sequences of IT-scATAC-seq adapters and library structure. Figure 8A, Indexed Tn5 transposome complex associated adapters, including indexed left adapters Q5XX (SEQ ID NO: 71) and indexed right adapters Q7XX (SEQ ID NO: 72) . The reverse strand of Tn5 Binding adapter is SEQ ID NO: 1. After assembly Tn5 with unique Q5XX or Q7XX, mix equal Tn5-Q5XX and Tn5-Q7XX to form a paired Tn5-Q5XX / Q7XX. The first round of cell indexing was introduced by different paired Tn5 tagmentation. Figure 8B, sequences of the first PCR primers H5XX / H7XX targeting to Q5XX (SEQ ID NO: 73) and Q7XX (SEQ ID NO: 74) , respectively. The second round of cell indexing was introduced by different combination of H5XX / H7XX. Figure 8C, sequences of the 3rd cell indexing primers T5XX / T7XX (SEQ ID NO: 75) / (SEQ ID NO: 76) targeting to H5XX / H7XX, respectively. The third round of cell indexing was introduced by different combination of T5XX / T7XX. Figure 8D, The ITscATAC-seq integrates the Illumina Truseq read 1 and read 2 adapters in the library structure ready for the standard massively parallel sequencing, which overcomes the inherent inconvenience of custom sequencing requirements and reduces substantial cost. Figure 8D shows example sequences H5XX (SEQ ID NO: 77) , T5XX (SEQ ID NO: 75) , H7XX (SEQ ID NO: 78) , and T7XX (SEQ ID NO: 79) . The double stranded sequence (SEQ ID NO: 80) represents an example sequence of a tagmentation insert sequence produced by the primers in Figure 8D.
[0039] Figure 9. Gating strategy for flow cytometry of DAPI stained nuclei.
[0040] Figures 10A-10D. Assessment of IT-scATAC-seq profiling data with bulk OmniATAC-seq . Figure 10A, Bulk OmniATAC-seq shows typical nucleosome periodicity pattern; Figure 10B, Average signal and heatmap cantered on the transcription start sites (TSS) showing ATAC-seq signals for bulk OmniATAC repeats; Figure 10C, Mapping rate of bulk OmniATAC repeats to genomic and mitochondrial regions; Figure 10D, FRiP of bulk OmniATAC repeats.
[0041] Figures 11A-11E. Quality control of two human PBMCs samples (Sample 1 cell pass quality control n = 8, 341, Sample 2 cell pass quality control n = 10, 933) . Figure 11A, TSS enrichment score plotted by the number of unique nuclear fragments for each cell. Figure 11B, Log 10 of the number of unique ATAC-seq fragments for Sample 1 PBMCs and Sample 2 PBMCs. Figure 11C, Violin plots of the distribution of single-cell TSS enrichment scores. Figure 11D, Nucleosomal periodicity indicated shown by the fragment size distribution of aggregate single-cell profiles of IT-scATAC-seq libraries of two PBMC samples. Figure 11E, Enrichment of normalized Tn5 insertions around the TSSs of the aggregate IT-scATAC single-cell profiles of two PBMC samples.
[0042] Figures 12A-12C. Comparative analysis of T cells versus B cells, monocytes versus other cell types, and NK cells versus other lymphocytes. Figure 12A. Volcano Plots showing the differentially accessible peaks between T cells (CD4 and CD8) and B cells ( and memory) , monocytes (CD14 and CD16) versus all other cell types, and NK cells compared to other lymphocytes (CD4 and CD8 T, and memory B) , highlighting the significantly upregulated and downregulated peak regions (False Discovery Rate, FDR ≤ 0.1; log2 fold-change, FC ≥ 1) . Figure 12B, Enriched motifs analysis shows motifs enriched in the up-regulated peaks (FDR ≤ 0.1, log FC ≥ 0.5) among differential accessible peaks identified in a, ranked and colored by the significance of enrichment (-log10 adjusted P value) . Figure 12C, Enrichment motif analysis shows motifs enriched in the down-regulated peaks (FDR ≤ 0.1, log FC ≤ -0.5) among differential accessible peaks identified in a, i.e. motifs enriched in marker peaks for B cells, cells other than monocytes, and non-NK lymphocytes, within each pairwise comparison.
[0043] Figure 13. Aggregate single-cell chromatin accessibility profiles and fragment counts of T, B, NK and erythroid lineage cells at corresponding marker gene loci.
[0044] Figure 14. Highly variable motifs associated with chromatin accessibility identified based on the per-cell motif activity score calculated by ChromVar algorithm.
[0045] Figures 15A-15C. Figure 15A, The timeline of IT-scATAC-seq library preparation for 10^4 cells; Figure 15B, The ratio of time-consuming on hands-On and hands-Off; Figure 15C, Estimated total cost for IT-scATAC-seq library preparation covering 10^4 cells. *The cost of FACS is calculated by the equipment use charge per hour, which varies between institutes and labs.DETAILED DESCRIPTION OF THE INVENTION
[0046] The disclosed method and compositions can be understood more readily by reference to the following detailed description of particular embodiments and the Example included therein and to the Figures and their previous and following description.
[0047] Disclosed is indexed Tn5 tagmentation-based single cell Assay for Transposase-Accessible Chromatin sequencing (IT-scATAC-seq) , a new strategy for chromatin accessibility analysis at single cell resolution with unlimited scale. This advance is based on the discovery of several techniques that can increase the efficiency, sensitivity, and throughput, and reduce the cost, of single cell Assay for Transposase-Accessible Chromatin sequencing. The techniques can include using multiple barcoded Tn5 for parallel indexed Tn5 tagmentation for first-round barcoding and distribution of cells or nuclei into individual wells of a plate preloaded with second-round barcodes. This can result in each well containing multiple cells or nuclei with distinct first-round barcodes. Performing a second round of barcoding can result in uniquely indexing for the cells or nuclei in each well. The new method is named indexed Tn5 tagmentation-based scATAC-seq (IT-scATAC-seq) . Using human cell lines, we demonstrated that this technology could achieve high library complexity, low mitochondrial contamination and high ATAC-seq signal enrichment around the transcription start site (TSS) at the single-cell level. We used IT-scATAC-seq to explore dynamics in chromatin structure as mouse embryonic stem cells transitioned out of naive pluripotency and uncovered cell-fate flexibility during early developmental stages. IT-scATAC-seq enabled a detailed examination of cellular diversity within human PBMCs, illustrating the regulatory effects of various transcription factors on the functional dynamics of hematopoietic lineage cells. These results demonstrate the versatility and effectiveness of IT-scATAC-seq in analyzing complex biological systems at the single-cell level. Without utilizing specialized equipment, this strategy allows a cost-effective library preparation with high-throughput capabilities by combining affordability and efficiency.
[0048] The disclosed methods can have one or more of the following benefits:
[0049] Scalable. IT-scATAC-seq can enhance the throughput by, for example, one order of magnitude compared with the plate-based scATAC-seq 9, increasing cell processing capacity from 102-3 to 105, achieving a level comparable to the droplet-based scATAC-seq 18 and 10x Genomics scATAC-seq. The methodology allows straightforward scalability by modularly increasing indexing at any process stage. In the demonstrated example, only 24 indexed Tn5 transposomes (6 Tn5-Q5XX and 4 Tn5-Q7XX) were assembled for the first-round indexing, 384 barcoded PCR for the second-round indexing (16 H5XX and 24 H7XX) , and 96 indexed Truseq primers (8 T5XX and 12 T7XX) . This design, which required the synthesis of only 70 oligos, is capable of capturing 884, 736 cells (24 x 384 x 96) in a massively parallel manner. With only a minimal amount of indexed Tn5 transposomes and one 384-well plate, IT-scATAC-seq easily achieves the capacity of 10X Genomics scATAC-seq, where 4,000 –5,000 cells are typically analyzed. In cases of complex tissue or organ atlas-scale profiling, aiming to map chromatin accessibility in 100,000 to 1,000,000 cells, IT-scATAC-seq can be expediently scaled up by incorporating additional indexed Tn5 transposons or PCR primers.
[0050] Accuracy. The doublet rate of IT-scATAC-seq only depends on the accuracy of cell sorting. This sets it apart from droplet-based and combinatorial indexing-based scATAC-seq, where misassignments and barcode collision are still challenging issues5, 18. Additionally, IT-scATAC-seq adopts parallel bulk tagmentation instead of single-cell individual tagmentation, effectively minimizing potential benchtop variations 6, 10.
[0051] Hands-on time. IT-scATAC-seq requires significantly less manual labor than plate-based scATAC-seq9. For example, capturing 5,000 cells using the plate-based method typically necessitates handling at least ten 384-well plates, whereas IT-scATAC-seq accomplishes the same with just one plate. Furthermore, using the Echo liquid handling system not only significantly reduces complex and labor-intensive pipetting but also lowers the risk of primer cross-contamination during PCR. The entire library preparation for 5,000 to 10,000 cells can be completed within a single day (Figure 15A) . Although FACS sorting is relatively time-consuming, the majority of processes are automated (Figure 15B) .
[0052] Cost consideration. IT-scATAC-seq reduces the per-cell cost by up to 100 times, depending on the number of cells profiled. The more cells processed, the lower the cost per cell. Unlike plate-based scATAC-seq9, which requires 3-10 μl of NEB Next High-Fidelity 2X PCR Master Mix per cell or nucleus, IT-scATAC-seq can be used to processes 24 or more indexed tagmentated nuclei in a single well of a 384-plate using only 1μl of PCR volume. This reduces the PCR reaction volume to merely 1 / 96 of that required for plate-based scATAC-seq9. As a result, the library preparation cost is substantially reduced to less than 0.1 HKD per cell (Figure 15C) , rendering it considerably more cost-effective than any existing scATAC-seq methodologies17. Moreover, all reagents required for IT-scATAC-seq are listed and can be readily prepared in-house with published protocols.
[0053] It is to be understood that the disclosed method and compositions are not limited to specific synthetic methods, specific analytical techniques, or to particular reagents unless otherwise specified, and, as such, can vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0054] Disclosed are compositions, systems, and methods for efficient, high throughput, and highly multiplexed forms of single-cell (sc) assay for transposase accessible chromatin (ATAC) sequencing (scATAC-seq) , referred to as indexed transposome scATAC-seq (IT-scATAC-seq) . The disclosed systems and methods make use of a set of N pairs of left and right barcoded tagmentation adapters. The barcoded tagmentation adapters each comprise a transfer strand and a non-transfer strand hybridized together, the hybridized transfer and non-transfer strands form a double stranded transposase recognition sequence and a single stranded barcoded adapter sequence, with the barcoded adapter sequence extending 5’ from the transfer strand.
[0055] The barcoded adapter sequence comprises an adapter barcode and an adapter sequence. In some forms, the adapter barcodes of each of the left and right barcoded tagmentation adapters in the set of pairs of left and right barcoded tagmentation adapters are distinct from each other. In some forms, all of the adapter sequences of all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence and all of the adapter sequences of all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.
[0056] In some forms, the set of N pairs of barcoded tagmentation adapters are comprised in a set of N sets of indexed transposomes. In some forms, for each of the N sets of indexed transposomes, the transposomes in the set comprises a tagmentation transposase and the same pair of left barcoded tagmentation adapter and right barcoded tagmentation adapter. In some forms, each of the N sets of indexed transposomes comprises a distinct pair of pair of left barcoded tagmentation adapters and right barcoded tagmentation adapters. In some forms, each of the N sets of indexed transposomes is distinguished by a pair of adapter barcodes that differs in each of the N sets of indexed transposomes.
[0057] In some forms, each of the N sets of indexed transposomes is independently formed by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes. In some forms, each of the sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters. This results in separate sets of indexed transposomes. In some forms, i *j = N, where each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.
[0058] In some forms, the set of N sets of indexed transposomes are comprised in a set of N distinct sets of tagmentated cells or nuclei. In some forms, the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by the pair of adapter barcodes that differs in each of the N sets of indexed transposomes. As an illustrative example, tagmentated cells / nuclei set N1 comprises adapter barcode pair (i1, j1) ; tagmentated cells / nuclei set N2 comprises adapter barcode pair (i1, j2) ; tagmentated cells / nuclei set N3 comprises adapter barcode pair (i1, j3) ; ... tagmentated cells / nuclei set Nj-1 comprises adapter barcode pair (i1, jj-1) ; tagmentated cells / nuclei set Nj comprises adapter barcode pair (i1, jj) ; tagmentated cells / nuclei set Nj+1 comprises adapter barcode pair (i2, j1) ; tagmentated cells / nuclei set Nj+2 comprises adapter barcode pair (i2, j2) ; tagmentated cells / nuclei set Nj+3 comprises adapter barcode pair (i2, j3) ; ... tagmentated cells / nuclei set N2j-1 comprises adapter barcode pair (i2, jj-1) ; tagmentated cells / nuclei set N2j comprises adapter barcode pair (i2, jj) ; tagmentated cells / nuclei set N2j+1 comprises adapter barcode pair (i3, j1) ; tagmentated cells / nuclei set N2j+2 comprises adapter barcode pair (i3, j2) ; tagmentated cells / nuclei set N2j+3 comprises adapter barcode pair (i3, j3) ; ... tagmentated cells / nuclei set N3j-1 comprises adapter barcode pair (i3, jj-1) ; tagmentated cells / nuclei set N3j comprises adapter barcode pair (i3, jj) ; ... tagmentated cells / nuclei set N(i-2) *j+1 comprises adapter barcode pair (ii-1, j1) ; tagmentated cells / nuclei set N (i-2) *j+2 comprises adapter barcode pair (ii-1, j2) ; tagmentated cells / nuclei set N (i-2) *j+3 comprises adapter barcode pair (ii-1, j3) ; ... tagmentated cells / nuclei set N (i-1) *j-1 comprises adapter barcode pair (ii-1, jj-1) ; tagmentated cells / nuclei set N (i-1) *j comprises adapter barcode pair (ii-1, jj) ; tagmentated cells / nuclei set N (i-1) *j+1 comprises adapter barcode pair (ii, j1) ; tagmentated cells / nuclei set N (i-1) *j+2 comprises adapter barcode pair (ii, j2) ; tagmentated cells / nuclei set N (i-1) *j+3 comprises adapter barcode pair (ii, j3) ; ... tagmentated cells / nuclei set Ni*j-1 comprises adapter barcode pair (ii, jj-1) ; and tagmentated cells / nuclei set Ni*j comprises adapter barcode pair (ii, jj) .
[0059] In some forms, the set of N distinct sets of tagmentated cells or nuclei are disposed in a set of K reaction chambers. In some forms, each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei. As an illustrative example, reaction chamber K1 has one cell / nucleus from each of tagmentated cell / nucleus sets N1, N2, N3, ... NN-1, and NN; reaction chamber K2 has one cell / nucleus from each of tagmentated cell / nucleus sets N1, N2, N3, ... NN-1, and NN; reaction chamber K3 has one cell / nucleus from each of tagmentated cell / nucleus sets N1, N2, N3, ... NN-1, and NN; ... reaction chamber KK-1 has one cell / nucleus from each of tagmentated cell / nucleus sets N1, N2, N3, ... NN-1, and NN; and reaction chamber KK has one cell / nucleus from each of tagmentated cell / nucleus sets N1, N2, N3, ... NN-1, and NN. In some forms, the K reaction chambers are comprised in L reaction groups, where each reaction group comprises M of the K reaction chambers.
[0060] In some forms, the set of K reaction chambers further comprise a set of P distinct pairs of forward and reverse barcoded first PCR primers. In some forms, a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is disposed in each of the M reaction chambers in a given reaction group of the L reaction groups. As an illustrative example, reaction chamber M1 of reaction group L1 uses barcoded first PCR primer pair P1; reaction chamber M2 of reaction group L1 uses barcoded first PCR primer pair P2; reaction chamber M3 of reaction group L1 uses barcoded first PCR primer pair P3; ... reaction chamber MM-1 of reaction group L1 uses barcoded first PCR primer pair PP-1; reaction chamber MM of reaction group L1 uses barcoded first PCR primer pair PP; reaction chamber M1 of reaction group L2 uses barcoded first PCR primer pair P1; reaction chamber M2 of reaction group L2 uses barcoded first PCR primer pair P2; reaction chamber M3 of reaction group L2 uses barcoded first PCR primer pair P3; ... reaction chamber MM-1 of reaction group L2 uses barcoded first PCR primer pair PP-1; reaction chamber MM of reaction group L2 uses barcoded first PCR primer pair PP; reaction chamber M1 of reaction group L3 uses barcoded first PCR primer pair P1; reaction chamber M2 of reaction group L3 uses barcoded first PCR primer pair P2; reaction chamber M3 of reaction group L3 uses barcoded first PCR primer pair P3; ... reaction chamber MM-1 of reaction group L3 uses barcoded first PCR primer pair PP-1; reaction chamber MM of reaction group L3 uses barcoded first PCR primer pair PP; ... reaction chamber M1 of reaction group LL-1 uses barcoded first PCR primer pair P1; reaction chamber M2 of reaction group LL-1 uses barcoded first PCR primer pair P2; reaction chamber M3 of reaction group LL-1 uses barcoded first PCR primer pair P3; ... reaction chamber MM-1 of reaction group LL-1 uses barcoded first PCR primer pair PP-1; reaction chamber MM of reaction group LL-1 uses barcoded first PCR primer pair PP; reaction chamber M1 of reaction group LL uses barcoded first PCR primer pair P1; reaction chamber M2 of reaction group LL uses barcoded first PCR primer pair P2; reaction chamber M3 of reaction group LL uses barcoded first PCR primer pair P3; ... reaction chamber MM-1 of reaction group LL uses barcoded first PCR primer pair PP-1; and reaction chamber MM of reaction group LL uses barcoded first PCR primer pair PP.
[0061] In some forms, all of the forward barcoded first PCR primers comprise the same forward first primer sequence. In some forms, the forward first primer sequence matches the forward first primer matching sequence. In some forms, all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence. In some forms, the reverse first primer sequence matches the reverse first primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.
[0062] In some forms, the L reaction groups further comprises a different pair of forward and reverse barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers disposed in the contents of the M reaction chambers in each of the L reaction groups. As an illustrative example, the contents of reaction chambers M1 , M2, M3, ... MM-1, and MM of reaction group L1 are amplified using barcoded second PCR primer pair Q1; the contents of reaction chambers M1 , M2, M3, ... MM-1, and MM of reaction group L2 are amplified using barcoded second PCR primer pair Q2; the contents of reaction chambers M1 , M2, M3, ... MM-1, and MM of reaction group L3 are amplified using barcoded second PCR primer pair Q3; ... the contents of reaction chambers M1 , M2, M3, ... MM-1, and MM of reaction group LL-1 are amplified using barcoded second PCR primer pair QQ-1; and the contents of reaction chambers M1 , M2, M3, ... MM-1, and MM of reaction group LL are amplified using barcoded second PCR primer pair QQ. In some forms, all of the forward barcoded second PCR primers comprise the same forward second primer sequence, wherein the forward second primer sequence matches the forward second primer matching sequence. In some forms, all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence, wherein the reverse second primer sequence matches the reverse second primer matching sequence.
[0063] In some forms, each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.
[0064] In some forms, the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination of the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise next-generation sequencing sequences.
[0065] In some forms, the disclosed methods involve producing N distinct sets of tagmentated cells or nuclei. In some forms, the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei, which results in production of N distinct sets of tagmentated cells or nuclei. In some forms, each of the N sets of tagmentated cells or nuclei is distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei.
[0066] In some forms, the disclosed method further involve sorting a single cell or nucleus of each of the N distinct sets of tagmentated cells or nuclei into K reaction chambers (e.g., wells) such that each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei. In some forms, the K reaction chambers are comprised in L reaction groups (e.g., microtiter plates) , where each reaction group comprises M of the K reaction chambers (e.g., M = K / L) .
[0067] In some forms, the disclosed methods further involve performing separate PCR amplifications of the contents of each of the K reaction chambers using a set of P distinct pairs of barcoded first PCR primers. In some forms, for each of the M reaction chambers in a given reaction group, a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is used.
[0068] In some forms, the disclosed methods further involve performing PCR amplification of the contents of the M reaction chambers in a given reaction group using a pair of barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers. In some forms, a single PCR amplification is performed in the pooled contents of all of the M reaction chambers in a given reaction group. In some forms, for each of the L reaction groups, a different pair of barcoded second PCR primers from the set of Q distinct pairs of barcoded second PCR primers is used.
[0069] In some forms, the disclosed methods further involve performing sequencing of the amplified products of the PCR amplification of the contents of the M reaction chambers.
[0070] In some forms, the disclosed methods further comprise independently forming each of the N sets of indexed transposomes by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes. In some forms, each of the N sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters, which results in formation of N separate sets of indexed transposomes. In some forms, i *j = N, where each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.
[0071] In some forms, each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers. In some forms, each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.
[0072] In some forms, all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence and wherein all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.
[0073] In some forms, all of the forward barcoded first PCR primers comprise the same forward first primer sequence. In some forms, the forward first primer sequence matches the forward first primer matching sequence. In some forms, all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence In some forms, the reverse first primer sequence matches the reverse first primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.
[0074] In some forms, all of the forward barcoded second PCR primers comprise the same forward second primer sequence. In some forms, the forward second primer sequence matches the forward second primer matching sequence. In some forms, all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence. In some forms, the reverse second primer sequence matches the reverse second primer matching sequence.
[0075] In some forms, the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination of the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise next-generation sequencing sequences.
[0076] Index variables:
[0077] N = number of distinct sets of tagmented cells or nuclei (also the number unique pairs of left and right barcoded tagmentation adapters (R = N) ; also the number of distinct combinations of the i distinct left barcoded tagmentation adapters and the j distinct right barcoded tagmentation adapters (i *j = R = N) ) .
[0078] K = number of reaction chambers into which a single tagmentated cell / nucleus from each of the N sets of tagmentated cells / nuclei is sorted.
[0079] L = number of reaction groups into which the K reaction chambers are divided.
[0080] M = number of reaction chambers in each of the L reaction groups (M = K / L) .
[0081] P = number of distinct barcoded first PCR primer pairs (also the number of distinct combinations of the k distinct forward barcoded first PCR primers and the l distinct reverse barcoded first PCR primers (k *l = P) ; also the number of reaction chambers in each of the L reaction groups (M = P) ) .
[0082] Q = number of distinct barcoded second PCR primer pairs (also the number of distinct combinations of the m distinct forward barcoded second PCR primers and the n distinct reverse barcoded second PCR primers (m *n = Q) ; also the number of reaction groups (L = Q) ) .
[0083] R = number unique pairs of left and right barcoded tagmentation adapters (R = N) ;
[0084] i = number of distinct left barcoded tagmentation adapters.
[0085] j = number of distinct right barcoded tagmentation adapters.
[0086] k = number of distinct forward barcoded first PCR primers.
[0087] l = number of distinct reverse barcoded first PCR primers.
[0088] m = number of distinct forward barcoded second PCR primers.
[0089] n = number of distinct reverse barcoded second PCR primers.
[0090] Table 3. List of Components
[0091] The disclosed systems and methods can make use of a set of N pairs of left and right barcoded tagmentation adapters 100. The barcoded tagmentation adapters can each comprise: a transfer strand 110 and a non-transfer strand 120 hybridized together, the hybridized transfer and non-transfer strands form:
[0092] a double stranded transposase recognition sequence 130 and
[0093] a single stranded barcoded adapter sequence 140, with the barcoded adapter sequence extending 5’ from the transfer strand.
[0094] The barcoded adapter sequence can comprises: an adapter barcode 150 and an adapter sequence 160.
[0095] In some forms, the adapter barcodes of each of the left and right barcoded tagmentation adapters in the set of pairs of left and right barcoded tagmentation adapters are distinct from each other.
[0096] In some forms, all of the adapter sequences of all of the left barcoded tagmentation adapters 104 comprise the same forward first primer matching sequence 170 and all of the adapter sequences of all of the right barcoded tagmentation adapters 106 comprise the same reverse first primer matching sequence 180.
[0097] In some forms, the set of N pairs of barcoded tagmentation adapters are comprised in a set of N sets of indexed transposomes 600. In some forms, for each of the N sets of indexed transposomes, the transposomes in the set comprises a tagmentation transposase 650 and the same pair of left barcoded tagmentation adapter and right barcoded tagmentation adapter. In some forms, each of the N sets of indexed transposomes comprises a distinct pair of pair of left barcoded tagmentation adapters and right barcoded tagmentation adapters. In some forms, each of the N sets of indexed transposomes is distinguished by a pair of adapter barcodes that differs in each of the N sets of indexed transposomes.
[0098] In some forms, each of the N sets of indexed transposomes is independently formed by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes. In some forms, each of the sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters. This results in separate sets of indexed transposomes. In some forms, i *j = N, where each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.
[0099] In some forms, the set of N sets of indexed transposomes are comprised in a set of N distinct sets of tagmentated cells or nuclei 700. In some forms, the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by the pair of adapter barcodes that differs in each of the N sets of indexed transposomes.
[0100] In some forms, the set of N distinct sets of tagmentated cells or nuclei are disposed in a set of K reaction chambers 800. In some forms, each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei. In some forms, the K reaction chambers are comprised in L reaction groups 900, where each reaction group comprises M of the K reaction chambers.
[0101] In some forms, the set of K reaction chambers further comprise a set of P distinct pairs of forward barcoded first PCR primers 204 and reverse barcoded first PCR primers 206. In some forms, a different pair of barcoded first PCR primers 200 from the set of P distinct pairs of barcoded first PCR primers is disposed in each of the M reaction chambers in a given reaction group of the L reaction groups.
[0102] In some forms, all of the forward barcoded first PCR primers comprise the same forward first primer sequence 270. In some forms, the forward first primer sequence matches the forward first primer matching sequence. In some forms, all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence 370.
[0103] In some forms, all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence 280. In some forms, the reverse first primer sequence matches the reverse first primer matching sequence. In some forms, all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence 380.
[0104] In some forms, the L reaction groups further comprises a different pair of forward barcoded second PCR primers 404 and reverse barcoded second PCR primers 406 from a set of Q distinct pairs of barcoded second PCR primers disposed in the contents of the M reaction chambers in each of the L reaction groups.
[0105] In some forms, all of the forward barcoded second PCR primers comprise the same forward second primer sequence 470, wherein the forward second primer sequence matches the forward second primer matching sequence.
[0106] In some forms, all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence 480, wherein the reverse second primer sequence matches the reverse second primer matching sequence.
[0107] In some forms, each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers.
[0108] In some forms, each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers.
[0109] In some forms, each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.
[0110] In some forms, the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination of the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise forward next-generation sequencing sequence 574 and reverse next-generation sequencing sequence 576.
[0111] Figures 8A-8D illustrate generalized examples of barcoded tagmentation adapters and barcoded PCR primers. Figure 8A show example sequences of Q5XX and Q7XX barcoded tagmentation adapters. The two strands of Q5XX are SEQ ID NOs: 71 and 1. The two strands of Q7XX are SEQ ID NOs: 72 and 1. Figure 8B shows example sequences of H5XX and H7XX barcoded first PCR primers. H5XX are SEQ ID NO: 73. H7XX are SEQ ID NO: 74. Figure 8C shows example sequences of T5XX and T7XX barcoded second PCR primers. T5XX are SEQ ID NO: 75. T7XX are SEQ ID NO: 76. Figure 8D shows example sequences H5XX (SEQ ID NO: 77) , T5XX (SEQ ID NO: 75) , H7XX (SEQ ID NO: 78) , and T7XX (SEQ ID NO: 79) . The double stranded sequence (SEQ ID NO: 80) represents an example sequence of a tagmentation insert sequence produced by the primers in Figure 8D.
[0112] 1. Barcoded Tagmentation Adapters
[0113] The barcoded tagmentation adapter (100) has two strands (atransfer strand (110) and a non-transfer strand (120) ) that a double stranded portion and a single stranded extension. The double stranded portion includes or consists of a transposase recognition sequence (130) . The double stranded portion can include a sequence that is not part of the transposase recognition sequence (130) , but that is not preferred. The transposase recognition sequence (130) is at the 5’ end of the non-transfer strand (120) and the 3’ end of the transfer strand (110) .
[0114] The single stranded extension of the barcoded tagmentation adapter (100) can be referred to as the barcoded adapter sequence (140) . The barcoded adapter sequence (140) includes or consists of two portions (an adapter barcode (150) and an adapter sequence (160) ) . The single stranded extension can include extra sequences that are not part of either the adapter barcode (150) or the adapter sequence (160) , but that is not preferred. Such extra sequences can occur at the 5’ end of the single stranded extension and / or between the adapter barcode (150) and the adapter sequence (160) .
[0115] The adapter barcode (150) can be any suitable length. Generally, the length of the adapter barcode (150) need only be long enough (a) to encode as many different barcodes as different barcoded tagmentation adapters (100) will be used together and (b) to be bindingly distinguishable from all other barcodes and all other sequences present in all of the barcoded tagmentation adapters (100) , barcoded first PCR primers (200) , and barcoded second PCR primers (400) as will be used together. The sequences of the adapter barcodes (150) can be freely chosen so long as they serve these purposes. Generally, adapter barcodes (150) can have 4 to 25 nucleotides, preferably 8 to 15 nucleotides.
[0116] The adapter sequence (160) includes or consists of the first primer matching sequence (170) . The first primer matching sequence (170) is at the 5’ end of the barcoded adapter sequence (140) . The first primer matching sequence (170) can be any suitable length. Generally, the length of the first primer matching sequence (170) need only be long enough to bind to and allow its complement to effectively and specifically facilitate priming by the corresponding barcoded first PCR primer (200) . The sequences of the adapter sequences (160) and first primer matching sequences (170) can be freely chosen so long as they serve this purpose. Generally, adapter sequences (160) and first primer matching sequences (170) can have 8 to 30 nucleotides, preferable 12 to 25 nucleotides.
[0117] 2. Barcoded First PCR Primers
[0118] The barcoded first PCR primer (200) includes three portions (afirst primer sequence (270) , a first barcode (250) , and a second primer matching sequence (370) ) . The first primer sequence (270) is at the 3’ end of the barcoded first PCR primer (200) and has a sequence matching the sequence of its corresponding first primer matching sequence (170) . The first barcode (250) is between the second primer matching sequence (370) and the first primer sequence (270) on the barcoded first PCR primer (200) . The second primer matching sequence (370) is at the 5’ end of the barcoded first PCR primer (200) and has a sequence matching the sequence of its matching second primer sequence (470) .
[0119] The barcoded first PCR primer (200) can include extra sequences that are not part of any of the first primer sequence (270) , first barcode (250) , and second primer matching sequence (370) . Such extra sequences can be between the first primer sequence (270) and the first barcode (250) or between the first barcode (250) and the second primer matching sequence (370) .
[0120] The first barcode (250) can be any suitable length. Generally, the length of the first barcode (250) need only be long enough (a) to encode as many different barcodes as different barcoded first PCR primers (200) will be used together and (b) to be bindingly distinguishable from all other barcodes and all other sequences present in all of the barcoded tagmentation adapters (100) , barcoded first PCR primers (200) , and barcoded second PCR primers (400) as will be used together. The sequences of the first barcodes (250) can be freely chosen so long as they serve these purposes. If any portion of a next-generation sequencing sequence is a part of a barcoded first PCR primer, that next-generation sequencing sequence would be included and not freely chosen. Generally, first barcodes (250) can have 4 to 25 nucleotides, preferably 8 to 15 nucleotides.
[0121] 3. Barcoded Second PCR Primers
[0122] The barcoded second PCR primer (400) includes three portions (asecond primer sequence (470) , a second barcode (450) , and all or a part of a next-generation sequencing sequence (570) ) . The second primer sequence (470) is at the 3’ end of the barcoded second PCR primer (400) and has a sequence matching the sequence of its corresponding second primer matching sequence (370) . The second barcode (250) is between the all or a portion of a next-generation sequencing sequence (570) and the second primer sequence (470) on the barcoded second PCR primer (200) .
[0123] The barcoded second PCR primer (400) can include extra sequences that are not part of any of the second primer sequence (470) , second barcode (450) , and next-generation sequencing sequence (570) . Such extra sequences can be between the second primer sequence (470) and the second barcode (450) or between the second barcode (450) and the next-generation sequencing sequence (570) . In some forms, the next-generation sequencing sequence (570) overlaps all or part of the second primer sequence (470) , the second barcode (450) , or both.
[0124] The second barcode (450) can be any suitable length. Generally, the length of the second barcode (450) need only be long enough (a) to encode as many different barcodes as different barcoded second PCR primers (400) will be used together and (b) to be bindingly distinguishable from all other barcodes and all other sequences present in all of the barcoded tagmentation adapters (100) , barcoded first PCR primers (200) , and barcoded second PCR primers (400) as will be used together. The sequences of the second barcodes (450) can be freely chosen so long as they serve these purposes. Any portion of a next-generation sequencing sequence that is a part of a barcoded second PCR primer would be included and would not be freely chosen. Generally, second barcodes (450) can have 4 to 25 nucleotides, preferably 8 to 15 nucleotides.
[0125] The sequences of the barcoded second PCR primers and the barcoded first PCR primers are generally chosen based on the sequencing platform used. Thus, the special primer and adapter sequences used by the chosen sequencing platform will be incorporated into the barcoded second PCR primers and the barcoded first PCR primers as needed for the sequencing platform to function on the barcoded and amplified tagmentated sequences. The special primer and adapter sequences used by the chosen sequencing platform that are incorporated into the barcoded second PCR primers and the barcoded first PCR primers are referred to herein as next-generation sequencing sequences.
[0126] In some forms, the barcoded first PCR primers (200) , the barcoded second PCR primers (400) , or a combination of the barcoded first PCR primers (200) and the barcoded second PCR primers (400) include all or a part of a next-generation sequencing sequence (570) .
[0127] Table 4 shows the sequences of examples of Q5XX, Q7XX, H5XX, H7XX, T5XX, and T7XX shown in Figure 8. Nucleotides 22-40 of SEQ ID NO: 71 (Q5XX) and nucleotides 23-41 of SEQ ID NO: 72 (Q7XX) are the transposase recognition sequence. Nucleotides 14-21 of SEQ ID NO: 71 (Q5XX) , nucleotides 15-22 of SEQ ID NO: 72 (Q7XX) , nucleotides 34-41 of SEQ ID NO: 73 (H5XX) , nucleotides 35-42 of SEQ ID NO:74 (H7XX) , nucleotides 30-37 of SEQ ID NO: 75 (T5XX) , nucleotides 25-32 of SEQ ID NO: 76) (T7XX) , nucleotides 32-39 of SEQ ID NO: 77 (H5XX) , nucleotides 33-40 of SEQ ID NO: 78 (H7XX) , and nucleotides 25-32 of SEQ ID NO: 79 (T7XX) are the barcodes of the barcoded tagmentation adapters and barcoded PCR primers.
[0128] Nucleotides 1-13 of SEQ ID NO: 71 (Q5XX) and nucleotides 1-14 of SEQ ID NO: 72 (Q7XX) are the first primer matching sequences of the barcoded tagmentation adapters, which match the first primer sequences of their respective barcoded first PCR primers. Nucleotides 42-54 of SEQ ID NO: 73 (H5XX) and nucleotides 43-56 of SEQ ID NO: 74 (H7XX) are the first primer sequences of the barcoded first PCR primers, which match the first primer matching sequences of their respective barcoded tagmentation adapter.
[0129] Nucleotides 1-22 of SEQ ID NO: 73 (H5XX) and nucleotides 1-23 of SEQ ID NO: 74 (H7XX) are the second primer matching sequences of the barcoded first PCR primers, which match the second primer sequences of their respective barcoded second PCR primers. Nucleotides 38-59 of SEQ ID NO: 75 (T5XX) and nucleotides 33-55 of SEQ ID NO: 76 (T7XX) are the second primer sequences of the barcoded second PCR primers, which match the second primer matching sequences of their respective barcoded first PCR primers.
[0130] Nucleotides 40-55 of SEQ ID NO: 77 (H5XX) and nucleotides 41-57 of SEQ ID NO: 78 (H7XX) are the first primer sequences of the barcoded first PCR primers, which match the first primer matching sequences of their respective barcoded tagmentation adapter. Nucleotides 1-20 of SEQ ID NO: 77 (H5XX) , nucleotides 1-21 of SEQ ID NO: 78 (H7XX) are the second primer matching sequences of the barcoded first PCR primers, which match the second primer sequences of their respective barcoded second PCR primers. Nucleotides 40-59 of SEQ ID NO: 75 (T5XX) and nucleotides 33-53 of SEQ ID NO: 79 (T7XX) are the second primer sequences of the barcoded second PCR primers, which match the second primer matching sequences of their respective barcoded first PCR primers.
[0131] Nucleotides 1-33 of SEQ ID NO: 73 (H5XX) , nucleotides 1-34 of SEQ ID NO: 74 (H7XX) , nucleotides 38-59 of SEQ ID NO: 75 (T5XX) , nucleotides 33-55 of SEQ ID NO: 76 (T7XX) , nucleotides 1-31 of SEQ ID NO: 77 (H5XX) , nucleotides 1-32 of SEQ ID NO: 78 (H7XX) , and nucleotides 33-53 of SEQ ID NO: 79 (H7XX) , are in the Illumina Read sequences.
[0132] Nucleotides 1-59 of SEQ ID NO: 80 correspond to the forward barcoded second PCR primer of SEQ ID NO: 75 (T5XX) . Nucleotides 38-70 of SEQ ID NO: 80 correspond to Illumina Read #1 sequence. Nucleotides 40-94 of SEQ ID NO: 80 correspond to the forward barcoded first PCR primer of SEQ ID NO: 77 (H5XX) . Nucleotides 79-121 of SEQ ID NO: 80 correspond to the left barcoded tagmentation adapter of SEQ ID NO: 81 (Q5XX) . Nucleotides 122-178 of SEQ ID NO: 80 correspond to the insert sequence. Nucleotides 179-222 of SEQ ID NO: 80 correspond to the right barcoded tagmentation adapter of SEQ ID NO: 82 (Q7XX) . Nucleotides 206-262 of SEQ ID NO: 80 correspond to the reverse barcoded first PCR primer of SEQ ID NO: 78 (H7XX) . Nucleotides 231-264 of SEQ ID NO: 80 correspond to Illumina Read #2 sequence. Nucleotides 244-296 of SEQ ID NO: 80 correspond to the reverse barcoded second PCR primer of SEQ ID NO: 79 (T7XX) .
[0133] The example H5XX, H7XX, T5XX, and T7XX primers include special sequences used by the Illumina sequencing platform (as a preferred example) , which sequences constitute next-generation sequencing sequences in the disclosed primers and systems. These sequences can be substituted to match special sequences of other sequencing platforms, such as BGI sequencers. Thus, all the example H5XX, H7XX, T5XX, and T7XX primers can be substituted to include sequences matching the special sequences of other sequencing platforms.
[0134] Table 4. Example Adapter and Primer Sequence
[0135] Table 5. Example Adapter and Primer Components
[0136] Examples
[0137] Example 1: Demonstration of an example of IT-scATAC-seq, a new strategy for chromatin accessibility analysis at single cell resolution with unlimited scale.
[0138] Single-cell chromatin accessibility sequencing (scATAC-seq) has emerged as an indispensable tool in single-cell genomics, joining with single-cell transcriptomics to uncover epigenetic mechanisms governing transcriptional regulation. However, the high cost, fluctuating performance, and limited capacity have hampered its wider use. To address these challenges, we developed the Indexed Tn5 tagmentation-based scATAC-seq (IT-scATAC-seq) . IT-scATAC-seq has a streamlined, robust, and semi-automated workflow that allows for scalable and cost-effective library preparation without sacrificing sensitivity. The optimized library structure ensures compatibility with all Illumina next-generation sequencing platforms (Illumina sequencers take up over 90% marketing portion, until BGI Genomics (华大基因) recently released its own sequencer) . Using this technology, we examined chromatin landscape changes in mouse ESCs exiting from naive pluripotency and revealed high cell fate plasticity during early embryogenesis. IT-scATAC-seq also facilitated a comprehensive analysis of cellular heterogeneity in human peripheral blood mononuclear cells (PBMCs) , depicting the roles of diverse transcription factors in hematopoiesis. These findings highlight the versatility and effectiveness of IT-scATAC-seq in analyzing complex biological systems at single-cell resolution.
[0139] Materials and Methods
[0140] Cell culture. The HEK293T and mouse fibroblast cell line NIH / 3T3 were routinely maintained in High-glucose Dulbecco’s modified Eagle’s medium (DMEM) containing 10%Fetal Bovine Serum (FBS) and 1%Penicillin / Streptomycin. The B6 murine ES Cell line was cultured on gelatin-coated dishes in 2i medium composed of High-glucose DMEM supplemented with 15%stem-cell qualified FBS, 2mM GlutaMAX, 1x non-essential amino acids (NEAA) , 0.1mM β-mercaptoethanol, 1000 U / ml recombinant mouse LIF (Merck Millipore) , 2i (1μM PD032591 and 3μM CHIR99021, MedChemExpress) and 1%Penicillin / Streptomycin. The cell culture basic medium and supplements were purchased from Thermofisher. All the cells were cultured at 37 ℃ in 5%CO2 and tested negative for mycoplasma infection using the PCR method by the Centre for PanorOmic Sciences, Li Ka Shing Faculty of Medicine.
[0141] Purification of transposase Tn5. The pTBX1-Tn5 plasmid was a kind gift from Dr. Rickard Sandberg (Addgene #60240) . Briefly, pTBX1-Tn5 plasmid was transformed into competent E. coli C3013 cells (NEB, Cat #C2527I) and induced with 250 μl 1M Isopropyl β-d-1-thiogalactopyranoside (IPTG) at 23℃ for 5 hours. Cell pellet was resuspended in 60ml HEGX buffer (20 mM HEPES buffer pH 7.2, 1.0 M NaCl, 1 mM ethylenediaminetetraacetic acid (EDTA) , 10%v / v glycerol, 0.2%v / v triton X-100 and 10mM PMSF) and sonicated using Covaris sonicator with 10 cycles of 30s on and 30s off, 40%duty. The cleared Tn5-CBD protein fraction was enriched with chitin resin (NEB, Cat #S6651S) the a cold room for 2 hours and further washed with 200 ml of HEGX buffer. The Tn5 protein was released by 100 mM dithiothreitol (DTT) cleavage, concentrated with PierceTM Protein 30K MWCO Concentrators and dialyzed twice in 1L 2X HEPES dialysis buffer (100 mM HEPES pH 7.2, 0.2M NaCl, 0.2 mM EDTA, 20%w / v glycerol, and 2 mM dithiothreitol (DTT) . After dialysis, the Tn5 was equilibrated with pure glycerol to 60%concentration. The final Tn5 was quantified by SDS-PAGE and Commassie Blue Staining using the fitting curve plotted by standard BSA. Tn5 was quantified as 1.6 μg / μl, approximately 30μM in this study.
[0142] Preparation of indexed Tn5 transposome complex. Dissolve the indexed adapters and Tn5 reverse adapters (ordered from IDT, detailed sequences are summarized in Tables 1 and 5) with annealing buffer (10 mM Tris-HCl pH 8.0, 50mM NaCl, 2mM EDTA) to make 200μM stock. Prepare 15 μl of individual adapter with 15 μl reverse adapter in 200 μl PCR tube and anneal in a thermocycler as follows: 98 ℃ for 10 min, and slowly cool down to 23 ℃with -0.1 ℃ / s. Mix the annealed adapter with 100 μl 30 μM Tn5 and 70 μl coupling buffer (100 mM HEPES-NaOH, 500 mM NaCl, 50%v / v Glycerol, 0.5 mM EDTA, 2 mM DTT) , and incubate in thermomixer at 25℃, 1000 rpm for one hour. The indexed Tn5 transposome was prepared by mixing 20 μl of the paired two Tn5- adapters with 80 μl coupling buffer and the resulting Tn5 transposome complex was 5 μM and can be stored at -20℃ without activity loss more than one year.
[0143] Table 1. Adapter and Primer Sequences
[0144] Quality control of assembled transposases. Prepare 1 μl 300 ng / μl genomic DNA, 4 μl 5xTAPS-DMF buffer (50 mM TAPS-NaOH pH 8.2, 25 mM MgCl2, 50%DMF) , 13 μl H2O, 2 μl assembled Tn5. Incubate at 55 ℃ for 10 min, followed by adding 2 μl 10X STOP buffer (2%SDS, 40 mM EDTA) and quench at 37℃ 15min to dissociate Tn5 from DNA. Add 5 μl 6x Loading dye and run 1.5%DNA gel. The majority of tagmentated DNA size were less than 1,000bp indicating the assembled transposases are quantified for downstream experiments. In this study, we randomly picked indexed Tn5 for quality examination.
[0145] Isolation of human peripheral blood mononuclear cells (PBMCs) . About 5 microliters of blood were taken from two donors with the informed consent given by guardians and human tissue procurement under the guidance of ethical regulations of Guangdong Provincial People's Hospital and The University of Hong Kong. The Peripheral Blood Mononuclear Cells were isolated using Ficoll-Paque based gradient separation and frozen at liquid nitrogen until usage. The frozen cells were rapidly thawed in a water bath at 37 ℃ and transferred to a 15 ml tube containing 10 ml prewarmed medium. The cell suspension was then centrifuged at 500g for 5 min at room temperature to pellet down. The cell pellet was resuspended in 1ml prewarmed medium and 10 μl were taken to check the cell viability. Make sure >90%cells were still alive before the formal nuclei extraction.
[0146] mESCs-epiblast differentiation. mESCs were cultured in 2i medium to 60-80%confluency and dissociated into single cells using 0.1%Trypsin. After washing twice with PBS buffer, the mESCs were then resuspended in fresh embryoid body media (withdraw of 2i and mLIF) and seeded on gelatin-coated plate. The spontaneously differentiated cells were collected on day 2 and ready for IT-scATAC-seq.
[0147] Species Mixing Experiment. For each single cell, metrics generated by SAMtools idxstats were used to calculate the fraction of reads mapped to human (hg10) and mouse (mm10) genomes. If the fraction mapped to human genome > 0.85, the cell was identified as human, if the fraction mapped to human genome > 0.15, the cell was identified as mouse; the cell otherwise is classified as a doublet. The same were calculated for high-quality cells (unique mapped reads > 3000) .
[0148] IT-scATAC-seq library preparation. The whole procedure was similar to previously reported plate-based scATAC-seq, except that multiple indexed Tn5 transposome complex were used in parallel. Briefly, the nuclei were prepared following OmniATAC protocol and resuspended in 0.33x PBS buffer. Next, 76 μl nuclei (~5x10^4) were aliquoted to several 1.5ml Eppendorf DNA Tubes, add 20 μl 5xTAPS-DMF buffer and 4 μl indexed Tn5 transposome complex. The tagmentation reactions were performed on a thermomixer at 37℃, 800 rpm for 30 min. Then, 500 μl stop buffer (1xPBS, 1%BSA and 20mM EDTA) was added to quench the reaction on ice for 10 min and transferred to FACS tubes. DAPI was added at a final concentration of 1μg / μl to stain the nuclei before sorting. During the tagmentation, 350nl lysis buffer (10mM Tris-HCl, 10mM NaCl, 0.2%SDS and 0.2μ / ml Proteinase K) were distributed to 384-plates by 550 Liquid Handler (Labcyte) and centrifuge at 2000 rpm for 3 min. Different index-tagmentated nuclei can be sorted into the same well. After sorting, the plates were centrifuged at 2000 rpm for 3 min and the nuclei were lysed at 55℃ for 10 min and 100nl 10%Triton X100 were added to quench SDS. Then, 25 nl 20 μM indexed forward and reverse primers (H5XX and H7XX) , as well as 0.5 μl NEB Next High-Fidelity 2X PCR Master Mix, were added to each well. The first round of amplification was performed following 72℃ 5min, 98 ℃ 30s; 12 cycles of 98 ℃ 20 s, 63℃ 30s; 72 ℃ 1min; 72 ℃5min, 4 ℃ hold. The PCR product was pooled by centrifuge, followed by purification using MinElute PCR purification kit and eluted with 50 μl nuclease-free H2O. The undesired fragments, primers and adapters were removed by Exo I digestion, 1.0 x AMPure XP beads selection, and eluted with 25 μl nuclease-free H2O. The Truseq P5 / P7 adapters containing different barcoded primers were added by the second PCR with another 3 amplification cycles. After another double AMPure XP beads selection (0.5x / 0.5x) , the libraries were sent for quality control and next-generation sequencing by ANOROAD GENOME.
[0149] Data pre-processing. Cutadapt 4.5 was used to remove TruSeq Index 1 (i7) Adapters and Index 2 (i5) Adapters at both 5’-and 3’-end of each read. The barcode sequences were then extracted from 5’-end of each read sequence and appended to read headers of the paired-end reads by Cutadapt 4.5 with --rename=' {id} CB: Z: {r1. adapter_name} {r2. adapter_name} ' -e 0.04 --no-indels --action=trim’ , and adapter sequences and name are specified in fasta files with parameters -g and -G. The trimmed and barcode-extracted reads were mapped to corresponding reference genome, including human (GRch38) for HEK293T and human PBMCs, mouse (mm10) for epiblast differentiation, and human (GRCh38) and mouse (mm10) hybrid genome assembly for species-mixing experiments, using BWA-MEM v. 0.7.17. The bam file is then sorted by the cell barcode (CB) tag and split into BAM file by CB using SAMtools 1.17. For the demultiplexed BAM file for each single cell, MarkDuplicates of Picard Tools 3.1.0 was used to mark and remove duplicated reads. The deduplicated BAM were then merged using SAMtools into a deduplicated single cell aggregate BAM file for downstream analysis. Using deduplicated single cell aggregates BAM file, accessible chromatin regions (peaks) were called using MACS2, with parameters -f BAMPE -g hs --shift -75 --extsize 150 --nomodel --call-summits --nolambda --keep-dup all -p 0.01 -B.
[0150] Bulk omniATAC-seq processing and visualization. Bulk HEK293T omniATAC-seq were processed as previously described8. CollectInsertSizeMetrics of Picard Tools 3.1.0 were used to calculate the fragment size. Deeptools (version 3.5.2) ’s were used to computeMatrix and plotHeatmap to visualize the enriched signal around ± 5 kb up and downstream to TSS region and to estimate TSS score.
[0151] IT-scATAC-seq library quality control and comparison with previous methods. Deduplicated single cell BAM files without any filtering were processed to generate single-cell quality control (QC) metrics as previously described in the plate-based approach. Briefly, Library size and deduplication rate was estimated using the metric file generated by Picard MarkDuplicates in the previous deduplication. SAMtools idxstats was used to calculate number of unique fragments and percentage of mitochondrial fragments. Mapping rates was estimated by SAMtools flagstats. Fraction of read in peaks (FRiP) was calculated by determining the ratio of reads in peak regions to total reads for each sample, using bedtools instersect 2.31.0 for peak-reads intersection (bedtools intersect -a $BAM -b $PEAKS -u -nonamecheck -ubam) and SAMtools (samtools view -c) for read counting. To compare these metrics with the plate-based and C1-based approaches, quality control metrics of K562 and mouse embryonic stem cell were obtained from https: / / github. com / dbrg77 / plate_scATAC-seq / . To compare with other scATAC approaches, quality metrics (after down sampling to 40,000 reads per cell) calculated for all individual samples were obtained from previous study17.
[0152] Visualization of IT-scATAC-seq signal and correlation analysis. For HEK293T single-cell aggregate and 200 randomly selected HEK293T single-cell profiles, we use Bamcoverage of Deeptools suite (version 3.5.2) was used to first normalize total reads to 10,000, 00 and generate BigWig and Bedgraph files with the parameters --scaleFactor 10,000,000 / reads_number --binSize 50. Integrative Genomics Viewer (version 2.11.1) was used to visualize the tracks. We use Deeptool’s multiBigwigSummary and plotCorrelation to calculate the Pearson correlation coefficient between the normalized single-cell aggregate, selected single-cell profiles and bulk omniATAC-seq of HEK293T. The Bedgraph were used to generate the pairwise correlation in read coverage between three indexed HEK293T IT-scATAC-seq libraries.
[0153] Analysis of IT-scATAC-seq EpiLC library. The BAM file of deduplicated single cell aggregates was converted to a fragment BED file with Tn5 insertion centering correction using fragment function in Sinto 0.10.0 with parameters --collapse_within. The fragment file was compressed using bgzip 1.18 and indexed by tabix 1.18. The R package ArchR was used to create an Arrow file using the fragment file, the quality control criteria were set as TSS > 5 and number of unique fragments > 1,000. The TSS enrichment score for each single cell was calculated at the same time. When creating the Arrow file, a ‘Title Matrix’ counting the number of fragments that fall into genome-wide 500-bp bins, and a ‘Gene Score Matrix’ counting calculating each gene’s accessibility score based on tile distance, gene size, and Tn5 insertions, and normalizing these scores across all genes. An ArchR project was subsequently created using the Arrow file for downstream analysis. We used the gene score in ‘Gene Score Matrix’ to infer the gene activity. Unsupervised clustering of the single-cell EpiLC data was modified from the previously described method27. Briefly, a selected panel of marker genes for ESC and three germ layers were obtained from25-27. For each of the four lineage types -embryonic stem cells (ESC) , endoderm, mesoderm, and ectoderm -we calculated the standard deviation of gene score across all single cells. We then identified the top 50 genes with the highest standard deviation as lineage-specific markers for each cell type. To perform lineage scoring, we normalized the gene scores of these marker genes for each cell, thereby mitigating the impact of differential gene accessibility levels on scoring. For each cell, we computed the average normalized gene score of its lineage markers to derive its lineage score. Unsupervised clustering using the ward. D method was performed to generate the heatmap that depicted the transient cell states characterized by the lineage state.
[0154] Dimensionality reduction, clustering analysis for human PBMCs. We use ArchR to generate Arrow files and created ArchR project as described earlier. Iterative latent semantic indexing was performed using ArchR’s function addIterativeLSI to reduce dimensions, and the harmony algorithm was used to correct the samples’ batch-effect using ‘addHarmony’ function. Cells were clustered using addClusters by using Suerat’s FindClusters method and then embedded using UMAP by the ‘addUMAP’ function. The marker genes were identified by the ‘getMarkerFeatures’ function with Gene Score Matrix calculated by ArchR. MAGIC algorithm was used by applying the ‘addImputedWieghts’ function to impute gene scores by smoothing signal across neighboring cells and were used to visualize selected lineage marker genes’ gene scores overlayed on the UMAP embedding. Cell identity was annotated by constrained cross-platform linkage of scATAC-seq cells with scRNA-seq cells with ArchR’s ‘addGeneIntegrationMatrix’ function using firstly with scRNA-seq dataset of hematopoietic differentiation and then with CITE-seq reference of PBMCs. By integrating the results with marker genes, the cell clusters were annotated and merged.
[0155] Marker peaks identification and marker motif analysis. For pseudo-bulk replicates, the chromatin accessible peaks set was created using the MACS2 function ‘addGroupCoverages. ’ Peaks were then called using the MACS2 function ‘addReproduciblePeakSet’ , and the ‘addPeakMatrix’ function appended the count matrix of the combined peakset to the Arrow file. The ‘addDeviationsMatrix’ function was then used to compute per-cell deviations across all motif annotations, utilizing the enriched marker peaks for the clusters to build the ‘MotifMatrix’ deviation matrix. After that, getMarkerFeatures and peakAnnoEnrichment function with cutoff at FDR <= 0.1 and Log2FC >= 0.5 were used to identify peaks that were specific to individual clusters as well as clusters that belong to the same cell lineages, respectively. The top motif deviation matrix was computed using ArchR's addDeviationsMatrix Based on the ChromVAR algorithm. With ‘getVarDeviations’ , the top variable motifs were found, and the MotifPlot function of the Signac R package was used to plot the logo of the top 30 variable motifs.
[0156] Results
[0157] Benchmark of IT-scATAC-seq.
[0158] The Tn5 transposase was purified according to the published protocol21, and its concentration was determined using the standard BSA plotting curve before assembly with indexed adapters (Figure 7A-7B) . In vitro tagmentation assay demonstrated high activity of in-house assembled indexed Tn5 transposome complexes (Figure 7C) .
[0159] In the IT-scATAC-seq, nuclei are isolated following the refined Omni-ATAC procedure to minimize mitochondrial DNA contamination22 and divided into equal parts for parallel bulk transposition reactions with indexed Tn5 complexes (number of reactions = N) (Figures 1 and 8A) . The transposed nuclei from each tagmentation reaction are assigned individually into 384-well plates via fluorescence-activated cell sorting (FACS) (Figure 9) . After sorting, each well houses N uniquely indexed nuclei. Nuclei are lysed in the pre-loaded buffer containing sodium dodecyl sulphate (SDS) and proteinase K. The lysis process is then quenched, followed by DNA amplification using indexed PCR primers pre-loaded for the second-round barcoding. Subsequently, PCR products are pooled for another round of PCR to incorporate the standard Illumina Truseq adapter, preparing them for next-generation sequencing (NGS) (Figures 1 and 8B-8D) . With the acoustic liquid transfer system Labcyte Echo550 Acoustic Liquid Handler, all steps in 384-well plates are robotically automated to avoid intricate pipetting.
[0160] The accuracy of IT-scATAC-seq was evaluated by a species-mixing experiment, where equal amounts of human cell line HEK293T and mouse cell line NIH3T3 were mixed and analyzed. Under a comparatively low sequencing depth, we captured all 4, 608 input cells. Of these, 4, 553 cells were predominantly marked with either mouse (n = 2, 249) or human (n = 2, 304) , with 55 cells identified as doublets. This rendered an accuracy rate of 98.6% (Figure 2A) . When filtered out the low-quality cells (number of unique fragments < 3000) , only 6 cells were classified as doublets (99.8%accuracy) (Figure 2A) , suggesting that increasing sequencing depth to capture more ATAC-seq fragments or using stringent cut-off could further improve the accuracy.
[0161] We proceeded with IT-scATAC-seq on the widely used HEK293T cell line. For the HEK293T IT-scATAC-seq, we utilized three distinctly indexed Tn5 transposome complexes for first-round indexing. Nuclei tagged with first-round indices were sorted into a 384-well plate for subsequent library preparation. From this process, three IT-scATAC-seq libraries were produced, labelled as index #1, index #2, and index #3, each consisting of 384 cells and collectively contributing to 1, 152 uniquely indexed cells. All 1, 152 cells met the established quality control criteria (TSS enrichment score > 5 and more than 1,000 unique fragments per cell) , with consistently high library complexity and signal-to-noise ratio across the indices –median number of unique fragments per cell 21, 975 for index #1, 53,626 for index #2, and 54, 136 for index #3; median TSS enrichment scores at 19.8 for index#1, 18.4 for index#2, and 19.2 for index #3 (Figure 2B) . Majority of cells (99.4%) were qualified even under more rigorous criteria (TSS > 7 and unique fragments > 10,000) (Figure 2B) . The fragment size distribution displayed clear nucleosome periodicity patterns (Figure 2C) . A strong enrichment around transcriptional start sites (TSS) was observed for sampled single-cell profiles and aggregate single-cell profiles (Figure 2D) . These results demonstrated the very high quality of IT-scATAC libraries.
[0162] In HEK293T, we performed bulk OmniATAC-seq for three biological repeats. These libraries exhibited high quality, characterized by a typical periodic fragment pattern, minimal mitochondrial contamination, high TSS scores, and high FRiP scores (Figure 10A-10D) , qualifying them as suitable reference libraries. The pseudo-bulk profiles of IT-scATAC-seq libraries (index #1, #2, and #3) revealed a robust correlation (Pearson correlation coefficient, r = 0.85) with the bulk omniATAC-seq library, and a significant similarity among themselves (r = 0.99) (Figure 2E) . Additionally, 200 randomly selected IT-scATAC profiles displayed high congruence with bulk data, with Pearson correlation coefficients between 0.49 and 0.94 (Table 2) . Moreover, these aggregated single-cell profiles closely resembled bulk signals, both in accessible genomic regions and at specific genomic loci (Figure 2F) .
[0163] Table 2. Correlation of randomly peaked 20 single cell ATAC profiles to aggregates and bulk OmniATAC repeats.
[0164] Comparison of IT-scATAC-seq with other scATAC-seq methods.
[0165] To further evaluate the library quality of IT-scATAC-seq, we compared HEK293T IT-scATAC data with plate-based scATAC-seq and Fluidigm C1 scATAC-seq9. Under relatively lower sequencing depth (median duplication rate 54-57%of IT-scATAC-seq versus nearly 95%of plate-based and C1 scATAC-seq) (Figure 3A) , IT-scATAC-seq showed 100%genomic mapping rate (Figure 3B) and the highest library complexity, as indicated by the higher number of unique fragments per cell (Figure 3C) . While the proportion of sequencing reads mapped to mitochondrial DNA was comparable (Figure 3D) , a significantly higher percentage of reads aligned with peaks of chromatin accessibility were observed in IT-scATAC-seq with a median FRiP score of 80% (Figure 3E) . Notably, IT-scATAC-seq demonstrated more consistent quality control metrics at the single-cell level, with data more tightly clustered around the median, unlike the broader variability seen in the other datasets.
[0166] A recent study has systematically benchmarked common scATAC-seq approaches, including Bio-Rad ddSEQ, HyDrop, and several variants of the 10x Genomics scATAC-seq assay17. An extended comparison with these methods revealed that IT-scATAC-seq either outperformed or matched the performance of commercial 10x Genomics scATAC-seq and Bio-Rad ddSEQ in key metrics, including unique fragments per cell, nuclear fragment proportion, and FRiP (Figure 3F-3H) . Collectively, these findings establish IT-scATAC-seq as a straightforward, robust, and highly sensitive approach for single-cell chromatin profiling.
[0167] IT-scATAC-seq detects high plasticity of cell fate during stem cell differentiation.
[0168] Stem cell differentiation is inherently dynamic, characterized by a spectrum of heterogeneous intermediate states23. We used the IT-scATAC-seq to examine the remodeling of epigenetic signatures occurring as mouse embryonic stem cells (ESCs) exit from pluripotency. mouse embryonic stem cells (ESCs) were subjected to a two-day differentiation period to primed epiblast-like stem cells (EpiLCs) , a transient interval that has acquired competence for differentiating towards downstream mesodermal (meso) , endodermal (endo) and ectodermal (ecto) lineages (Figure 4A) . Analysis of the EpiLCs IT-scATAC-seq library showed that 4, 167 out of 4, 608 cells (90.4%) passed our stringent quality control parameters, including a threshold for TSS score (> 5) and a minimum of 1,000 unique fragments (Figure 4B) . From these cells, we harvested a total of 131.81 million fragments, with the fragment size distribution displaying a typical nucleosomal pattern and an enrichment of signal around the TSS region (Figures 4C-4D) . Despite a relatively modest sequencing depth, indicated by a 44%duplication rate (Figure 4E) , we observed a 98%read alignment rate and a median of 18, 058 unique fragments per cell, indicating high library complexity (Figures 4F & 4J) . Additionally, cells demonstrated an average TSS enrichment score of 14.35, low mitochondrial contamination (median 1.62%) and high FRiP score (median 0.69) (Figures 4G-4I) . These results showcased the very high quality of IT-scATAC-seq library preparation.
[0169] Previous research demonstrated that the EpiLCs are competent to differentiate into all three germ layers24. However, the mechanisms by which ESCs transit to EpiLCs and how gene cascades are selectively activated to determine the cell fate have not been fully elucidated by scRNA-seq alone25. Gene score quantifies local chromatin within and around a gene, weighted by distance and gene size, and normalized across the entire gene region to infer the potential regulatory impact on gene expression. We leveraged a targeted panel of lineage marker genes 25, 26; and calculated lineage scores for each cell based on the average marker activity, similar to previously described itChIP-seq27. Unsupervised clustering of 4, 617 single cells, guided by normalized accessibility profiles for lineage-specific markers, revealed 10 distinct clusters (Figure 4K) . These clusters entailed a spectrum of cellular state, ranging from naive ESCs with pronounced accessibility in ESC marker regions to cells exhibiting increased accessibility across both ESC and multiple lineage markers, suggestive of priming for germ layer differentiation, and cells with commitment to specific germ layers (Figure 4K) .
[0170] Notably, a substantial number of cells occupied intermediate states, including those with relatively low accessibility for all four categorical markers compared with primed ESCs, indicative of formative ESCs (Figure 4K) . Additionally, cells with transitional combinations of marker gene accessibility, such as meso-endo-ecto (n=171) and endo-ecto (n=120) , underlined the multifaceted nature of to epiblast-like transition. Interestingly, we detected a group of cells exhibiting uniformly lower gene activity of all three germ-layer markers as well as ESC markers than any other groups. The pseudo-temporal trajectory plots, which were generated based on germ-layer versus ES scores, elucidated an increase in mesoderm accessibility scores concurrent with the exit from the naive ESC state, while endoderm and ectoderm accessibility scores showed phases of stability or minor decline bore exit pluripotency (Figure 4L) , implying a potential epigenetic constriction before definitive lineage commitment. These observations collectively echoed the concept of cell fate plasticity, where cells manifest the potential to transit between states and adapt to developmental cues through dynamic epigenetic remodeling28-30. The simultaneous activation of multiple lineage markers post-ESC state exiting substantiated this plasticity, highlighting the cells' adaptability and the non-fixed nature of development. These results also highlighted the prowess and need for advanced high-resolution single-cell technology like IT-scATAC-seq in dissecting regulatory mechanisms underlining dynamic and transient cell states.
[0171] IT-scATAC-seq dissects cellular heterogeneity across high-throughput human peripheral blood mononuclear cells.
[0172] One significant advantage of IT-scATAC-seq is that it offers an easy and streamlined approach for scaling up to profile over 10,000 cells. To assess its efficacy in resolving diverse cell types and dissecting the epigenomic heterogeneity at a single-cell level within complex primary tissues, IT-scATAC-seq was applied to cryopreserved human PBMCs from the peripheral blood of two donors. Despite shallow sequencing depth (~10,000 raw reads per cell) , a total of 19, 274 cells were retained after filtering low-quality nuclei, consisting of 8, 341 derived from Sample 1 and 10, 933 cells derived from Sample 2 (Figure 11A) . The median number of unique fragments for each sample was 3, 281 and 3, 817, and the median TSS enrichment score was 24.96 and 22.91, respectively (Figure 11B & 11C) . The combined single-cell insert size distributions from both samples revealed discernible nucleosome banding patterns and high signal enrichment around the TSS region (Figures 11D-11E) . After batch effect correction, uniform manifold approximation and projection (UMAP) analysis and visualization revealed the presence of 17 distinct immune cell populations (Figure 5A) . We discerned and merged the clusters by integrating single-cell gene expression profiles from two human PBMCs and hemopoiesis scRNA-seq datasets31, 32. This resulted in nine cell clusters that were identified as B memory (n = 677) , B (n = 1, 416) , CD4 T (n = 5, 422) , CD8 T (n = 2, 840) , CD14 monocyte (n = 2, 038) , CD16 monocytes (n = 1, 289) , natural killer (NK, n = 4, 828) , erythroid (n = 141) and common lymphoid progenitors (n = 623) subtypes (Figure 5B) . By overlaying the per-cell gene activity on UMAP embedding, we showed the aggregated accessibility of cell-lineage-specific genes, including PAX5, MS4A1 and EBF1 for B cells; CD3G, IL7R, CD44 for T cell lineages; CD16a (FGCR3A) , NKG7 and IL2RB for NK cells; and CD14, CEBPB, and CCR2 for monocytes (Figure 5C) , corresponding to the interpreted cluster identities. For each cluster, we called peaks and created a union set of 68, 975 reproducible accessible peaks. A total of 9, 935 peaks were correlated with gene expression (Pearson correlation r > 0.45 and adjusted p-value < 0.1) , indicating a high consistency between peak and genes (Figure 5D) . From the called peak set, we identified 7, 208 differentially accessible peak regions (FDR ≤ 0.1, LogFC ≥ 1) (Figures 6A-6B and 12A) . Lineage-specific accessibility of CD3, MS4A1, NKG7 and GATA1 were found for T cell, B cell, NK cell and erythroid, respectively (Figure 13) .
[0173] For differentially accessible regions, ChromVAR was used to calculate motif activity, and 870 motifs were identified as associated with variability among different cell populations (Figures 6C and 14) . Motif enrichment analysis discovered that RUNX family genes (RUNX1, RUN2, RUNX3) , CBFB, TCF7, and FOXO family (FOXO3, FOXO4) were among top-ranked differentially enriched in differential peaks of T-lineage cells relative to B-lineage cells (Figures 6C, 6D, and 12B) . Ets family members SPI1 and SPIB were previously found important in the regulation of genes important for B cell antigen receptor signalling33. Consistently, SPI1 and SPIB, together with other transcription factors, including RELA, NKB1, IRF1, and POUF2F, were found as the top enriched TFs in B cells (Figures 6C, 6D, and 12C) . SMARCC1, FOS, JUN and C / EBP families demonstrated increased accessibility in monocytes; whereas NK cells, when compared to other lymphocytes, were enriched for RUNX2, RUNX1, MGA, EOMES and TBX family binding motifs (Figures 6C, 6D, 12B, and 12C) . These findings are in concordance with those observed at the single-cell transcriptome34, 35 and bulk scale36, underscoring that IT-scATAC signatures can adequately characterize cell type-specific gene-regulatory programs.
[0174] Discussion
[0175] In this paper, we report a simple and robust method, IT-scATAC-seq, that achieves high-throughput single-cell chromatin accessibility at a very low cost. IT-scATAC-seq mainly involves four steps: 1) the assembly of indexed Tn5 transposome; 2) parallel bulk nuclei tagmentation; 3) sorting different indexed nuclei into the same well for barcoded PCR; 4) sample pooling and PCR for Illumina Truseq adapter addition. For 104 cells, the whole processing of IT-scATAC-seq can be finished within one day, and the quality of IT-scATAC-seq is better or at least comparable to the robust plate-based scATAC-seq 9 and commercial 10x Genomics ATAC-seq17.
[0176] Compared to the plate-based approach9, IT-scATAC-seq notably enhances the throughput to encompass thousands or even tens of thousands of cells while significantly reducing both labor and consumable costs in downstream procedures. The utilization of the automated Labcyte Echo550 Acoustic Liquid Handler during the second indexing step obviates the necessity for intricate pipetting, therefore substantially mitigating the risk of primer cross-contamination. Additionally, the library structure contains the commonly used Truseq adapters, making it compatible with massively parallel sequencing. These improvements substantially reduce the sequencing cost. Details of improvement can be found elsewhere herein.
[0177] Here, we applied IT-scATAC-seq to examine the chromatin accessibility profile of stem cell differentiation and biological samples of high heterogeneity. IT-scATAC-seq profiling of mESCs differentiation offered profound insights into the regulatory dynamics during the transition out of pluripotency. The analysis showed that cells were predominantly in a transitional intermediate state with diversified chromatin accessibility profiles for the ESC and germ-layer signature genes. These findings aligned with the cell-fate plasticity and underscored the dynamic gene regulation schemes during early development. In human PBMC analysis, IT-scATAC-seq demonstrated its capability to distinguish various lineages of hematopoietic cells. Through motif analysis, we elucidated the unique transcription-factor-driven regulatory mechanisms characterizing each blood cell type. Further co-accessibility analysis may help reveal the cooperative and coordinated cis-regulatory modules, uncovering how they shape the immune landscape.
[0178] The high-quality and robust data produced by IT-scATAC-seq position this method as a powerful tool in elucidating the regulatory features and transcriptional regulation mechanisms of novel and rare cell populations. It may facilitate our comprehension of cellular heterogeneity across a spectrum of biological and medical scenarios.
[0179] Moreover, the strategy of bulk cells / nuclei indexed tagmentation and single-cell / nucleus sorting used in IT-scATAC-seq also opens a new direction for single-cell omics study. It can be easily extended to other single-cell omics alone or combined, such as single-cell whole genome sequencing (scWGS-seq) 37, single-cell Cleavage Under Targets and Tagmentation (scCUT&Tag-seq) 38 and single-cell HiC-seq 39 or multimodal profiling11, 14-16, 40-42.
[0180] References
[0181] 1. Lai, B., Gao, W., Cui, K., Xie, W., Tang, Q., Jin, W., Hu, G., Ni, B., and Zhao, K. (2018) . Publisher Correction: Principles of nucleosome organization revealed by single-cell micrococcal nuclease sequencing. Nature 564, E17. 10.1038 / s41586-018-0690-1.
[0182] 2. Jin, W., Tang, Q., Wan, M., Cui, K., Zhang, Y., Ren, G., Ni, B., Sklar, J., Przytycka, T.M., Childs, R., et al. (2015) . Genome-wide detection of DNase I hypersensitive sites in single cells and FFPE tissue samples. Nature 528, 142-146. 10.1038 / nature15740.
[0183] 3. Li, Y.E., Preissl, S., Hou, X., Zhang, Z., Zhang, K., Qiu, Y., Poirion, O.B., Li, B., Chiou, J., Liu, H., et al. (2021) . An atlas of gene regulatory elements in adult mouse cerebrum. Nature 598, 129-136. 10.1038 / s41586-021-03604-1.
[0184] 4. Stergachis, A.B., Neph, S., Reynolds, A., Humbert, R., Miller, B., Paige, S.L., Vernot, B., Cheng, J.B., Thurman, R.E., Sandstrom, R., et al. (2013) . Developmental fate and cellular maturity encoded in human regulatory DNA landscapes. Cell 154, 888-903. 10.1016 / j. cell. 2013.07.020.
[0185] 5. Cusanovich, D.A., Daza, R., Adey, A., Pliner, H.A., Christiansen, L., Gunderson, K.L., Steemers, F.J., Trapnell, C., and Shendure, J. (2015) . Multiplex single cell profiling of chromatin accessibility by combinatorial cellular indexing. Science 348, 910-914. 10.1126 / science. aab1601.
[0186] 6. Buenrostro, J.D., Wu, B., Litzenburger, U.M., Ruff, D., Gonzales, M.L., Snyder, M.P., Chang, H.Y., and Greenleaf, W.J. (2015) . Single-cell chromatin accessibility reveals principles of regulatory variation. Nature 523, 486-490. 10.1038 / nature14590.
[0187] 7. Buenrostro, J.D., Giresi, P.G., Zaba, L.C., Chang, H.Y., and Greenleaf, W.J. (2013) . Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods 10, 1213-1218. 10.1038 / nmeth. 2688.
[0188] 8. Grandi, F.C., Modi, H., Kampman, L., and Corces, M.R. (2022) . Chromatin accessibility profiling by ATAC-seq. Nat Protoc 17, 1518-1552. 10.1038 / s41596-022-00692-9.
[0189] 9. Chen, X., Miragaia, R.J., Natarajan, K.N., and Teichmann, S.A. (2018) . A rapid and robust method for single cell chromatin accessibility profiling. Nat Commun 9, 5345. 10.1038 / s41467-018-07771-0.
[0190] 10. Mezger, A., Klemm, S., Mann, I., Brower, K., Mir, A., Bostick, M., Farmer, A., Fordyce, P., Linnarsson, S., and Greenleaf, W. (2018) . High-throughput chromatin accessibility profiling at single-cell resolution. Nature Communications 9. ARTN 364710.1038 / s41467-018-05887-x.
[0191] 11. Xie, Y., Zhu, C., Wang, Z., Tastemel, M., Chang, L., Li, Y.E., and Ren, B. (2023) . Droplet-based single-cell joint profiling of histone modifications and transcriptomes. Nat Struct Mol Biol 30, 1428-1433. 10.1038 / s41594-023-01060-1.
[0192] 12. Stuart, T., Hao, S., Zhang, B.J., Mekerishvili, L., Landau, D.A., Maniatis, S., Satija, R., and Raimondi, I. (2023) . Nanobody-tethered transposition enables multifactorial chromatin profiling at single-cell resolution. Nature Biotechnology 41, 806-+. 10.1038 / s41587-022-01588-5.
[0193] 13. Liu, Z., Chen, Y., Xia, Q., Liu, M., Xu, H., Chi, Y., Deng, Y., and Xing, D. (2023) . Linking genome structures to functions by simultaneous single-cell Hi-C and RNA-seq. Science 380, 1070-1076. 10.1126 / science. adg3797.
[0194] 14. Xiong, H., Luo, Y., Wang, Q., Yu, X., and He, A. (2021) . Single-cell joint detection of chromatin occupancy and transcriptome enables higher-dimensional epigenomic reconstructions. Nat Methods 18, 652-660. 10.1038 / s41592-021-01129-z.
[0195] 15. Mimitou, E.P., Lareau, C.A., Chen, K.Y., Zorzetto-Fernandes, A.L., Hao, Y., Takeshima, Y., Luo, W., Huang, T.S., Yeung, B.Z., Papalexi, E., et al. (2021) . Scalable, multimodal profiling of chromatin accessibility, gene expression and protein levels in single cells. Nat Biotechnol 39, 1246-1258. 10.1038 / s41587-021-00927-2.
[0196] 16. Ma, S., Zhang, B., LaFave, L.M., Earl, A.S., Chiang, Z., Hu, Y., Ding, J., Brack, A., Kartha, V.K., Tay, T., et al. (2020) . Chromatin Potential Identified by Shared Single-Cell Profiling of RNA and Chromatin. Cell 183, 1103-1116 e1120. 10.1016 / j. cell. 2020.09.056.
[0197] 17. De Rop, F.V., Hulselmans, G., Flerin, C., Soler-Vila, P., Rafels, A., Christiaens, V., Gonzalez-Blas, C.B., Marchese, D., Caratu, G., Poovathingal, S., et al. (2023) . Systematic benchmarking of single-cell ATAC-sequencing protocols. Nat Biotechnol. 10.1038 / s41587-023-01881-x.
[0198] 18. Lareau, C.A., Duarte, F.M., Chew, J.G., Kartha, V.K., Burkett, Z.D., Kohlway, A.S., Pokholok, D., Aryee, M.J., Steemers, F.J., Lebofsky, R., and Buenrostro, J.D. (2019) . Droplet-based combinatorial indexing for massive-scale single-cell chromatin accessibility. Nat Biotechnol 37, 916-924. 10.1038 / s41587-019-0147-6.
[0199] 19. Zhang, K., Hocker, J.D., Miller, M., Hou, X., Chiou, J., Poirion, O.B., Qiu, Y., Li, Y.E., Gaulton, K.J., Wang, A., et al. (2021) . A single-cell atlas of chromatin accessibility in the human genome. Cell 184, 5985-6001 e5919. 10.1016 / j. cell. 2021.10.024.
[0200] 20. Domcke, S., Hill, A.J., Daza, R.M., Cao, J., O'Day, D.R., Pliner, H.A., Aldinger, K.A., Pokholok, D., Zhang, F., Milbank, J.H., et al. (2020) . A human cell atlas of fetal chromatin accessibility. Science 370. 10.1126 / science. aba7612.
[0201] 21. Picelli, S., Bjorklund, A.K., Reinius, B., Sagasser, S., Winberg, G., and Sandberg, R. (2014) . Tn5 transposase and tagmentation procedures for massively scaled sequencing projects. Genome Res 24, 2033-2040. 10.1101 / gr. 177881.114.
[0202] 22. Corces, M.R., Trevino, A.E., Hamilton, E.G., Greenside, P.G., Sinnott-Armstrong, N.A., Vesuna, S., Satpathy, A.T., Rubin, A.J., Montine, K.S., Wu, B., et al. (2017) . An improved ATAC-seq protocol reduces background and enables interrogation of frozen tissues. Nat Methods 14, 959-962. 10.1038 / nmeth. 4396.
[0203] 23. Zu, S., Li, Y.E., Wang, K., Armand, E.J., Mamde, S., Amaral, M.L., Wang, Y., Chu, A., Xie, Y., Miller, M., et al. (2023) . Single-cell analysis of chromatin accessibility in the adult mouse brain. Nature 624, 378-389. 10.1038 / s41586-023-06824-9.
[0204] 24. Hayashi, K., and Surani, M.A. (2009) . Self-renewing epiblast stem cells exhibit continual delineation of germ cells with epigenetic reprogramming in vitro. Development 136, 3549-3556. 10.1242 / dev. 037747.
[0205] 25. Kurimoto, K., Yabuta, Y., Hayashi, K., Ohta, H., Kiyonari, H., Mitani, T., Moritoki, Y., Kohri, K., Kimura, H., Yamamoto, T., et al. (2015) . Quantitative Dynamics of Chromatin Remodeling during Germ Cell Specification from Mouse Embryonic Stem Cells. Cell Stem Cell 16, 517-532. 10.1016 / j. stem. 2015.03.002.
[0206] 26. Ibarra-Soria, X., Jawaid, W., Pijuan-Sala, B., Ladopoulos, V., Scialdone, A., D.J., Tyser, R.C.V., Calero-Nieto, F.J., Mulas, C., Nichols, J., et al. (2018) . Defining murine organogenesis at single-cell resolution reveals a role for the leukotriene pathway in regulating blood progenitor formation. Nature Cell Biology 20, 127-+. 10.1038 / s41556-017-0013-z.
[0207] 27. Ai, S., Xiong, H., Li, C.C., Luo, Y., Shi, Q., Liu, Y., Yu, X., Li, C., and He, A. (2019) . Profiling chromatin states using single-cell itChIP-seq. Nat Cell Biol 21, 1164-1172. 10.1038 / s41556-019-0383-5.
[0208] 28. Perino, M., and Veenstra, G.J. (2016) . Chromatin Control of Developmental Dynamics and Plasticity. Dev Cell 38, 610-620. 10.1016 / j. devcel. 2016.08.004.
[0209] 29. Kutateladze, T.G., Gozani, O., Bienz, M., and Ostankovitch, M. (2017) . Histone modifications for chromatin dynamics and cellular plasticity. J Mol Biol 429, 1921-1923. 10.1016 / j. jmb. 2017.06.001.
[0210] 30. Yadav, T., Quivy, J.P., and Almouzni, G. (2018) . Chromatin plasticity: A versatile landscape that underlies cell fate and identity. Science 361, 1332-1336. 10.1126 / science. aat8950.
[0211] 31. Hao, Y., Hao, S., Andersen-Nissen, E., Mauck, W.M., 3rd, Zheng, S., Butler, A., Lee, M.J., Wilk, A.J., Darby, C., Zager, M., et al. (2021) . Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529. 10.1016 / j. cell. 2021.04.048.
[0212] 32. Granja, J.M., Klemm, S., McGinnis, L.M., Kathiria, A.S., Mezger, A., Corces, M.R., Parks, B., Gars, E., Liedtke, M., Zheng, G.X.Y., et al. (2019) . Single-cell multiomic analysis identifies regulatory programs in mixed-phenotype acute leukemia. Nat Biotechnol 37, 1458-1465. 10.1038 / s41587-019-0332-7.
[0213] 33. Garrett-Sinha, L.A., Hou, P., Wang, D., Grabiner, B., Araujo, E., Rao, S., Yun, T.J., Clark, E.A., Simon, M.C., and Clark, M.R. (2005) . Spi-1 and Spi-B control the expression of the Grap2 gene in B cells. Gene 353, 134-146. 10.1016 / j. gene. 2005.04.009.
[0214] 34. Thoms, J.A.I., Truong, P., Subramanian, S., Knezevic, K., Harvey, G., Huang, Y., Seneviratne, J.A., Carter, D.R., Joshi, S., Skhinas, J., et al. (2021) . Disruption of a GATA2-TAL1-ERG regulatory circuit promotes erythroid transition in healthy and leukemic stem cells. Blood 138, 1441-1455. 10.1182 / blood. 2020009707.
[0215] 35. Georgolopoulos, G., Psatha, N., Iwata, M., Nishida, A., Som, T., Yiangou, M., Stamatoyannopoulos, J.A., and Vierstra, J. (2021) . Discrete regulatory modules instruct hematopoietic lineage commitment and differentiation. Nat Commun 12, 6790. 10.1038 / s41467-021-27159-x.
[0216] 36. Goode, D.K., Obier, N., Vijayabaskar, M.S., Lie, A.L.M., Lilly, A.J., Hannah, R., Lichtinger, M., Batta, K., Florkowska, M., Patel, R., et al. (2016) . Dynamic Gene Regulatory Networks Drive Hematopoietic Specification and Differentiation. Dev Cell 36, 572-587. 10.1016 / j. devcel. 2016.01.024.
[0217] 37. Fan, X., Yang, C., Li, W., Bai, X., Zhou, X., Xie, H., Wen, L., and Tang, F. (2021) . SMOOTH-seq: single-cell genome sequencing of human cells on a third-generation sequencing platform. Genome Biol 22, 195. 10.1186 / s13059-021-02406-y.
[0218] 38. Kaya-Okur, H.S., Wu, S.J., Codomo, C.A., Pledger, E.S., Bryson, T.D., Henikoff, J.G., Ahmad, K., and Henikoff, S. (2019) . CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Commun 10, 1930. 10.1038 / s41467-019-09982-5.
[0219] 39. Ramani, V., Deng, X., Qiu, R., Gunderson, K.L., Steemers, F.J., Disteche, C.M., Noble, W.S., Duan, Z., and Shendure, J. (2017) . Massively multiplex single-cell Hi-C. Nat Methods 14, 263-266. 10.1038 / nmeth. 4155.
[0220] 40. Clark, S.J., Argelaguet, R., Kapourani, C.A., Stubbs, T.M., Lee, H.J., Alda-Catalinas, C., Krueger, F., Sanguinetti, G., Kelsey, G., Marioni, J.C., et al. (2018) . scNMT-seq enables joint profiling of chromatin accessibility DNA methylation and transcription in single cells. Nat Commun 9, 781. 10.1038 / s41467-018-03149-4.
[0221] 41. Zachariadis, V., Cheng, H., Andrews, N., and Enge, M. (2020) . A Highly Scalable Method for Joint Whole-Genome Sequencing and Gene-Expression Profiling of Single Cells. Mol Cell 80, 541-553 e545. 10.1016 / j. molcel. 2020.09.025.
[0222] 42. Cao, J., Cusanovich, D.A., Ramani, V., Aghamirzaie, D., Pliner, H.A., Hill, A.J., Daza, R.M., McFaline-Figueroa, J.L., Packer, J.S., Christiansen, L., et al. (2018) . Joint profiling of chromatin accessibility and gene expression in thousands of single cells. Science 361, 1380-1385. 10.1126 / science. aau0730
[0223] It is understood that the disclosed method and compositions are not limited to the particular methodology, protocols, and reagents described as these can vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present invention which will be limited only by the appended claims.
[0224] Disclosed are materials, compositions, and components that can be used for, can be used in conjunction with, can be used in preparation for, or are products of the disclosed method and compositions. These and other materials are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these materials are disclosed that while specific reference of each various individual and collective combinations and permutation of these compounds may not be explicitly disclosed, each is specifically contemplated and described herein. For example, if a tagmentation adapter is disclosed and discussed and a number of modifications that can be made to a number of molecules including the tagmentation adapter are discussed, each and every combination and permutation of tagmentation adapter and the modifications that are possible are specifically contemplated unless specifically indicated to the contrary. Thus, if a class of molecules A, B, and C are disclosed as well as a class of molecules D, E, and F and an example of a combination molecule, A-D is disclosed, then even if each is not individually recited, each is individually and collectively contemplated. Thus, is this example, each of the combinations A-E, A-F, B-D, B-E, B-F, C-D, C-E, and C-F are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. Likewise, any subset or combination of these is also specifically contemplated and disclosed. Thus, for example, the sub-group of A-E, B-F, and C-E are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. Further, each of the materials, compositions, components, etc. contemplated and disclosed as above can also be specifically and independently included or excluded from any group, subgroup, list, set, etc. of such materials. These concepts apply to all aspects of this application including, but not limited to, steps in methods of making and using the disclosed compositions. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific embodiment or combination of embodiments of the disclosed methods, and that each such combination is specifically contemplated and should be considered disclosed.
[0225] It must be noted that as used herein and in the appended claims, the singular forms “a, ” “an, ” and “the” include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to “a tagmentation adapter” includes a plurality of such tagmentation adapters, reference to “the tagmentation adapter” is a reference to one or more tagmentation adapters and equivalents thereof known to those skilled in the art, and so forth.
[0226] “Optional” or “optionally” means that the subsequently described event, circumstance, or material may or may not occur or be present, and that the description includes instances where the event, circumstance, or material occurs or is present and instances where it does not occur or is not present.
[0227] Unless the context clearly indicates otherwise, use of the word “can” indicates an option or capability of the object or condition referred to. Generally, use of “can” in this way is meant to positively state the option or capability while also leaving open that the option or capability could be absent in other forms or embodiments of the object or condition referred to. Unless the context clearly indicates otherwise, use of the word “may” indicates an option or capability of the object or condition referred to. Generally, use of “may” in this way is meant to positively state the option or capability while also leaving open that the option or capability could be absent in other forms or embodiments of the object or condition referred to. Unless the context clearly indicates otherwise, use of “may” herein does not refer to an unknown or doubtful feature of an object or condition.
[0228] Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, also specifically contemplated and considered disclosed is the range from the one particular value and / or to the other particular value unless the context specifically indicates otherwise. Similarly, when values are expressed as approximations, by use of the antecedent “about, ” it will be understood that the particular value forms another, specifically contemplated embodiment that should be considered disclosed unless the context specifically indicates otherwise. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint unless the context specifically indicates otherwise. It should be understood that all of the individual values and sub-ranges of values contained within an explicitly disclosed range are also specifically contemplated and should be considered disclosed unless the context specifically indicates otherwise. Finally, it should be understood that all ranges refer both to the recited range as a range and as a collection of individual numbers from and including the first endpoint to and including the second endpoint. In the latter case, it should be understood that any of the individual numbers can be selected as one form of the quantity, value, or feature to which the range refers. In this way, a range describes a set of numbers or values from and including the first endpoint to and including the second endpoint from which a single member of the set (i.e. a single number) can be selected as the quantity, value, or feature to which the range refers. The foregoing applies regardless of whether in particular cases some or all of these embodiments are explicitly disclosed.
[0229] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of skill in the art to which the disclosed method and compositions belong. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present method and compositions, the particularly useful methods, devices, and materials are as described. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such disclosure by virtue of prior invention. No admission is made that any reference constitutes prior art. The discussion of references states what their authors assert, and applicants reserve the right to challenge the accuracy and pertinency of the cited documents. It will be clearly understood that, although a number of publications are referred to herein, such reference does not constitute an admission that any of these documents forms part of the common general knowledge in the art.
[0230] Although the description of materials, compositions, components, steps, techniques, etc. can include numerous options and alternatives, this should not be construed as, and is not an admission that, such options and alternatives are equivalent to each other or, in particular, are obvious alternatives.
[0231] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the method and compositions described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
A method of single-cell (sc) assay for transposase accessible chromatin (ATAC) (scATAC) , the method comprising:(a) producing N distinct sets of tagmentated cells or nuclei, wherein the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei is distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei, whereby the N distinct sets of tagmentated cells or nuclei are produced, wherein each of the N sets of tagmentated cells or nuclei is distinguished by a pair of adapter barcodes that differs in each of the N sets of tagmentated cells or nuclei;(b) sorting a single cell or nucleus of each of the N distinct sets of tagmentated cells or nuclei into K reaction chambers (e.g., wells) such that each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei, wherein the K reaction chambers are comprised in L reaction groups (e.g., microtiter plates) , wherein each reaction group comprises M of the K reaction chambers (e.g., M = K / L) ;(c) performing separate PCR amplifications of the contents of each of the K reaction chambers using a set of P distinct pairs of barcoded first PCR primers, wherein, for each of the M reaction chambers in a given reaction group, a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is used; and(d) performing PCR amplification of the contents of the M reaction chambers in a given reaction group using a pair of barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers, wherein, for each of the L reaction groups, a different pair of barcoded second PCR primers from the set of Q distinct pairs of barcoded second PCR primers is used.The method of claim 1 further comprising:(e) performing sequencing of the amplified products of step (d) .The method of claim 1 or 2, wherein each of the N sets of indexed transposomes is independently formed by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes, wherein each of the N sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters, whereby N separate sets of indexed transposomes are formed, wherein i *j = N, each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.The method of any one of claims 1-3, wherein each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers,wherein each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers, andwherein each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.The method of claim 3 or 4, wherein all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence and wherein all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.The method of claim 5, wherein all of the forward barcoded first PCR primers comprise the same forward first primer sequence, wherein the forward first primer sequence matches the forward first primer matching sequence, wherein all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence,wherein all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence, wherein the reverse first primer sequence matches the reverse first primer matching sequence, wherein all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.The method of claim 6, wherein all of the forward barcoded second PCR primers comprise the same forward second primer sequence, wherein the forward second primer sequence matches the forward second primer matching sequence,wherein all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence, wherein the reverse second primer sequence matches the reverse second primer matching sequence.The method of any one of claims 5-7, wherein the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination the forward barcoded first PCR primers, the forward barcoded second PCR primers comprise next-generation sequencing sequences.A set of N pairs of left and right barcoded tagmentation adapters,wherein the barcoded tagmentation adapters each comprise a transfer strand and a non-transfer strand hybridized together,wherein the hybridized transfer and non-transfer strands form a double stranded transposase recognition sequence and a single stranded barcoded adapter sequence, wherein the barcoded adapter sequence extends 5’ from the transfer strand,wherein the barcoded adapter sequence comprises an adapter barcode and an adapter sequence,wherein the adapter barcodes of each of the left and right barcoded tagmentation adapters in the set of pairs of left and right barcoded tagmentation adapters are distinct from each other,wherein all of the adapter sequences of all of the left barcoded tagmentation adapters comprise the same forward first primer matching sequence, and wherein all of the adapter sequences of all of the right barcoded tagmentation adapters comprise the same reverse first primer matching sequence.The set of N pairs of barcoded tagmentation adapters of claim 9 comprised in a set of N sets of indexed transposomes, wherein, for each of the N sets of indexed transposomes, the transposomes in the set comprises a tagmentation transposase and the same pair of left barcoded tagmentation adapter and right barcoded tagmentation adapter, wherein each of the N sets of indexed transposomes comprises a distinct pair of pair of left barcoded tagmentation adapters and right barcoded tagmentation adapters, wherein each of the N sets of indexed transposomes is distinguished by a pair of adapter barcodes that differs in each of the N sets of indexed transposomes.The set of N sets of indexed transposomes of claim 10, wherein each of the N sets of indexed transposomes is independently formed by bringing into contact one of a set of i distinct left barcoded tagmentation adapters, one of a set of j distinct right barcoded tagmentation adapters, and a tagmentation transposase under conditions that promote formation of transposomes, wherein each of the sets of indexed transposomes is formed using a distinct combination of one of the i distinct left barcoded tagmentation adapters and one of the j distinct right barcoded tagmentation adapters, whereby separate sets of indexed transposomes are formed, wherein i *j = N, wherein each of the i and j sets of barcoded tagmentation adapters is distinguished by an adapter barcode that differs in each of the i and j sets of barcoded tagmentation adapters.The set of N sets of indexed transposomes of claim 10 or 11 comprised in a set of N distinct sets of tagmentated cells or nuclei, wherein the cells or nuclei of each of the N distinct sets of tagmentated cells or nuclei are distinguished by the pair of adapter barcodes that differs in each of the N sets of indexed transposomes.The set of N distinct sets of tagmentated cells or nuclei of claim 12 disposed in a set of K reaction chambers, wherein each of the K reaction chambers contains a single cell or nucleus from each of the N sets of tagmentated cells or nuclei, wherein the K reaction chambers are comprised in L reaction groups, wherein each reaction group comprises M of the K reaction chambers.The set of K reaction chambers of claim 13 further comprising a set of P distinct pairs of forward and reverse barcoded first PCR primers, wherein a different pair of barcoded first PCR primers from the set of P distinct pairs of barcoded first PCR primers is disposed in each of the M reaction chambers in a given reaction group of the L reaction groups,wherein all of the forward barcoded first PCR primers comprise the same forward first primer sequence, wherein the forward first primer sequence matches the forward first primer matching sequence, wherein all of the forward barcoded first PCR primers comprise the same forward second primer matching sequence,wherein all of the reverse barcoded first PCR primers comprise the same reverse first primer sequence, wherein the reverse first primer sequence matches the reverse first primer matching sequence, wherein all of the reverse barcoded first PCR primers comprise the same reverse second primer matching sequence.The L reaction groups of claim 14 further comprising a different pair of forward and reverse barcoded second PCR primers from a set of Q distinct pairs of barcoded second PCR primers disposed in the contents of the M reaction chambers in each of the L reaction groups,wherein all of the forward barcoded second PCR primers comprise the same forward second primer sequence, wherein the forward second primer sequence matches the forward second primer matching sequence,wherein all of the reverse barcoded second PCR primers comprise the same reverse second primer sequence, wherein the reverse second primer sequence matches the reverse second primer matching sequence.The PCR primers of claim 15, wherein each of the adapter barcodes are distinct from each other and from each of the barcodes in the barcoded first PCR primers and each of the barcodes in the barcoded second PCR primers,wherein each of the barcodes in the barcoded first PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded second PCR primers, andwherein each of the barcodes in the barcoded second PCR primers are distinct from each other and from each of the adapter barcodes and each of the barcodes in the barcoded first PCR primers.The PCR primers of claim 15 or 16, wherein the forward barcoded first PCR primers, the forward barcoded second PCR primers, or a combination the forward barcoded first PCR primers and the forward barcoded second PCR primers comprise next-generation sequencing sequences.
Citation Information
Patent Citations
High-sensitivity method for detecting chromatin state and genome information
CN117802209A
Single cell whole genome libraries and combinatorial indexing methods of making thereof
US20180023119A1
Methods and compositions for molecular interaction mapping using transposase
WO2023081863A1