Method for multiplexed multiome and use thereof

The method addresses high costs and batch effects in single-cell multi-omics by using sample-specific barcoded transposons for multiplexed analysis, achieving efficient and cost-effective simultaneous profiling of gene expression and chromatin accessibility.

WO2026112003A1PCT designated stage Publication Date: 2026-05-28CHILDRENS NAT MEDICAL CENT
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHILDRENS NAT MEDICAL CENT
Filing Date
2025-11-17
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing single-cell multi-omics methods are limited by high costs and technical batch effects due to separate processing of biological samples, which introduces variability and makes it difficult to distinguish biological differences from technical artifacts.

Method used

A method for multiplexed analysis involving sample-specific barcoded transposons that tag genomic DNA with unique barcodes, allowing pooling and concurrent processing of multiple samples in a single reaction, thereby reducing costs and eliminating batch effects.

Benefits of technology

The method achieves significant cost reduction and eliminates technical batch effects, enabling simultaneous profiling of gene expression and chromatin accessibility from multiple samples with high demultiplexing efficiency and doublet detection, maintaining data quality and compatibility with existing commercial equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055756_28052026_PF_FP_ABST
    Figure US2025055756_28052026_PF_FP_ABST
Patent Text Reader

Abstract

A method for multiplexed analysis of single nuclei is provided. The method pools nuclei from multiple independent biological samples before performing droplet-based capture for simultaneous single-nucleus RNA sequencing (snRNA-seq) and ATAC sequencing (snATAC-seq). The method involves using custom Tn5 transposomes loaded with sample-specific barcoded oligonucleotides to tag open chromatin within the nuclei of each sample. After the barcoding transposition reaction, the nuclei from all samples are pooled and processed in a single reaction for nuclei capture, sequencing library preparation, and sequencing. Subsequent data pre-processing allows for the computational demultiplexing of reads based on the sample-specific barcode.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR MULTIPLEXED MULTIOME AND USE THEREOFCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 722,083, filed November 19, 2024, entitled "METHOD FOR MULTIPLEXED MULTIOME AND USE THEREOF. " The entire contents of this provisional application are hereby incorporated by reference in their entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under NS 116418 awarded by the National Institutes of Health. The government has certain rights in the invention.REFERENCE TO SEQUENCE LISTING

[0003] In accordance with 37 CFR §1.52(e)(5) and with 37 CFR §1.831, the specification makes reference to a Sequence Listing submitted electronically as a .xml file named 19819. OOOIWOUL xml copy, created and filed herewith is 14000 bytes in size. The entire contents of the Sequence Listing are hereby incorporated by reference.BACKGROUND

[0004] The present disclosure relates to methods for analyzing biological samples, and more specifically to multiplexing methods for simultaneous single-cell or single-nucleus genomic and transcriptomic analysis. In recent years, the field of single-cell multi-omics, which allows for the concurrent measurement of multiple modalities such as the transcriptome and chromatin accessibility7from the same single cell, has seen rapid advancements.

[0005] Commercially available systems enable the simultaneous profiling of the transcriptome (via single nuclei RNA sequencing or snRNA-seq) and chromatin accessibility7(via Assay for Transposase-Accessible Chromatin using sequencing or snATAC-seq). This is often referred to as "multiome" sequencing technology. One prominent example is the Chromium Next GEM Single Cell Multiome ATAC + Gene Expression platform from 10X Genomics. In this system, single nuclei are encapsulated in microfluidic droplets with barcoded gel beads. Within each droplet, messenger RNA(mRNA) is captured for transcriptomic analysis, and a transposase enzy me inserts sequencing adapters into open chromatin regions for epigenomic analysis. All molecules from a single nucleus share a common cell barcode, allowing data to be linked back to the original nucleus after sequencing.

[0006] Despite the power of these technologies, their adoption is limited by significant drawbacks. A primary limitation is the prohibitive cost. For example, performing a single sn-Multiome reaction using existing commercial kits can cost approximately $3,000 per sample. For biological experiments that require multiple samples from different conditions or time points, often with several replications, the total cost can quickly become prohibitive for many research laboratories. A 48-sample experiment, for instance, could incur reagent costs alone of approximately $140,000. Furthermore, the potential output of a single reaction is often under-utilized when only a small number of nuclei are available as input, exacerbating the cost-inefficiency. This problem is magnified with the advent of higher-throughput capture platforms capable of processing millions of nuclei per run.

[0007] Another significant challenge with processing samples individually is the introduction of technical batch effects. When each biological sample is processed in a separate reaction, minor variations in reagents, handling, and instrument performance can introduce non-biological variability into the data. This can confound downstream analysis and make it difficult to distinguish true biological differences between samples from technical artifacts, particularly when comparing samples across different experimental conditions or replications.

[0008] Therefore, there exists a need in the art for a method that allows for the pooling and concurrent processing of multiple biological samples in a single reaction. Such a sample multiplexing strategy would significantly reduce the per-sample cost, minimize technical batch effects, and improve the overall efficiency and scalability of single-cell multi-ome sequencing experiments. The present disclosure addresses these and other deficiencies in the prior art.SUMMARY

[0009] Provided herein is a method for multiplexed analysis of biological samples that transforms single-cell multi-omics by enabling cost reduction and elimination of technical batch effects. The method involves providing a plurality of biological samples, eachcontaining nuclei from diverse sources including different species, individuals, tissues, time points, or experimental conditions.

[0010] In one embodiment and for each biological sample, the method performs a separate transposition reaction using a transposome complex comprising a sample-specific barcoded transposon that functions to tag genomic DNA within the nuclei with a unique barcode corresponding to the sample of origin. This barcoding enables subsequent computational identification and assignment of sequencing reads to their originating samples.

[0011] Following transposition, the method involves pooling the barcoded nuclei from the plurality of biological samples to create a pooled nuclei suspension and subjecting the pooled nuclei suspension to a single droplet-based capture reaction to co-encapsulate individual nuclei with barcoded gel beads. This pooling approach allows multiple samples to be processed together, dramatically reducing per-sample costs while eliminating technical batch effects.

[0012] In some embodiments, the analysis comprises simultaneous single-nucleus RNA sequencing (snRNA-seq) and single-nucleus Assay for Transposase- Accessible Chromatin using sequencing (snATAC-seq). This dual-modality analysis enables researchers to measure both gene expression and chromatin accessibility from the same nucleus while multiplexing multiple samples in a single reaction, providing comprehensive molecular profiling capabilities.

[0013] In some embodiments, the transposome complex comprises aTn5 transposase enzyme, which provides the catalytic activity for simultaneous DNA cleavage and adapter ligation during the transposition reaction. The Tn5 transposase recognizes specific mosaic end sequences and efficiently inserts the barcoded oligonucleotides into accessible chromatin regions.

[0014] In some embodiments, the sample-specific barcoded transposon comprising a doublestranded oligonucleotide with precise length specifications. The first strand of the oligonucleotide can have an overall length of less than about 60 base pairs to attenuate or prevent concatenation, which reduces or eliminates artifacts that interfere with sequencing library preparation.

[0015] In some embodiments, the first strand of the oligonucleotide has an overall length of about 42 to about 45 base pairs, which reduces or eliminates concatenation while maintaining full transposition functionality. In some embodiments, oligonucleotides of about 60 base pairs or longer form concatenated products with approximately 60 base pairperiodicity, creating ladder patterns on TapeStation analysis that cannot be readily removed bv standard size selection methods.

[0016] In one embodiment, a sequence structure comprises 5'- GTGCTCTTCCGATCT[X]nAGATGTGTATAAGAGACAG-3', wherein [X]n represents a unique barcode corresponding to the sample of origin (SEQ ID NO:1). This structure can provide balanced functionality for transposition, PCR amplification, and computational demultiplexing.

[0017] In some embodiments, the method encompasses generating a gene expression sequencing library and a chromatin accessibility sequencing library from molecules captured on the barcoded gel beads followed by sequencing the gene expression and chromatin accessibility libraries to produce sequencing reads.

[0018] In some embodiments, computationally demultiplexing the sequencing reads from the chromatin accessibility7library based on the unique barcode to assign each read to its sample of origin. This process can achieve greater than about 90% efficiency, with over about 90% of raw ATAC reads containing recognizable MuMu barcodes.

[0019] In some embodiments, the method further includes assigning the gene expression sequencing reads to their sample of origin by associating a cell barcode present on the gene expression reads with a cell barcode present on the demultiplexed chromatin accessibility reads. This association can leverage the shared cell barcodes from gel beads to link multimodal data from the same nucleus.

[0020] In some embodiments, identity ing and computationally removing data corresponding to doublet events, wherein a doublet event is identified by the association of a single cell barcode with more than one unique barcode. For example, validation demonstrates doublet detection rates of approximately 5% between samples from the same species and up to 20% between samples from different species, consistent with expected droplet capture statistics.

[0021] In some embodiments, there is provided one or more, preferably several, oligonucleotides for sample multiplexing in transposon-based assays comprising a central sample-specific barcode sequence and one or more mosaic end (ME) sequences recognized by a transposase. For example, an overall length of less than about 60 base pairs, such as less than 60 base pairs and such as from about 42 to about 45 base pairs.

[0022] In some embodiments, the transposase is Tn5 transposase, which recognizes the mosaic end sequences and efficiently catalyzes transposition.

[0023] In some embodiments, the specific sequence structure 5'- GTGCTCTTCCGATCT[X]nAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 1) is provided herein.

[0024] In some embodiments, the transposome complex comprises atransposase enzyme and a double-stranded nucleic acid construct bound to the transposase enzyme, where the first strand comprises a sample-specific barcode sequence and has an overall length of less than about 60 base pairs. In some embodiments, the transposase enzyme is preferably Tn5 transposase, providing the catalytic function for efficient transposition.

[0025] In some embodiments, kits for performing multiplexed multi-omic analysis of single nuclei containing a plurality of distinct barcoded oligonucleotides, wherein each oligonucleotide is as described herein. The kits can, in some embodiments include a transposase enzyme for assembly with the plurality of barcoded oligonucleotides to form a plurality of active transposomes, for instance a naked Tn5 transposase. Additional components, in certain embodiments, comprise nuclei lysis buffer, nuclei resuspension buffer, and transposition buffer as well as primers for sequencing library preamplification or final library amplification.

[0026] In some embodiments, provided herein are complete systems for performing multiplexed analysis including plurality of distinct transposome complexes, microfluidic devices configured to perform droplet-based capture reactions, and computer-readable media comprising instructions for demultiplexing sequencing data.

[0027] In some embodiments, the computational component performs identifying the unique sample-specific barcode sequence within chromatin accessibility sequencing reads and assigning the chromatin accessibility reads and associated gene expression reads to a sample of origin based on the identified barcode. Additional embodied features include identify ing and removing data corresponding to doublet events by detecting association of single cell barcodes with multiple unique sample-specific barcodes.

[0028] In some embodiments, the method can achieve quantitative performance benchmarks including demultiplexing efficiency of at least about 90%, with greater than about 90% of chromatin accessibility reads containing recognizable sample-specific barcodes. Speciesmixing validation can demonstrate nuclei exhibit at least about 70% alignment to their expected species genome and less than about 30% alignment to non-target species genomes.

[0029] In some embodiments, the method produces sequencing libraries with size distribution profiles substantially identical to libraries generated by standard non-multiplexed protocols, ensuring compatibility with existing commercial equipment and analysis software. The approach eliminates the need for antibody-based cell hashing or complex combinatorial indexing, providing a streamlined solution integrated directly into the ATAC-seq workflow

[0030] A variety of additional inventive aspects will be set forth in the description that follows. The inventive aspects can relate to individual features and to combinations of features. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the broad inventive concepts of the disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings, which are incorporated in and constitute a part of the description, illustrate several aspects of the present disclosure. A brief description of the drawings is as follows:

[0032] FIG. 1 shows Tn5 transposome concatenate during reaction when transposon is 60 base pairs (bp) in size. FIGS. A and B show TapeStation analysis traces showing unexpected products (rectangle with red dashed line) with periodicity of roughly 60bp in the presence (A) or absence (B) of genomic DNA (rectangle with green dashed line), which the transposon is at 60bp. (C) No concatenation is observed using the MuMu barcode oligomers, which are 45 bp in size.

[0033] FIGs. 2 A and B show ty pical TapeStation traces of finished ATAC libraries produced using (A) the current protocol or (B) the original 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression protocol.

[0034] FIGs. 3A-K demonstrated sample multiplexing and downstream data analyses. (A) Sankey diagram showing the breakdown of all ATAC reads acquired through sequencing at each data preprocessing step. The percentage out of the total ATAC reads for each category7is labeled in parentheses. ME, mosaic end; M, million reads. (B) Histograms showing the distribution of percentages of RNA or ATAC reads mapped to mouse or pig genome from each nucleus. Sample 1 and 3 are collected from a mouse neocortex at embryonic day (E) 16, while 2 and 4 are from a pig neocortex at gestation week 10. Dotted vertical lines indicate 50% mapping rate (“cutoff”). Additionally, an E15.5 mouse neocortex sample (“10X Multiome”) processed using the original 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit was included for comparison. (C) Venn diagram showing number of cell barcodes presented in eachsample. Cell barcodes presented in more than one sample are likely “doublets” and are removed from downstream analyses. (D) Violin plots showing distributions of RNA read counts, detected genes, and percentage of reads mapped to mitochondrial genome in nuclei passing quality control. MT, mitochondria. (E and F) Uniform manifold approximation and projection constructed with gene expression data showing single nuclei profiles from the mouse (C) or the pig (D) samples. Colors represent cell clusters identified through the Seurat analysis pipeline. The cell clusters are ordered by size (number of nuclei) from largest to smallest. (G and H) Violin plots showing the most specifically expressed genes associated with each cell cluster of the mouse (E) or the pig (F) samples. Clusters are arranged by size from largest to smallest. (I) Violin plots showing distributions of ATAC read counts, ATAC features, transcription start site enrichment, and nucleosome signal. TSS, transcription start site. (J and K) Chromatin accessibility profiles around gene PTPRZ1 in each cell cluster of the mouse (J) or the pig (K) samples. In dark grey. PTPRZ1 expression levels are shown on the right as violin plots. The accumulative chromatin accessibility profile from all nuclei is shown below, followed by a diagram of PTPRZ1 gene body. All identified ATAC peaks (“Peaks”) within the region are shown as vertical lines. A peak in the upstream region (“Links”) is identified in both mouse and pig genome that is correlated with gene expression of PTPRZ1. Norm. exp. level, normalized expression level.

[0035] FIG. 4 is a graphical representation of an embodiment of the present disclosure showing each step from nuclei extraction to data analysis, including exemplary times for the steps.DETAILED DESCRIPTION

[0036] "Transposome" means a ribonucleoprotein complex comprising (i) a transposase enzyme, preferably Tn5 transposase, and (ii) a double-stranded oligonucleotide substrate bound to said transposase enzy me, wherein the oligonucleotide substrate comprises mosaic end sequences recognized by the transposase and is capable of being inserted into target DNA through transposition. In preferred embodiments, the transposome complex is assembled by incubating the transposase enzyme with annealed oligonucleotides at defined ratios to form an active complex capable of simultaneous DNA cleavage and adapter ligation.

[0037] "Sample-specific barcode" or "unique barcode" means a defined nucleotide sequence, for example, 6 to 10 base pairs in length, that serves as a unique identifier for a particular biological sample within a multiplexed reaction. Each sample-specific barcode is distinct from all other barcodes used in the same multiplexed experiment and is incorporated into genomic DNA fragments during the transposition reaction to enable subsequent computational assignment of sequencing reads to their sample of origin.

[0038] "Barcoded transposon" means a double-stranded oligonucleotide construct designed for assembly with a transposase enzyme, comprising (i) a central sample-specific barcode sequence, (ii) flanking mosaic end sequences recognized by the transposase, and (iii) having a total length of less than about 60 base pairs, for example, from 42 to 45 base pairs, to minimize concatenation artifacts.

[0039] "Multiplexing" or "sample multiplexing" means the process of combining nuclei or cells from two or more distinct biological samples that have been individually tagged with unique sample-specific barcodes, followed by processing the combined samples together in a single droplet-based capture reaction, library preparation, and sequencing workflow. This approach enables the simultaneous analysis of multiple independent samples while reducing per-sample costs and eliminating technical batch effects. "Doublet" or "doublet event" means an artifact occurring when a single droplet or reaction vessel contains nuclei from two or more different biological samples, identifiable by the association of a single cell barcode with more than one unique sample-specific barcode. Doublets can be computationally detected and removed from analysis to improve data quality7. For example, experimental validation demonstrates doublet detection rates of approximately 5% between samples from the same species and up to 20% between samples from different species. Doublets with the same sample-specific barcode can be recognized computationally. For example, this can be accomplished byfirst training a deep neural network model on doublets with more than one unique barcode. Once trained, this model can be applied to predict doublets with the same barcode based on either the RNA or the ATAC features.

[0040] "Demultiplexing" means the computational process of assigning sequencing reads back to their originating biological samples based on the sample-specific barcodes incorporated during the transposition reaction. This process involves identifying and extracting sample-specific barcodes from sequencing reads, sorting reads into samplespecific datasets, and associating gene expression data with chromatin accessibility data through shared cell barcodes. Successful demultiplexing typically can achieve, for example, greater than 90% efficiency, with over 90% of chromatin accessibility sequencing reads containing recognizable barcodes.

[0041] "Concatenation" means the undesired formation of tandem repeats or linked sequences of transposon oligonucleotides during the transposition reaction, typically occurring when oligonucleotides exceed, for example, 60 base pairs in length. Concatenation products appear as periodic ladder patterns with approximately 60 base pair periodicity when analyzed by capillary electrophoresis and interfere with sequencing library preparation by creating artifacts that cannot be removed by standard size selection methods4.

[0042] "Droplet-based capture" means a microfluidic process wherein individual nuclei are co-encapsulated with barcoded gel beads in aqueous droplets surrounded by oil, creating Gel-Beads in Emulsion (GEMs), to enable the capture and barcoding of molecules from single cells or nuclei for subsequent sequencing library preparation.

[0043] "Cross-platform compatibility" means the ability of the multiplexed method to produce sequencing libraries having size distribution profiles and analytical characteristics substantially identical to libraries generated by standard non-multiplexed protocols, thereby enabling processing using existing commercial equipment, reagents, and analysis software without requiring platform-specific modifications.

[0044] "Cell barcode" means a unique nucleotide sequence incorporated onto molecules during the droplet-based capture reaction that serves to identify all molecules (e.g., mRNA-derived cDNA and transposed genomic DNA fragments) originating from a single nucleus or cell. The cell barcode is provided by the barcoded gel bead and is distinct from the sample-specific barcode, allowing molecules from the same nucleus to be computationally linked together after sequencing while maintaining the ability to assign that nucleus to its sample of origin.

[0045] "Unique molecular identifier (UMI)" means a short, random nucleotide sequence added to individual molecules during the reverse transcription step of the droplet-based capture reaction. UMIs enable the identification and removal of PCR amplification artifacts by allowing the distinction between multiple copies of the same original molecule versus multiple different original molecules, thereby providing more accurate quantification of gene expression levels and chromatin accessibility fragments.

[0046] "Alignment rate" means the percentage of sequencing reads that successfully map to a reference genome or transcriptome during bioinformatics analysis. In the context of the multiplexed method, alignment rates serve as a quality control metric and validation tool, with, for example, alignment rates of approximately 90% when reads are mapped to their appropriate species genome and significantly lower rates (e.g., approximately 20%) when mapped to phylogenetically distant species genomes.

[0047]

[0048] "Transcription start site (TSS) enrichment" means a quality control metric that measures the relative enrichment of chromatin accessibility sequencing reads at annotated transcription start sites compared to background levels. TSS enrichment values can be used to assess the quality' of ATAC-seq data, with higher values indicating better signal- to-noise ratios and successful capture of regulatory regions.

[0049] "Nucleosome signal" means a quality control metric derived from the size distribution of chromatin accessibility sequencing fragments that reflects the nucleosomal organization of chromatin. The nucleosome signal can be calculated based on the ratio of mono-nucleosome-sized fragments (e.g., approximately 147 base pairs plus linker DNA) to nucleosome-free fragments, providing an indication of chromatin structure preservation during sample processing.

[0050] "ATAC features" means the number of distinct accessible chromatin regions or peaks identified per nucleus in single-nucleus ATAC sequencing data. This metric serves as a quality control measure for individual nuclei, with higher feature counts generally indicating better data quality and more comprehensive capture of chromatin accessibility information.

[0051] "Demultiplexing efficiency" means the percentage of sequencing reads that contain recognizable sample-specific barcodes and can be confidently assigned to their originating biological samples during computational analysis.

[0052] "Read counts" means the total number of sequencing reads assigned to a particular nucleus, gene, or genomic region after quality filtering and alignment. Read counts serveas a unit for quantitative analysis in both gene expression (RNA read counts) and chromatin accessibility (ATAC read counts) measurements, with appropriate distributions for downstream biological interpretation.

[0053] "Quality control metrics" means a set of computational measurements used to assess the technical quality7of single-nucleus multi-omic data, including but not limited to: RNA read counts per nucleus, number of detected genes per nucleus, percentage of reads mapping to mitochondrial genomes, ATAC read counts per nucleus, number of ATAC features per nucleus, TSS enrichment, and nucleosome signal. These metrics enable identification and filtering of low-quality nuclei to ensure robust dow nstream analysis.

[0054] The present disclosure provides a method, compositions, and a kit for the multiplexed analysis of single cells or single nuclei, which allows for the simultaneous profiling of multiple molecular modalities, such as gene expression (transcriptome) and chromatin accessibility (epigenome), from multiple distinct biological samples in a single reaction. This method, referred to herein as Multiplexed Multiome (MuMu), addresses critical limitations of cost and technical batch effects associated with prior art methods.

[0055] The method of the present disclosure is distinct from and provides significant advantages over other known multiplexing strategies in the art. While other methods for single-cell multiplexing exist, they ty pically rely on different mechanisms that are more complex, less direct, or not readily compatible with simultaneous ATAC and gene expression profiling of single nuclei.

[0056] For example, some methods employ combinatorial fluidic indexing, where samples are split across multiple wells of a plate, subjected to a first round of barcoding, pooled, and then split again for a second round of barcoding. While effective for some applications like RNA-seq, this process is substantially more complex and labor-intensive than the present methodology, which requires only a single barcoding reaction for each sample prior to a single pooling step. The MuMu method's simplicity7reduces potential for sample loss and user error.

[0057] Other multiplexing techniques, often referred to as "cell hashing," use antibodies or lipid molecules (McGinnis, C.S., Patterson, D.M., Winkler, J. et al. MULTI-seq: sample multiplexing for single-cell RNA sequencing using lipid-tagged indices. Nat Methods 16, 619-626 (2019)) conjugated to barcoded oligonucleotides to "stain" cells or nuclei from different samples with a sample-specific tag. These methods, however, depend on the presence of ubiquitous surface proteins that can be targeted by the antibodies, which may not be consistently present or accessible on isolated nuclei. Moreover, they introduce anentirely separate experimental step (antibody incubation and washing), adding time and potential variability to the workflow. In contrast, the present disclosure integrates the sample barcoding step directly into the ATAC-seq transposition reaction, a step that is already native to the multi-ome workflow. This elegant integration eliminates the need for extra reagents (like antibodies and lipid molecules), separate incubation steps, or complex fluidic handling, providing a more streamlined, cost-effective, and robust solution for multiplexing. The direct ligation of a sample barcode to open chromatin fragments is a novel and non-obvious approach that distinguishes the MuMu method from these other strategies.

[0058] In a typical prior art single-cell multi-ome workflow, such as the one employed by the 10X Genomics Chromium platform, each biological sample is processed individually. Nuclei from a single sample are transposed in bulk, and then individual nuclei are encapsulated in Gel-Beads in Emulsion (GEMs). Within each GEM, a unique cell barcode from the gel bead is attached to both mRNA-derived cDNA and transposase-accessible DNA fragments, linking these two modalities to a single nucleus. However, because each biological sample is processed in a separate, expensive reaction, this workflow is costly and susceptible to batch effects.

[0059] The embodiments described herein address and, in some embodiments overcomes, these limitations by introducing a sample-specific barcode during the initial transposition reaction, prior to the pooling of samples and droplet-based single-nucleus capture. Nuclei from multiple independent biological sources (e.g.. Sample 1 , Sample 2, etc.) are each transposed separately using a transposome carry ing a unique, sample-identifying barcode. After this barcoding step, the nuclei from all samples are pooled and processed in a single droplet-based capture reaction. This multiplexed approach dramatically reduces reagent costs and eliminates batch-to-batch variability between samples, as all samples are processed together under identical conditions from the point of pooling onwards.

[0060] The method for multiplexed analysis of biological samples enables significant cost reduction and elimination of technical batch effects. The method involves providing multiple biological samples each containing nuclei, performing separate transposition reactions on each sample using sample-specific barcoded transposons to tag genomic DNA with unique barcodes, pooling the barcoded nuclei, and subjecting the pooled nuclei to a single droplet-based capture reaction for co-encapsulation with barcoded gel beads.

[0061] In preferred embodiments, the method enables simultaneous single-nucleus RNA sequencing (snRNA-seq) and single-nucleus ATAC sequencing (snATAC-seq), allowingresearchers to analyze both gene expression and chromatin accessibility from the same nucleus while multiplexing multiple samples in a single reaction.

[0062] The method utilizes a transposome complex comprising aTn5 transposase enzyme and sample-specific barcoded transposons. One important embodiment is the use of double-stranded oligonucleotides, preferably with an overall length of less than about 60 base pairs, preferably from about 42 to about 45 base pairs, to attenuate and / or prevent concatenation artifacts that interfere with library preparation. One exemplary oligonucleotide sequence structure corresponds to 5'- GTGCTCTTCCGATCT[X]nAGATGTGTATAAGAGACAG-3', where [X]n represents the unique sample-specific barcode (SEQ ID NO: 1).

[0063] The method can further comprise generating both gene expression and chromatin accessibility sequencing libraries from molecules captured on the barcoded gel beads, followed by sequencing to produce reads containing both cell barcodes and samplespecific barcodes.

[0064] One advantage of the method is the ability to computationally demultiplexing sequencing reads from the chromatin accessibility library based on the unique sample barcodes, thereby assigning each read to its originating sample. Gene expression reads can then be assigned to samples by associating cell barcodes betw een the demultiplexed ATAC and GEX data.

[0065] The method provides an elegant solution for identifying and removing doublet events by detecting when a single cell barcode is associated with more than one unique sample barcode, indicating that multiple nuclei from different samples were captured in the same droplet.

[0066] The oligonucleotides are specifically designed for sample multiplexing in transposonbased assays, comprising a central sample-specific barcode sequence flanked by mosaic end (ME) sequences recognized by transposase, and in certain preferred embodiments have a length of less than about 60 base pairs to reduce, attenuate or prevent concatenation.

[0067] Kits are embodiment that can be used to perform multiplexed multi-omic analysis containing multiple distinct barcoded oligonucleotides with unique sample-specific barcodes, optionally including transposase enzymes such as naked Tn5 transposase, buffers for nuclei isolation and transposition, and primers for I i brar amplification.

[0068] Further, complete systems comprising multiple distinct transposome complexes with unique sample barcodes, microfluidic devices for droplet-based capture, and computer-readable media with instructions for demultiplexing sequencing data by identifying sample-specific barcodes and assigning reads to their samples of origin may be utilized.

[0069] The method can be compatible with existing commercial single-cell platforms without requiring modifications to standard operating procedures, producing sequencing libraries with size distributions substantially identical to non-multiplexed protocols.

[0070] While optimized for simultaneous snRNA-seq and snATAC-seq, the multiplexing strategy can be adapted for other applications including standalone single-nucleus ATAC- seq, single-cell CUT&TAG, and bulk ATAC-seq multiplexing.

[0071] The multiplexed analysis method can achieve exceptional computational demultiplexing performance, wherein at least about 90% of the chromatin accessibility sequencing reads contain recognizable unique barcodes and are successfully assigned to their sample of origin during the computational demultiplexing process. This high efficiency can be achieved through the design of the sample-specific barcoded transposons and optimized transposition conditions, ensuring that the vast majority of transposed genomic DNA fragments carry identifiable sample barcodes that can be computationally extracted and used for accurate sample assignment.

[0072] The computational demultiplexing process has quantitative accuracy with sample assignment rates of from, e.g., 90% to 95% based on barcode recognition and read mapping algorithms. This precision level can be maintained through pattern recognition algorithms that identify and extract sample-specific barcodes from sequencing reads while accounting for potential sequencing errors and variations in barcode quality.

[0073] The transposition reaction preferably achieves superior performance metrics, with greater than about 90% of raw chromatin accessibility reads containing identifiable sample-specific barcodes following the transposition step.

[0074] The method can obtain high doublet identification capabilities with species-dependent detection rates that reflect the biological reality of droplet capture statistics. When multiplexing samples from the same species, doublet events are identified at rates of from about 3% to about 8%, while multiplexing samples from different species yields doublet detection rates of from about 15% to about 25%. These differential rates can provide valuable quality control information and enable researchers to optimize nuclei concentration and capture conditions based on the specific experimental design.

[0075] The method can provide validation capabilities through species-mixing experiments, wherein nuclei exhibit at least about 70% alignment to their expected species genome and less than 30% alignment to non-target species genomes. This validation approachleverages phylogenetic differences to provide data for assessing multiplexing fidelity, with correctly assigned samples typically achieving approximately 90% species-specific alignment rates while showing only about 20% cross-species mapping due to evolutionary divergence.

[0076] The simultaneous single-nucleus RNA sequencing and single-nucleus ATAC sequencing approach can achieve integrated performance metrics comprising: (a) demultiplexing efficiency of at least about 90% for chromatin accessibility reads; (b) doublet detection rates of about 5% to about 20% depending on species origin; and (c) species-specific alignment validation with greater than about 70% accuracy.

[0077] The multiplexed analysis demonstrates measurable performance advantages with quantitative metrics including demultiplexing efficiency above about 90%, doublet identification capability with species-dependent rates, and cross-contamination levels maintained, e.g., below 0.3% of total reads.

[0078] The computational demultiplexing process can incorporate automated qualityassessment features that flag samples for review when demultiplexing efficiency falls below about 85% or when cross-contamination exceeds predetermined thresholds.

[0079] The method can include doublet detection algorithms that identify and remove doublet events exceeding expected statistical ranges based on droplet capture probabilities. By comparing observed doublet rates to theoretical expectations, the system can detect systematic errors in sample preparation or processing and maintain high data quality standards through statistical validation of capture efficiency.

[0080] An aspect of the disclosure is a novel barcoded oligonucleotide for use in the transposon assembly. An exemplary oligonucleotide comprises a central, sample-specific barcode sequence flanked by sequences recognized by a transposase, such as the mosaic end (ME) sequences for the Tn5 transposase.

[0081] In an embodiment, the overall length of the barcoded oligonucleotide is designed to be less than about 60 base pairs (bp) to reduce or prevent transposon concatenation. In certain embodiments, longer transposon oligonucleotides, for example 60 bp in length have a tendency to form concatenations or repeats during the transposition reaction. These concatenated products, which appear as a periodic ladder of fragments on a TapeStation trace interfere with subsequent sequencing library preparation and can be difficult to remove by size selection without also losing desired genomic fragments.

[0082] By reducing the oligonucleotide length, this concatenation issue is resolved. In certain embodiments, the barcoded oligonucleotide has a length of approximately 42 bp to 45 bp.

[0083] In one non-limiting embodiment, a 42 bp barcoded oligonucleotide has a structure corresponding to the sequence: 5'-GTGCTCTTCCGATCT[XXXXXXXX]AGATGTGTATAAGAGACAG-3', wherein "[XXXXXXXX]" represents a variable, sample-specific barcode sequence (SEQ ID NO: 1). The length of this barcode sequence can be varied (e.g., 6, 8, or 10 bp) to accommodate the desired number of multiplexed samples, with appropriate adjustments to the overall oligonucleotide length to remain below the 60 bp threshold. A plurality of such oligonucleotides, each with a unique barcode sequence, can be synthesized to multiplex a corresponding number of samples. This oligonucleotide is designed to be annealed to a reverse complement oligonucleotide (e.g., ME rev: / 5Phos / CTGTCTCTTATACACATCT) (SEQ ID NO:2) before being assembled with a transposase enzyme, such as Tn5, to form a functional, barcoded transposome.

[0084] The transposome design has been validated to produce sequencing libraries with size distribution profiles substantially identical to those generated by standard 10X Genomics protocols. TapeStation analysis demonstrates that libraries prepared using the MuMu protocol exhibit trace characteristics indistinguishable from non-multiplexed controls, confirming that the sample multiplexing does not compromise library quality or introduce technical artifacts.

[0085] The computational demultiplexing process leverages the unique sample-specific barcodes inserted during the transposition reaction. Following paired-end sequencing with dual indexing, the raw sequencing data contains ATAC reads that can carry three barcode elements: (1) the cell barcode from the gel bead, (2) the unique molecular identifier (UMI), and (3) the sample-specific barcode from the MuMu transposon.

[0086] The demultiplexing pipeline begins with the identification and extraction of samplespecific barcodes from the ATAC sequencing reads using bioinformatics tools such as cutadapt. The process involves recognizing the mosaic end sequences flanking the sample barcodes and splitting reads accordingly.

[0087] In certain embodiments, the computational workflow comprises several sequential steps:1 . Adapter Removal and Read Splitting: Raw FASTQ files are processed to remove sequencing adapters and split reads based on the presence of sample-specific barcode sequences. The process utilizes pattern recognition to identify reads containing the expected barcode structure.2. Sample-Specific File Generation: Reads are sorted into sample-specific files based on the identified barcodes, creating separate datasets for each multiplexed sample while maintaining the cell barcode and UMI information.3. Gene Expression Assignment: Since GEX and ATAC data from the same nucleus share identical cell barcodes, gene expression reads are assigned to their samples of origin by associating cell barcodes between the demultiplexed ATAC data and the combined GEX dataset4. Quality Control and Validation: The demultiplexing efficiency and crosscontamination rates are assessed through alignment analysis and species-specific validation when applicable.

[0088] A unique advantage of the system is its ability to identify doublet events — single droplets containing nuclei from multiple samples. Doublets are computationally detected by identifying cell barcodes associated with more than one unique sample barcode. Doublet detection rates can be approximately about 5% between samples from the same species and up to about 20% between samples from different species, consistent with expected droplet capture statistics.

[0089] The demultiplexed data can maintain quality metrics comparable to standard singlecell multi-omics datasets, including appropriate distributions of RNA read counts, detected gene counts, ATAC read counts, transcription start site (TSS) enrichment, and nucleosome signals. This confirms that the multiplexing approach does not compromise the biological information content or technical quality of the resulting datasets.

[0090] The demultiplexed sample-specific datasets can be made to be compatible with standard single-cell analysis software, including cellranger-arc for initial processing and downstream analysis tools such as Seurat and Signac. This compatibility can faciliate integration into existing analytical workflows without requiring specialized software or modified analysis procedures.

[0091] The computational demultiplexing process can be scalable, accommodating various numbers of multiplexed samples while maintaining accuracy and efficiency.

[0092] The methodology7described herein, in certain embodiments, provides a validation method for confirming multiplexing accuracy that leverages phylogenetic differences between species. This validation approach can include providing nuclei from biological samples of at least two different species, performing separate transposition reactions on nuclei from each species using distinct sample-specific barcoded transposons, pooling the barcoded nuclei and subjecting them to droplet-based capture, generating sequencing libraries and sequencing reads, computationally mapping the sequencing reads toreference genomes of each species, and validating multiplexing accuracy by confirming that nuclei assigned to each sample exhibit preferential mapping to their corresponding species genome.

[0093] The species-mixing validation method can enable detection of cross-contamination by identify ing nuclei that map predominantly to a genome different from their assigned sample barcode. In certain embodiments, nuclei exhibiting greater than about 50% of reads mapping to an opposing species genome are identified as contaminated and removed from analysis, providing a quantitative threshold for data quality control.

[0094] One advantage of the species-mixing approach is its ability7to validate doublet detection capability' by identifying cell barcodes associated with sample-specific barcodes from different species and confirming that detected doublet rates fall within expected ranges for droplet-based capture systems.

[0095] The validation method can provide quantitative benchmarks for system performance, confirming that at least about 90% of sequencing reads contain recognizable samplespecific barcodes and are correctly assigned to their originating samples. For instance, species-mixing validation can demonstrate demultiplexing accuracy by confirming that nuclei from a first species exhibit at least about 70% alignment to the first species genome and less than about 30% alignment to a second species genome, while nuclei from the second species exhibit the reciprocal mapping pattern .

[0096] The species-mixing validation can serve as a comprehensive qualify control assay for detecting barcode swapping, cross-contamination, and systematic errors in multiplexed sample processing. This can provide a standardized method for qualify control that can be applied to any multiplexed multi-omic analysis system by processing control samples from at least two phylogenetically distinct species, analyzing genome mapping patterns of the resulting data, and using the mapping patterns to validate system performance and detect technical artifacts.

[0097] The species-mixing validation method can be implemented using human and mouse tissue samples. Despite their closer phylogenetic relationship, rat and mouse samples can serve as a validation pair. Alternatively, other species include, for example, macaque, marmoset, pig, sheep, cow, dog or other mammals. Avian, fish Drosophila, C. elegans, plants, fungi, bacterial, yeast and other can be combined for offering practical advantages for agricultural genomics, veterinary research, or comparative physiology studies.

[0098] In certain embodiments, provided herein is a computer-implemented method for demultiplexing multiplexed single-nucleus sequencing data, comprising: (a) receivingsequencing data files containing chromatin accessibility reads, wherein each read comprises a cell barcode, a unique molecular identifier, and a sample-specific barcode incorporated during transposition; (b) computationally identifying and extracting samplespecific barcodes from the chromatin accessibility reads using pattern recognition algorithms; (c) sorting the reads into sample-specific datasets based on the identified barcodes: (d) assigning gene expression reads to their sample of origin by associating cell barcodes between the demultiplexed chromatin accessibility data and combined gene expression data; and (e) generating sample-specific output files compatible with standard single-cell analysis software.

[0099] In certain embodiments, the computational identification of sample-specific barcodes achieves a demultiplexing efficiency of at least about 90%, with over about 90% of chromatin accessibility sequencing reads containing recognizable barcodes and being successfully assigned to their originating samples. In certain embodiments, the pattern recognition algorithms utilize sequence alignment tools including cutadapt to identify mosaic end sequences flanking the sample-specific barcodes and split reads accordingly.

[0100] In certain embodiments, provided herein is a software system for processing multiplexed multiome sequencing data, comprising: (a) a preprocessing module configured to remove sequencing adapters and identify sample-specific barcode sequences from raw FASTQ files; (b) a demultiplexing engine configured to sort sequencing reads based on barcode identity using configurable pattern matching algorithms; (c) a quality control module configured to calculate performance metrics including demultiplexing efficiency and cross-contamination rates; and (d) an output generator configured to produce sample-specific files compatible with existing single-cell analysis pipelines including cellranger-arc.

[0101] The software system, in an embodiment, includes a doublet detection module configured to identify cell barcodes associated with more than one unique sample barcode and computationally remove corresponding data from dow nstream analysis.

[0102] In certain embodiments, the preprocessing module implements a multi-step pipeline comprising: (i) adapter removal using cutadapt with specified length parameters, (ii) read splitting to separate barcode and genomic sequences, and (iii) barcode-specific file generation using configurable FASTA reference files.

[0103] In other embodiments, a computer-implemented validation method for multiplexed analysis systems is provided and includes (a) receiving sequencing data from samples containing nuclei from at least two phylogenetically distinct species; (b) computationallymapping sequencing reads to reference genomes of each species using alignment algorithms; (c) calculating species-specific alignment rates for each computationally defined nucleus; (d) identifying cross-contamination by detecting nuclei that map predominantly to a genome different from their assigned sample barcode; and (e) validating multiplexing accuracy by confirming that nuclei exhibit preferential alignment to their expected species genome.

[0104] In certain embodiments, the validation can confirm that nuclei from a first species exhibit at least 70% alignment to the first species genome and less than 30% alignment to a second species genome, while nuclei from the second species exhibit the reciprocal mapping pattern.

[0105] In certain embodiments, a computer-implemented validation method can also include computationally calculating doublet detection rates and can confirm that rates fall within ranges of 3-8% between samples from the same species and 15-25% between samples from different species.

[0106] In certain embodiments, a computer-implemented quality control system for multiplexed single-nucleus analysis is provided that includes (a) a metrics calculation engine configured to compute performance indicators including demultiplexing efficiency, alignment rates, doublet detection rates, and barcode recognition accuracy;(b) a validation module configured to compare calculated metrics against predetermined thresholds and identify protocol failures; (c) a reporting system configured to generate quality control summaries and performance visualizations; and (d) an alert system configured to flag datasets that do not meet established qualify standards. In certain embodiments, the metrics calculation engine is configured to generate standard single-cell qualify control metrics including RNA read counts per nucleus, detected gene counts, ATAC read counts per nucleus, transcription start site enrichment, and nucleosome signal.

[0107] In certain embodiments, there is provided a software interface system for integrating multiplexed analysis with existing single-cell platforms, comprising: (a) format conversion modules configured to generate output files compatible with cellranger-arc. Seurat, and Signac analysis workflows; (b) metadata management systems configured to maintain sample identify and barcode associations throughout the analysis pipeline; (c) batch processing capabilities configured to handle multiple multiplexed datasets simultaneously; and (d) standardized output formatting ensuring seamless integration with downstream analysis tools. In certain embodiments, the format conversion modules generate demultiplexed sample-specific datasets that maintain qualify metrics comparableto standard single-cell multi-omics datasets and are fully compatible with existing analytical workflows without requiring specialized software modifications.

[0108] In certain embodiments, there is provide a computer-implemented method for barcode-based sample assignment, comprising: (a) implementing a hierarchical data structure to store and query sample-specific barcode sequences; (b) applying fuzzy matching algorithms with configurable error tolerance to account for sequencing errors in barcode identification; (c) utilizing hash table-based lookup methods for efficient barcode-to-sample mapping; (d) implementing parallel processing algorithms to handle high-throughput sequencing data; and (e) generating assignment confidence scores for each read based on barcode match quality. In certain embodiments, this method can also be configured for implementing real-time monitoring algorithms that track demultiplexing progress and provide interim performance statistics during processing.

[0109] In certain embodiments, there is provided a machine learning-enhanced system for multiplexed analysis optimization, comprising: (a) pattern recognition algorithms trained to identify optimal barcode sequences and predict concatenation likelihood; (b) predictive models configured to estimate demultiplexing efficiency based on experimental parameters; (c) automated quality assessment algorithms that classify datasets as passing or failing based on learned qualify patterns; and (d) adaptive threshold optimization systems that adjust quality control parameters based on dataset characteristics. In certain embodiments, this method can also include error handling procedures for processing corrupted or unexpected sequence patterns, including: (a) detecting barcode sequences with qualify scores below predetermined thresholds; (b) identifying sequence patterns that deviate from expected barcode structures; (c) implementing fuzzy matching algorithms with configurable error tolerance to recover partially corrupted barcodes; and (d) generating error reports documenting the frequency and types of sequence anomalies encountered.

[0110] In certain embodiments, comprising adaptive qualify control measures can be used and can include one or more of (a) dynamically adjusting barcode recognition stringency based on overall sequencing qualify metrics; (b) implementing secondary pattern recognition algorithms when primary barcode identification fails; (c) flagging reads with ambiguous barcode assignments for manual review or alternative processing; and (d) maintaining detailed logs of error recovery' actions for qualify assurance purposes .[Ollljln certain embodiments, there is provided a computer-implemented method for recovering sample assignment information from degraded sequencing data, comprising: (a) analyzing read quality scores along the length of sample-specific barcode regions: (b) implementing sliding window algorithms to identify high-confidence subsequences within low-quality reads; (c) using probabilistic models to predict most likely barcode identity based on partial sequence information; (d) cross-referencing cell barcode associations to validate sample assignments when direct barcode identification fails; and (e) implementing confidence scoring systems that classify' assignment reliability. The probabilistic models can utilize Bayesian inference to calculate assignment probabilities based on observed nucleotide patterns, expected barcode frequencies, and sequencing error rates specific to the sequencing platform employed. Contamination detection and correction procedures can be utilized and can include: (a) identify ing reads containing multiple distinct sample-specific barcode sequences indicating potential crosscontamination; (b) detecting unexpected barcode combinations that deviate from experimental design parameters; (c) implementing statistical algorithms to distinguish genuine doublets from barcode switching artifacts; (d) automatically flagging samples with contamination rates exceeding predetermined thresholds; and (e) providing contamination correction recommendations based on contamination patterns identified.

[0112] The multiplexed multi-ome analysis system can incorporate automated error recovery mechanisms designed to maintain high data quality and processing reliability throughout the computational pipeline. The system can use real-time monitoring algorithms that continuously detect processing anomalies during demultiplexing operations, enabling immediate identification of issues such as corrupted sequence data, unexpected barcode patterns, or systematic processing errors1.

[0113] When primary processing algorithms encounter difficulties or failures, the system can automatically activates backup processing pipelines that utilize alternative algorithmic approaches to maintain data processing continuity. These backup systems employ machine learning models specifically trained to recognize and classify different types of sequence corruption, achieving classification accuracy of at least 95% for distinguishing between recoverable and non-recoverable sequence data.

[0114] The error recovery framework can include adaptive parameter adjustment systems that dynamically optimize processing parameters based on real-time assessment of data quality characteristics. The system can also incorporate automated reporting mechanismsthat generate comprehensive summary statistics of all error recovery' actions, providing detailed documentation for quality assurance and reproducibility purposes.

[0115] The computational system can use comprehensive quality validation procedures specifically designed for error-corrected sample assignments. These procedures can include sophisticated cross-validation algorithms that verify sample assignments using multiple independent evidence sources, providing robust confirmation of demultiplexing accuracy.

[0116] The validation framework can include confidence intervals for assignment accuracy based on sequence qualify metrics, enabling quantitative assessment of data reliability. Statistical testing procedures can validate that error correction procedures maintain expected doublet detection rates and species-specific alignment patterns, ensuring that corrective measures do not compromise the biological integrity of the results.

[0117] Cross-validation algorithms can utilize alignment patterns to reference genomes as independent validation of sample assignments, with experimental validation confirming that error-corrected assignments and can generate comprehensive quality reports that quantify the impact of error correction on final data integrity and provide specific recommendations for data inclusion or exclusion based on established quality assessment criteria.

[0118] The multiplexed analysis system can use a fault-tolerant computational pipeline architecture designed to handle the complex processing requirements of multiplexed single-cell data. The pipeline can use modular error handling systems that isolate processing failures to prevent pipeline-wide crashes, ensuring that individual component failures do not compromise the entire analysis w orkflow- '.

[0119] The architecture can include sophisticated checkpoint and recovery mechanisms that enable resumption of processing from specific failure points, minimizing data loss and computational resource waste. Parallel processing capabilities can distribute error-prone computations across multiple processing threads, improving both performance and reliability while maintaining data integrity.

[0120] The system can use advanced error correction algorithms designed to maximize data recovery from partially corrupted or degraded sequencing information. These algorithms can include consensus sequence determination methods that analyze multiple partial barcode observations within the same cell to reconstruct complete barcode sequences.

[0121] Neighborhood analysis algorithms can identify and correct isolated assignment errors by analyzing local barcode consistency patterns, leveraging the spatial relationshipsbetween related sequences to improve assignment accuracy. Temporal analysis of sequencing can run metrics enables identification and compensation for systematic sequencing errors that may occur during specific phases of the sequencing process *.

[0122] The error correction framework can include specialized barcode reconstruction algorithms that can assemble complete barcodes from fragmented sequence information.

[0123] An exemplary embodiment of the multiplexed multiome method is illustrated below. The protocol is compatible with standard droplet-based single-cell platforms, such as the 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit, with key modifications to enable multiplexing.

[0124] Transposon and Transposome Preparation: The process begins with the preparation of custom barcoded transposons. For each biological sample to be multiplexed, a unique barcoded oligonucleotide is annealed with a reverse complement oligonucleotide (e.g., ME_rev). In parallel, a standard transposon (e.g., ME-A / ME-Rev) is also prepared. Each unique barcoded transposon is then incubated with a transposase enzyme (e.g., naked Tn5 transposase) and the standard transposon at an equal molar ratio to form a functional transposome complex. This transposome is now "loaded" and ready to simultaneously cut genomic DNA and ligate the barcoded adapter sequence. The assembled transposomes can be stored at -20°C for future use.

[0125] Nuclei Extraction: High-quality nuclei are extracted from each biological sample independently. Samples may be, for example, frozen tissues or cell suspensions. The samples are thawed in a chilled lysis buffer (e.g., containing Tris-HCl, NaCl, MgC12, and detergents such as IGEPAL CA-630, Tween-20, and Digitonin) to permeabilize the cell membrane while keeping the nuclear membrane intact. Following lysis, nuclei are washed, resuspended in a suitable buffer (e.g.. Nuclei Resuspension Buffer with BSA), filtered to remove clumps, and counted using a hemacytometer to accurately determine the concentration. The concentration is adjusted to a target level (e.g., about 1,000 nuclei / pL) to optimize for single-nucleus capture efficiency in the downstream steps.

[0126] Barcoded Transposition Reaction: This is the multiplexing step. For each biological sample, a separate transposition reaction is performed. A defined number of nuclei (e.g., 5,000-10,000) are incubated with the corresponding sample-specific barcoded transposome prepared in step 1. The reaction is typically carried out at 37°C for approximately 1 hour. During this incubation, the transposome enters the nuclei andinserts the barcoded oligonucleotides into open chromatin regions of the genome. Each sample's genomic DNA is thus uniquely "tagged" with its barcode of origin.

[0127] Nuclei Pooling and Droplet Capture: After the individual transposition reactions are complete, the nuclei from all samples are collected and pooled into a single tube. From this point forward, all samples are processed together in a single reaction volume. The pooled nuclei are then loaded onto a microfluidic chip (e.g.. a 10X Genomics Chromium chip) along with barcoded gel beads, reagents for reverse transcription, and partitioning oil. This generates GEMs, where the majority of droplets contain at most one nucleus and one gel bead. Inside each GEM, the nucleus is lysed, and both the pre-mRNAs and the barcoded genomic DNA fragments are captured on the gel bead. A reverse transcription reaction then adds the gel bead's unique cell barcode and a unique molecular identifier (UMI) to all molecules.

[0128] Library Preparation and Sequencing: Following droplet capture, the emulsion is broken, and the pooled, barcoded molecules are collected. The material undergoes a preamplification step to generate sufficient material for both gene expression (GEX) and ATAC library construction. The pre-amplified product is then split. One portion is used to generate the final ATAC-seq library through PCR, which adds sequencing adapters. The other portion is used to create the final GEX sequencing library'. The resulting ATAC and GEX libraries are subjected to quality control, for example using an Agilent TapeStation. The libraries are then sequenced on a compatible platform using a paired-end, dual-index configuration.

[0129] Data Pre-processing and Demultiplexing: After sequencing, the raw data files (FASTQs) are processed to assign each sequencing read back to its sample of origin. The ATAC sequencing reads contain three key barcodes: the cell barcode (from the gel bead), the UMI, and the sample-specific barcode (from the MuMu transposon). A bioinformatics pipeline, for instance using tools like ' cutadapf , is used to identify and extract the sample-specific barcode from the ATAC reads. The reads are then sorted into separate files based on this barcode. This demultiplexing step computationally separates the data from the pooled samples. Since the GEX and ATAC data from the same nucleus share a common cell barcode, the GEX data can be assigned to its sample of origin by association with the now-demultiplexed ATAC data. The demultiplexed data for each sample can then be processed using standard analysis software, such as cellranger-arc.Examples

[0130] The protocol below describes the specific steps for sample multiplexing before performing a sn-Multiome run with 10X Genomics Chromium Single Cell Gene Expression + ATAC kit. Sample multiplexing in this protocol refers to the action of combining single nuclei from multiple independent biological sources before performing droplet-based single nuclei capture and sequencing library preparation in a single reaction as if all the nuclei come from the same sample, significantly reducing the cost of analyzing multiple independent samples. Single cell multi-omics technologies have rapidly advanced in recentyears. Baysoy, A., Bai, Z., Satija, R., and Fan, R. (2023). The technological landscape and applications of single-cell multi-omics. Nature Reviews Molecular Cell Biology 2023 24: 10 24, 695-713. https: / / doi.org / 10.1038 / s41580- 023-00615-w. Simultaneous single nuclei RNA and ATAC sequencing technology (referred to hereafter as multiome sequencing technology), in particular, is well established and commercially available. Single Cell Multiome ATAC + Gene Expression - Official lOx Genomics Support. However, the cost to perform one sn-Multiome reaction is very high. Moreover, biological experiments usually require multiple samples from different time points or conditions, in multiple replications. As a result, the cost to complete a whole experiment quickly rises to a prohibitive amount for many laboratories. Furthermore, when only a small number of nuclei is available as input, the potential output of a sn-Multiome reaction is often under-utilized, exacerbating cost-inefficiency. With the release of high through-put single cell capture platforms such as the Chromium X series (10X Genomics), whose capture capacity goes up to millions of cells or nuclei per capture, a reliable sample multiplexing strategy is now7even more relevant and valuable. Thus, we improved the utility of droplet-based sn-Multiome by redesigning the Tn5 transposon sequence to allow the insertion of a sample specific barcode during the initial transposition reaction of the ATAC workflow and prior to single nuclei capture.

[0131] The sample-specific barcode can be reliably detected in the fastq files containing sequencing reads and easily incorporated into read alignment steps for demultiplexing purposes. While the current protocol is written based on and specifically for the 10X Genomics Chromium Single Cell ATAC + GEX Multiome kit, it can be adapted to other droplet-based sn-Multiome strategies, such as single nucleus ATAC-seq and CUT&TAG with minor modifications. Wu, S.J., Furlan, S.N., Mihalas, A.B., Kaya-Okur, H.S., Feroze. A.H.. Emerson, S.N., Zheng, Y.. Carson, K.. Cimino, P.J.. Keene. C.D.. et al. (2021). Single-cell CUT&Tag analysis of chromatin modifications in differentiation andtumor progression. Nature Biotechnology 2021 39:7 39, 819-824. https: / / doi.org / 10.1038 / s41587-021-00865-z. The same approach can also be used to multiplex bulk ATAC-seq samples.

[0132] Transposon preparation

[0133] Timing: 30min

[0134] Mix 5pL ME-A (lOOpM) with 5pL ME-Rev (lOOpM) and 40pL TE buffer in a PCR tube.

[0135] Mix 5pL of each MuMu barcoding oligomer (lOOpM) with 5pL ME-Rev (lOOpM) and 40pL TE buffer in PCR strips or plates.

[0136] The number of MuMu barcoding oligomers depends on the number of samples to be multiplexed. Prepare at least one barcoding oligomer for one sample. More than one barcoding oligomer can be used for one sample.

[0137] Put all solutions on thermocycler and run the following program:

[0138] MuMu transposome assembly

[0139] Timing: 35min

[0140] Combine the following reagents in a PCR tube and mix by pipetting:The volumes used in this table are typically enough for four reactions, assuming around 10,000 input nuclei per reaction.[Of 41] Incubate at 25°C in a thermocycler for 30min.

[0142] The assembled transposome can be kept at -20°C for at least four weeks.

[0143] Prepare Pre-Amplification primers

[0144] Timing: 5min

[0145] Resuspend Pre-Amplification PCR primers to 100 pM.

[0146] Mix together lOpL of each primer and 60pL of TE buffer (lOOpL total).Key resources table• a Supplied as a 10% solution• b Supplied as a 10% solution; store at 4C• c Supplied as a 2% solution in DMSO; (store at -20C)

[0147] Materials and equipment setupLysis buffer

[0148] One can also use reagents provided by 10X Genomics Chromium Nuclei Isolation Kit.Nuclei Resuspension Buffer

[0149] One can also use Nuclei Resuspension Buffer provided in the 10X Genomics Chromium Next GEM Single Cell ATAC + Gene Expression Multi ome kit.

[0150] Step-by-step method details

[0151] Extract nuclei from each biological sample

[0152] Timing: 30min

[0153] This step extracts and isolates nuclei from frozen samples. The concentration of the final nuclei suspension is adjusted to a level that maximizes capture efficiency while minimizing the formation of doublets (capturing more than one nucleus per droplet).

[0154] Retrieve tissue or cells in a cry opreservation vial from -80°C freezer and immediately add 500pl of lysis buffer to the vial.

[0155] Incubate for lOmin on ice.

[0156] After incubation, add 500pl nuclei resuspension buffer (RSB) and gently pipette 5-10 times until no visible cell lumps.

[0157] Adjust volume to 5ml and centrifuge at 800g for 5min to collect the nuclei.

[0158] Resuspend nuclei in RSB at appropriate volume (usually l-2mL).

[0159] Filter resuspended nuclei into 5mL Polystyrene Round-Bottom Tube with Cell- Strainer Cap

[0160] Count nuclei with a hemacytometer.

[0161] Mix 5pL nuclei suspension with 5pL Trypan Blue Solution 0.4% by gently pipetting at least 10 times.

[0162] Load the mixture into the chamber on one side of the hemacytometer.

[0163] Count the number of nuclei under a microscope and calculate the concentration.

[0164] Adjust nuclei concentration to about 1,000 nuclei / pL.

[0015] Transposing genomic DNA in the extracted nuclei with MuMu barcode

[0166] Timing: Ih lOmin

[0167] This step transposes the open chromatin regions of genome within each nucleus and simultaneously adds the sample specific barcode. At the end of this step, the nuclei from multiple barcoded samples are combined and processed in a single reaction. As an example, we assume an experiment with four biological samples.

[0168] For each sample, combine the following reagents in a 1.5 mL microcentrifuge tube separately:

[0169] Note: If multiplexing more than four samples, adjust the number of input nuclei from each sample so that the final combined number does not exceed the maximum input amount required by the specific platform.

[0170] Incubate mixtures at 37°C for Ih on a thermocycler.

[0171] During incubation, prepare Master Mix.

[0172] Combine reagents according to 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression User Guide, Section 2. la.

[0173] Add 10 pL RSB to Master Mix and keep mixture on ice

[0174] If using 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit, bring out Gel Bead from -80 °C freezer to equilibrate to room temperature.

[0175] Collect post-transposition nuclei by centrifugation.

[0176] After the reaction, the nuclei in each tube are pelleted by centrifugation at 800g for 5 min.

[0177] Record the tubes' orientation in the centrifuge. A pellet may not be visible but nuclei will be collected on the outside wall of the tube after centrifugation.

[0178] Using a P20 pipette, remove all supernatant by pipetting from the opposite side of the pellet.

[0179] Gently wash the outside wall of each centrifuge tube using the master mix prepared in step 3 and combine sequentially.

[0180] Keep master mix containing all the transposed nuclei on ice.

[0181] Capture

[0182] Timing: 2h 30min

[0183] This step encapsules single nuclei within the droplet emulsions. Pre-mRNA and transposed genomic DNA fragments from each nucleus are released by lysis and captured on a gel bead. During reverse transcription, cell barcodes and unique molecular identifier (UMI) are added to mRNA (cDNA) and to genomic DNA fragments. Genomic DNA fragments and cDNA within the same droplet share the same cell barcode. Finally, the emulsion is broken and all molecules are collected for amplification in the next step.

[0184] The single cell capture is conducted faithfully following 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression User Guide, Section 2.2 to Section 3.2.

[0185] The product can be stored at 4 °C over night before proceeding to the next step.

[0186] Pre-amplification

[0187] Timing: 35min

[0188] This step amplifies the genomic DNA fragments and cDNA from the previous step. Specifically for genomic DNA fragments, additional sequences are added at the end as primer binding sites for sequencing library amplification.

[0189] Perform PCR reaction as follows:PCR reaction master mix

[0190] Do not use any “hot start” polymerase since the PCR reaction begins with an initial extension step.

[0191] PCR cycling conditions

[0192] Post pre-amplification cleanup: Follow 10X Genomics Chromium Next GEM Single Cell Multi ome ATAC + Gene Expression User Guide, Section 4.3 to perform cleanup.

[0193] Can be stored at 4°C for up to 72 h or at -20°C for long-term storage.

[0194] ATAC Sequencing Library Amplification

[0195] Timing: 20 min

[0196] This step amplifies the genomic DNA fragments and adds sequencing adaptors at the end.

[0197] Perform PCR reaction as follows:PCR reaction master mixPCR cycling conditions

[0198] ATAC Sequencing Library Post-amplification Cleanup

[0199] Timing: 20 min

[0200] In this step, SPRISelect beads are used to purify fragments with a double-sided size selection strategy. The final library consists of fragments mostly from open chromatin regions of the genome, with a minor population of single and multi-nucleosome fragments.

[0201] Cleanup amplified ATAC sequencing library with SPRIselect beads.

[0202] Vortex to resuspend SPRIselect reagent. Add 60 pl SPRIselect reagent (0.6X) to each sample. Pipette mix.

[0203] Incubate 5 min at room temperature.

[0204] Place on the magnet until the solution clears.

[0205] Transfer 150 pl supernatant to a new strip tube. DO NOT discard the supernatant.

[0206] Vortex to resuspend SPRIselect reagent.

[0207] Add 65 pl SPRIselect reagent (1.25X) to each sample (supernatant). Pipette mix.

[0208] Incubate 5 min at room temperature.

[0209] Place on the magnet until the solution clears.

[0210] Remove the supernatant.

[0211] Add 300 pl 80% ethanol to the pellet. Wait 30 sec.

[0212] Remove the ethanol.

[0213] Add 200 pl 80% ethanol to the pellet. Wait 30 sec.

[0214] Remove the ethanol.

[0215] Centrifuge briefly. Place on the magnet.

[0216] Remove remaining ethanol.

[0217] Remove from the magnet. Immediately add 20.5 pl Buffer EB. Pipette mix.

[0218] Incubate 2 min at room temperature.

[0219] Centrifuge briefly. Place on the magnet until the solution clears.

[0220] Transfer 20 pl sample to a new tube strip.

[0221] Pause point: Store at 4°C for up to 72 h or at -20°C for long-term storage.

[0222] Gene Expression Sequencing Library Preparation

[0223] Timing: 2 days

[0224] This step amplifies the cDNA library, which is subsequently fragmented. Then, the cDNA fragments are ligated with sequencing adaptors and amplified again through PCR reaction to generate the final sequencing library.

[0225] Follow 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression User Guide, Section 6 and 7 to generate gene expression sequencing library.

[0226] Sequencing Library Quality Control

[0227] Timing: 30 min

[0228] This is a routine quality' control step before libraries are sequenced. Successful ATAC and gene expression sequencing libraries show typical size profiles of their corresponding library type. Refer to 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression User Guide, Appendix: Agilent TapeStation Traces section for more details.

[0229] An Agilent TapeStation with the High Sensitivity D5000 kits is used to perform quality control on the final ATAC and gene expression sequencing libraries. Follow the manufacturer’s instructions depending on the exact instrument and kit used (TapeStation online protocol).

[0230] Sequencing

[0231] Timing: 2-4 weeks

[0232] Sequencing both gene expression and ATAC libraries with a paired-end, dual index protocol;Note: The i7 index corresponds to the NEB index and is only relevant during the fastq file generation step. The NEB index is 6-bp in length. If 8-bp is required, append “AT” at the end.

[0233] Provide the NEB index sequence from step 16 to generate fastq files with bcl2fastq function.

[0234] Three fastq.gz files will be generated after sequencing. The file names should be adjusted according to cellranger arc requirement as follows(‘X’ indicates a single digit):

[0235] Read 1 : [Run name]_SX_LXXX_Rl_XXX.fastq.gz

[0236] 15 index: [Run name]_SX_LXXX_R2_XXX.fastq.gz

[0237] Read 2: [Run name]_SX_LXXX_R3_XXX.fastq.gz

[0238] Pre-process MuMu barcode sequences before read alignemnt

[0239] Timing: 2h

[0240] In this step, the fastq files are processed to prepare for alignment with the cellrangerarc function.

[0241] (Optional) Remove sequencing adaptors from R1 and R3.(SEQ ID NO:9)1 . Split fastq file of R3 into two temporary files corresponding to [temporary]_SX_LXXX_l1_XXX.fastq.gz that contains MuMu barcode sequence (“11”) and the other one contains ATAC sequence (“R3”).(SEQ ID NO:10)Note: The R3 file ([Run name]_SX_LXXX_R3_XXX. fastq.gz) is used twice as input for cutadapt to simplify the read splitting process. Refer to cutadapt user guide for more information (REF).Generate sample specific fasta files corresponding to each MuMubarcode:Note: {name} will be replaced by sample name supplied in the MuMu_barcodes.fa, which can be modified by the user. Unzip [Run name]_SX_LXXX_R1_XXX.fastq.gz, [Run name]_SX_LXXX_R2_XXX.fastq.gz and [Run name]_{name}_ SX_LXXX_I1_XXX. fastq. gz split the fastq files by MuMu barcodes usingfastq_pair function.Note: fastq_pair outputs four files. Only the .paird.fq file of R1 (or R2) is useful here and needs to be renamed according to cellranger-arc format requirement of input files.Note: all inputs need to be unzipped fastq files. Make sure you have four sets of ATAC fastq files for each of the MuMu barcoded samples:[Run name]_{name}_SX_LXXX_R1_XXX.fastq.gz[Run name]_{name}_SX_LXXX_R2_XXX.fastq.gz [Run name]_{name}_SX_ LXXX_ R3_XXX. fastq. gz[Run name]_{name}_SX_LXXX_l1_XXX. fastq. gz Run “cellranger-arc count” program for each sample, using the sample specific ATAC fastq files and combined gene expression fastq files as input: cellranger-arc count \-reference={genome_reference} \-id=[Run name]_{name} \--libraries=libraries.csv6. (Optional) Run cellranger-arc aggr to aggregate all samples into one dataset.

[0242] Expected outcomes

[0243] The sn-Multiome ATAC library' generated with the MuMu protocol is identical to ones that are generated with the standard 10X Genomics Chromium Next GEM Single Cell Multiome ATAC + Gene Expression workflow (Figure 2). Sometimes a smaller peak (less than 150 bp) appears in the final library-, yvhich represents adaptor dimer and should not affect sequencing or downstream analyses.

[0244] Quantification and statistical analysis (optional)

[0245] As a demonstration of the MuMu protocol, yve processed tissue samples obtained from a mouse neocortex at El 6 and a pig neocortex at gestation yveek (GW) 10. The nuclei from the two pieces of tissue samples were extracted. We then performed transposition reaction yvith the MuMu barcode oligomers as the transposon on two separate pools of nuclei from each tissue sample. The two mouse nuclei pools correspond to sample 1 and sample 3, while the pig nuclei pools correspond to sample 2 and sample 4, respectively. Following the MuMu preparation and bioinformatics preprocessing pipeline, more than 90% of all the raw ATAC reads contained correct MuMu barcodes and were confidently assigned to their corresponding sample (Figure 3A). We first assessed the rate at which MuMu barcodes mix with one another. When RNA or ATAC reads from any nucleus are mapped to the appropriate genome reference (mouse-to- mouse or pig-to-pig), the alignment rate is around 90%. However, when reads are mapped to the opposing reference, the alignment rate is only around 20%, due to heterology’ between mouse and pig genomes. We demonstrated this pattern by mapping a mouse sn- Multiome dataset (" I OX Multiome”) generated in-house using the original 10X Genomics Multiome protocol to both the mouse and the pig genome references (Figure 3B). Based on this observation, yve reasoned that if there is frequent mixing between MuMu barcodes, one yvould expect a substantial shift in the distribution of mapping ratio and that more nuclei yvould be assigned to one species but have their reads mapped to the genome of the opposing species. Our analysis indicated that all the nuclei displayed the correct pattern of mapping ratio based on the percentage of ATAC reads mapped to the two genomes, suggesting that there is no mixing or interchange of the MuMu barcodes afterthe transposition reactions. While RNA reads from a small number (4.3%) of the total nuclei mapped mainly to the opposing genome, none of these nuclei had a mapping rate of the ATAC reads to the opposing genome at more than 50%, an average of 20. 1% (SE±0.1%). The erroneously mapped RNA reads were most likely due to ambient RNA. This phenomenon can be used to exclude this minority of nuclei. Together, our data demonstrated that the MuMu barcodes faithfully and specifically labeled the samples they were assigned to. The nuclei wi th more than 50% of either RNA or ATAC reads mapped to the opposing genome were removed from downstream analyses.

[0246] A unique advantage of the MuMu barcoding system is its ability' to unambiguously detect doublets, due to its reliability. Thus, we next examined how many cell barcodes (CB) were assigned to more than one sample as an indication of doublets. We found 144 CB shared between Sample 1 and Sample 3, as well as 571 CB shared between Sample 2 and Sample 4 (Figure 3C). The doublet rates were therefore about 5% between the mouse samples and about 20% between the pig samples. While the doublet ratio of the pig samples were higher than the mouse samples, it was still within the expected range of a droplet based single cell experiment (Xi, N.M., and Li, J.J. (2021). Benchmarking Computational Doublet-Detection Methods for Single-Cell RNA Sequencing Data. Cell Syst 12, 176-194.e6.), and may reflect an underestimation of the nuclei concentration during the extraction procedures for one of the pig samples (Sample 4) due to human error. We removed any CB that was assigned to more than one sample from downstream analyses.

[0247] The remaining nuclei displayed compatible distributions of RNA read counts, detected gene counts and percentage of RNA reads mapped to mitochondrial genome to standard single cell RNA-seq datasets. We then performed the Seurat (v5) analyses pipeline (Hao, Y, Stuart, T, Kowalski, M.H., Choudhary, S., Hoffman, P., Hartman, A., Srivastava, A., Molla, G., Madad, S., Fernandez-Granda, C., et al. (2023). Dictionary' learning for integrative, multimodal and scalable single-cell analysis. Nature Biotechnology 2023 42:2 42, 293-304.) and identified cell clusters in uniform manifold approximation and projection graphs (Figure 3E and F). The cell clusters from either species expressed the expected marker genes appropriate for the neocortical developmental period. Notably, El 6 in mouse and GW10 in pig mark the beginning gliogenesis in the two species (Figure 3G and H). As expected, we identified oligodendrocyte precursor cell populations (Mm_2 and Ss_8) in both datasets, whichwere marked by a specific expression of PTPRZ1. The ATAC data also displayed compatible quality to standard single cell ATAC-seq datasets, reflected by the distributions of ATAC read count, number of ATAC features, transcription start site enrichment and nucleosome signal. We then examined the chromatin accessibility landscape around the gene body of PTPRZ1, we observed an enhanced openness of the PTPRZ1 locus. More importantly, in both species, we were able to identify an ATAC peak about 200,000bp upstream of the gene that correlated with the PTPRZ1 RNA expression (Figure 3 J and K).6 The analyses thus identify’ a potentially common regulator}’ element of PTPRZ1 shared by the two species.

[0248] Current state-of-the-art ATAC-seq protocols, including the one described here, utilize transposon oligomers with two different sequences (mixed at equal molarity), so that DNA fragments can be tagged with different oligomers on each end to serve as primer binding sites for the PCR amplification. As a result, however, about half of the transposition events generate DNA fragments with the same oligomer sequences at both ends, and these "homo-lagged" fragments cannot be amplified. Thus, 50% of the transposed DNA fragments are lost by default. Penkov, D., Zubkova, E., and Parfyonova, Y. (2023). Tn5 DNATransposase in Multi-Omics Research. Methods Protoc 6, 24. https: / / doi.org / 10.3390 / MPS6020024.

[0249] This is not a significant drawback in bulk-tissue ATAC-seq because there are presumably many copies of the genome from the large amount of nuclei with the same open chromatin configuration. However, at the single nucleus level, each genome is considered unique, which means the open chromatin fragments from each nucleus are also unique. The loss of 50% fragment from each single nucleus is irreversible and poses a challenge on the interpretation of final sequencing data.

[0250] A limitation specific to the current protocol is the randomness in the sequencing read distribution and nuclei distribution among the MuMu barcoded samples, which is affected by many factors. The sequencing and nuclei distributions are accessible after sequencing. Thus, it is important to accurately estimate nuclei concentration during the extraction steps and practice precise techniques in pipetting nuclei between steps. The transposition reaction, especially the amount of MuMu transposome used, should also be carefully tested beforehand to achieve the optimal number of fragments from open chromatin regions.

[0251] TroubleshootingProblem 1 :

[0252] Failed to see a multimodal ATAC sequencing library profile. Only one peak can be detected at around 200 bp (Step 19).

[0253] This type of library profile usually indicates degradation of chromatin structure, resulting from poor quality of starting material. It is also possible that too much transposome is used. Another possible cause is an error during ATAC sequencing library post-amplification cleanup.

[0254] Potential solution:

[0255] Collect fresh tissue. Be sure to snap-freeze sample as soon as possible in liquid nitrogen and transfer it to -80 °C right away after freezing for long term storage.

[0256] Make sure to perform all nuclei extraction procedure on ice. Centrifuge nuclei at 4 °C until transposition step.

[0257] When in doubt, perform a bulk ATAC-seq with sample in question to assess quality before proceeding to sn-Multiome.

[0258] Try reducing the amount of transposome during the transposition step.

[0259] During ATAC sequencing library’ post-amplification cleanup, make sure SPRISelect beads are well mixed every time and use the recommended ratio of beads to sample.

[0260] Problem 2:

[0261] Nuclei clump after extraction and cannot be dissociated further (Step 1-8).

[0262] This is usually caused by over-lysis resulting in the release of genomic DNA from broken nuclei, due to excessive pipetting or strong lysis buffer.

[0263] Potential solution:

[0264] Pipette gently during nuclei extraction steps and reduce pipetting frequency depending on tissue type.

[0265] Reduce the amount of detergent in the lysis buffer.

[0266] Add DNase I to lysis buffer may help digest free-floating genomic DNA and reduce tangling.

[0267] Problem 3:

[0268] The final concentration of ATAC sequencing library is low (Step 14-16).

[0269] This usually indicates an overestimation of the number of captured nuclei.

[0270] Potential solution:

[0271] Quantification of starting number of nuclei: The initial estimation of nuclei concentration needs to be accurate. Carefully count and calculate the concentration of nuclei after extraction.

[0272] Sufficiently resuspend nuclei containing solution to homogeneity before mixing transposition reaction.

[0273] Carefully collect nuclei after transposition with master mix.

[0274] Further amplify sequencing library7with Illumina P5 and P7 primers.Table 1 : Pre-amplification primers

[0275] A unique advantage of this multiplexing strategy7is the ability to unambiguously identify and remove doublet events (i.e., a single droplet containing nuclei from two different samples). Because each sample has a unique barcode, a doublet containing nuclei from two different samples will be associated with two different sample barcodes. These identified doublets can be computationally filtered out, leading to higher quality final data.

[0276] Furthermore, the quality7of the demultiplexed data was found to be high and comparable to data generated using standard non-multiplexed methods. The distributionsfor metrics like RNA read counts, detected genes per nucleus, ATAC read counts, and Transcription Start Site (TSS) enrichment are all within expected ranges for high-quality single-cell multi-ome data. Downstream analysis of the data successfully identified distinct cell clusters in both the mouse and pig datasets. Moreover, the method enabled the correlation of gene expression with chromatin accessibility7at specific gene loci.

Claims

CLAIMSWhat is claimed is:

1. A method for multiplexed analysis of biological samples, the method comprising: a. providing a plurality of biological samples, each sample containing nuclei; b. for each of the plurality of biological samples, performing a separate transposition reaction on the nuclei using a transposome complex comprising a samplespecific barcoded transposon, thereby tagging genomic DNA within the nuclei with a unique barcode corresponding to the sample of origin; c. pooling the barcoded nuclei from the plurality of biological samples to create a pooled nuclei suspension; and d. subjecting the pooled nuclei suspension to a single droplet-based capture reaction to co-encapsulate individual nuclei with barcoded gel beads.

2. The method of claim 1, further comprising simultaneous single-nucleus RNA sequencing (snRNA-seq) and single-nucleus Assay for Transposase-Accessible Chromatin using sequencing (snATAC-seq).

3. The method of claim 1, wherein the transposome complex comprises a Tn5 transposase enzyme.

4. The method of claim 1, wherein the sample-specific barcoded transposon comprises a double-stranded oligonucleotide, wherein a first strand of the oligonucleotide has an overall length of less than about 60 base pairs to prevent concatenation.

5. The method of claim 4, wherein the first strand of the oligonucleotide has an overall length of from about 42 to about 45 base pairs.

6. The method of claim 4, wherein the first strand of the oligonucleotide comprises a sequence corresponding to 5'-GTGCTCTTCCGATCT[X]nAGATGTGTATAAGAGACAG- 3', wherein [X]n represents the unique barcode corresponding to the sample of origin.

7. The method of claim 2, further comprising:a. generating a gene expression sequencing library and a chromatin accessibility' sequencing library from molecules captured on the barcoded gel beads; and b. sequencing the gene expression and chromatin accessibility7libraries to produce sequencing reads.

8. The method of claim 7, further comprising computationally demultiplexing the sequencing reads from the chromatin accessibility library based on the unique barcode to assign each read to its sample of origin.

9. The method of claim 8, further comprising assigning the gene expression sequencing reads to their sample of origin by associating a cell barcode present on the gene expression reads with a cell barcode present on the demultiplexed chromatin accessibility reads.

10. The method of claim 8, further comprising identifying and computationally removing data corresponding to doublet events, wherein a doublet event is identified by the association of a single cell barcode with more than one unique barcode.

11. The method of claim 1, wherein the plurality7of biological samples are selected from the group consisting of: samples from different species, samples from different individuals of the same species, samples from different tissues of the same individual, samples from different time points, and samples from different experimental conditions.

12. The method of claim 10, wherein the doublet events are identified at a rate of from about 1% to about 25% of total cell barcodes when multiplexing samples from the same species.

13. The method of claim 10, wherein the doublet events are identified at a rate of from about 3% to about 8% when multiplexing samples from a first species and from about 15% to about 25% when multiplexing samples from a second species.

14. The method of claim 10, wherein the method detects doublet events comprising nuclei from different biological samples with at least 95% specificity7.

15. The method of claim 8, wherein at least 90% of the chromatin accessibility sequencing reads contain a recognizable unique barcode and are assigned to their sample of origin.

16. The method of claim 8, wherein from about 90% to about 95% of the chromatin accessibility7sequencing reads contain the unique barcode and are demultiplexed to their originating biological sample.

17. The method of claim 8, wherein the computational demultiplexing achieves a sample assignment accuracy of at least about 90% based on barcode recognition.

18. The method of claim 2, wherein the method has a demultiplexing efficiency of at least about 90% of chromatin accessibility reads containing recognizable barcodes; and adoublet detection rates of from about 1% to about 25% of total captured nuclei.

19. The method of claim 1, wherein the multiplexed analysis provides quantitative performance metrics comprising demultiplexing efficiency of greater than about 90% and doublet identification capability7distinguishing nuclei from different samples with at least about 95% accuracy.

20. A transposome complex for multiplexed analysis of biological samples, the complex comprising: a. a transposase enzy me; and b. a double-stranded nucleic acid construct bound to the transposase enzyme, the construct comprising a first strand and a second strand, wherein the first strand comprises a sample-specific barcode sequence and has an overall length of less than about 60 base pairs.

21. A system for performing multiplexed analysis of biological samples, the system comprising: a. a plurality of distinct transposome complexes, each complex comprising a transposase enzyme and a double-stranded oligonucleotide having a unique sample-specific barcode sequence, wherein the oligonucleotide has an overall length of less than about 60 base pairs;b. a microfluidic device configured to perform a droplet-based capture reaction on a pooled suspension of nuclei previously barcoded with the plurality of distinct transposome complexes; and c. a non-transitory computer-readable medium comprising instructions that, when executed by a processor, perform a method of demultiplexing sequencing data by: i. identifying the unique sample-specific barcode sequence within chromatin accessibility sequencing reads derived from the pooled suspension; and ii. assigning the chromatin accessibility reads and associated gene expression reads to a sample of origin based on the identified barcode.

22. The system of claim 21, wherein the multiplexed analysis comprises simultaneous single-nucleus RNA sequencing and single-nucleus Assay for Transposase-Accessible Chromatin using sequencing.

23. The system of claim 21 , wherein the instructions further cause the processor to identify and remove data corresponding to doublet events by detecting an association of a single cell barcode with more than one unique sample-specific barcode.

24. A method for multiplexed analysis of biological samples, comprising: a. means for tagging genomic DNA within nuclei of a plurality of distinct biological samples, wherein each sample is tagged with a unique barcode; b. pooling the nuclei after the tagging; and c. means for analyzing the pooled nuclei in a single reaction.

25. The method of claim 24, wherein the means for tagging comprises a plurality of transposome complexes, each complex comprising a transposase and an oligonucleotide construct having one of the unique barcodes.

Citation Information

Patent Citations

  • Single cell analysis of transposase accessible chromatin

    US20180340170A1

  • Methods and compositions for multiplexing single cell and single nuclei sequencing

    US20210171938A1

  • Multiplexed droplet-based sequencing using natural genetic barcodes

    WO2020123536A1

  • Method and kit for labeling nucleic acid molecules

    WO2022143221A1