Use of droplet single-cell epigenomic profiling for patient stratification
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HIFIBIO SAS
- Filing Date
- 2024-04-11
- Publication Date
- 2026-08-07
Smart Images

Figure 0007902217000001 
Figure 0007902217000002 
Figure 0007902217000003
Abstract
Description
Technical Field
[0001] The present invention is in the fields of molecular biology, drug resistance, and microfluidics. In particular, the present invention relates to methods for assaying nucleic acids within microfluidic droplets for the diagnosis and / or prediction of drug resistance and patient stratification. The present invention also encompasses nucleic acid sequences / structures for use in profiling the epigenetic and transcriptome states of single cells within biological samples.
Background Art
[0002] Eukaryotic genomes are organized into chromatin that not only packs DNA but also enables the regulation of DNA metabolism (replication, transcription, repair, recombination). Current challenges are therefore to understand (i) how functional chromatin domains are established in the nucleus, (ii) how chromatin structure / information is dynamic through assembly, disassembly, modification, and remodeling mechanisms, and (iii) how these events are involved / maintained in the establishment, progression, and recurrence of diseases. Understanding these events enables the identification of new mechanisms of disease progression and new therapeutic targets, as well as the control of the effects of therapeutic molecules.
[0003] Furthermore, a fundamental research challenge in biology is understanding how hundreds of distinct cell types arise from the same genetic material in multicellular organisms. Many different cell types cannot be explained by genetics alone, but rather by additional information that can bridge phenotype to genotype. In 1942, Conrad H. Waddington coined the term epigenetics as "the branch of biology that studies the causal interactions between genes and their phenotypic products." This additional layer of "epigenetic information" is conserved in the form of chemical modifications to both the DNA and histone proteins that make up chromatin. Epigenetic mechanisms through chromatin modification regulate gene expression, form specific chromatin landscapes, and thereby make it possible to predict cell type and tissue identity.
[0004] DNA and histone modifications are involved in various DNA-based processes by acting as recognition sites for effector proteins that can read information, and by stabilizing their binding to chromatin. A large number of histone modifications allow for strict control of chromatin structure and great flexibility in regulating DNA-based processes. This diversity leads to crosstalk between histones that can be modified simultaneously at different sites (Wang et al. 2008, Nature Genetics 40(7):897-903).
[0005] Histone modifications can positively or negatively influence each other. Furthermore, communication between histone modifications also exists between other chromatin modifications, such as DNA methylation, and all of them are involved in fine-tuning the overall regulation of biological function (Du et al. 2015, Nature Reviews Molecular Cell Biology, 16(9):519-532).
[0006] DNA and histone modifications contribute to defining epigenomic signatures within distinct chromatin states, and they highly indicate cell type and tissue identity. Genome-wide profiling of these marks can be used to understand the overall landscape of genomic regulation and, furthermore, to distinguish epigenomic differences in relation to normal and diseased cellular states (Consortium Epigenomics 2015, Nature 518(7539):317-329). However, the current state of chromatin profiling techniques does not allow for the study of cellular heterogeneity or the detection of intercellular variability in chromatin states.
[0007] To map epigenetic modifications, epigenetic markers / erasers, and factors that play a role in chromatin state genome-wide, 2D and 3D organization using traditional ChIP-seq methods requires a large number of cells to produce high-quality binding site profiles. Several studies have shown optimized ChIP-seq protocols that reduce input material from millions of cells to hundreds of cells without compromising resolution in detecting enriched or depleted regions (Adli et al. 2010, Nature Methods 7(8):615-618; Brind'Amour et al. 2015, Nature Communications 6:6033; Ma et al. 2018, Science Advances 4(4):eaar8187). However, these methods only produce averaged snapshots of modification states and do not provide insight into epigenetic heterogeneity.
[0008] Profiling histone modifications at single-cell resolution remains a challenge, partly because the level of noise associated with nonspecific binding during immunoprecipitation tends to increase with small amounts of starting material. While immunoprecipitation of chromatin from a single cell is technically feasible, it yields highly variable results.
[0009] Chromatin derived from isolated single cells can be pre-indexed with a unique and distinctive DNA sequence (barcode), and then combined with indexed chromatin from several to several thousand cells for bulk immunoprecipitation, similar to traditional ChIP-seq protocols. This method avoids the problems associated with large experimental noise in immunoprecipitation of small amounts of input material while preserving single-cell information. Indeed, since the barcode is unique to one cell, after sequencing, each read can be attributed to its cell of origin. However, as with other single-cell techniques involving molecular indexing, there is a potential for only indexed nucleosomes to be amplified and sequenced.
[0010] In this regard, Rotem developed the Drop-ChIP technique, which combines a chromatin indexing method with droplet-based microfluidics to profile histone modifications in thousands of cells (Rotem et al. 2015, Nature Biotechnol. 33(11):1165-1172). The droplet method provides a versatile tool for performing single-cell assays. Following the steps of compartmentalizing, lysing, and fragmenting chromatin with micrococcal nucleases within the droplet, the droplet is merged one-to-one with a second population of droplets containing DNA barcodes, enabling chromatin indexing at the single-cell level.
[0011] While Drop-ChIP has revealed distinct chromatin states within a population of embryonic stem cells, the single-cell information was limited to a few hundred unique enriched loci per cell due to the low effectiveness of chromatin indexing or the limited recovery of indexed nucleosomes. In particular, the Drop-ChIP technique suffers from two main limitations that can negatively impact the amount of information recovered per cell. First, only symmetrically indexed nucleosomes can be amplified and become part of the sequencing library. This requirement dramatically increases the rigor of the system, imposing a strong selection on nucleosomes (i.e., only those ligated with barcodes at both ends). Second, amplification of indexed nucleosomes relies solely on a very large number of polymerase chain reaction (PCR) cycles, thereby increasing the likelihood of amplification bias and errors.
[0012] Spontaneous, hereditary, or induced chromatin heterogeneity in untreated cells (or cells derived from untreated subjects) may be a crucial molecular component in the acquisition of drug resistance, regardless of the mechanism of action of cancer treatment. Many types of cancer are initially sensitive to chemotherapy but can become resistant over time through these and other mechanisms. However, while the methods of drug resistance may be disease-specific, others may be evolutionarily conserved. The emergence of resistance to treatment (including chemotherapy and targeted therapy) is a major challenge in the treatment of diseases, including cancer. Genetic heterogeneity within untreated tumors is now considered a critical determinant of resistance. Furthermore, non-hereditary, and in particular, transcriptional and epigenetic mechanisms are expected to play a role in the adaptation of cancer cells facing environmental, metabolic, or therapeutic relationship stresses (Rathert, P. et al. Nature 525, 543-547, (2015); Kim, C. et al. Cell 173, 879-893 e813, (2018)). While chromatin structure regulation via histone modification is a major epigenetic mechanism and regulator of gene expression, the contribution of chromatin features to tumor heterogeneity and evolution remains unknown. [Overview of the Initiative] [Problems that the invention aims to solve]
[0013] In one aspect of the present invention, it is disclosed that a rare population of cells within an untreated drug-sensitive tumor exhibits chromatin characteristics equivalent to those of resistant cells. The inventors have developed a droplet microfluidic technique for profiling the chromatin landscape of thousands of cells at single-cell resolution with coverage of up to 10,000 loci / cell.
[0014] Given the limitations described above that affect methods known in the art, it is clear that there is a need for improved methods for diagnosing and / or predicting emerging or current drug resistance to determine patient stratification, where drug resistance is associated with distinct chromatin states and the use of single-cell epigenetic profiling in microfluidic droplets is required. [Means for solving the problem]
[0015] One aspect of the present invention relates to a method for diagnosing and / or predicting drug resistance, wherein the chromatin state of a single cell is profiled in a cell obtained from a subject using a microfluidic system, and the method is as follows: a. A step of providing at least one droplet of a first type, Here, the first type of droplet is, i. Biological elements, ii. Dissolution buffer, and iii. Nuclease, including b. A step of collecting the first type of droplet under conditions that temporarily inactivate the nuclease, c. A step of incubating the first type of droplet described above, thereby reactivating the nuclease, d. Providing at least one droplet of a second type, wherein the droplet of the second type comprises a nucleic acid sequence. e. A step of merging the first and second types of droplets to generate a third type of droplet, f. Incubating the third type of droplet described above to thereby bind the nucleic acid sequence to one or more target genomic regions, g. Sequencing one or more of the target genomic regions mentioned above. Includes.
[0016] Another aspect of the present invention relates to a method for identifying one or more target genomic regions (s) using a microfluidic system, the method comprising the following steps: a. providing at least one droplet of a first type, wherein the droplet of the first type comprises i. a biological element, ii. a lysis buffer, and iii. a nuclease ; b. collecting the droplet of the first type under conditions that temporarily inactivate the nuclease, c. incubating the droplet of the first type to thereby reactivate the nuclease, d. providing at least one droplet of a second type, wherein the droplet of the second type comprises a nucleic acid sequence, e. merging the droplets of the first and second types to thereby generate a droplet of a third type, f. incubating the droplet of the third type to thereby bind the nucleic acid sequence to at least one target genomic region (s), g. sequencing the at least one target genomic region (s). ;
[0017] A further aspect of the present invention relates to a. at least one index sequence, b. a sequencing adapter, and c. at least one protecting function located at the 3'-end and / or 5'-end of a nucleic acid sequence;
[0018] A further aspect of the present invention relates to a method of using a nucleic acid sequence according to a particular embodiment of the present invention in the prediction, diagnosis or prediction / judgment of a subject, for example the drug resistance of a subject.
[0019] Another aspect of the present invention relates to a method of using nucleic acid sequences according to a particular embodiment of the present invention in profiling the epigenetic state of a sample obtained from a subject. Further aspects include the use of epigenetic state profiles or information in predicting a subject, for example, in diagnosing or predicting / determining any drug resistance of the subject. [Brief explanation of the drawing]
[0020] [Figure 1] The following describes a microfluidic workflow according to one aspect of the present invention: (A) Cells are compartmentalized in a 45 pl droplet with reagents necessary for lysis and chromatin fragmentation. In parallel, hydrogel beads containing DNA barcodes are encapsulated in a 100 pl droplet with ligation reagents. The two emulsions are reinjected into a fusion device to pair the barcode drop (100 pl) and the nucleosome drop (45 pl) asymmetrically, and fusion is triggered by an electric field. The fused droplets are scanned one by one with a laser beam to analyze the composition of each droplet in real time. (B) The emulsion from the fused droplets is collected for barcoding of nucleosomes in the droplets. The barcodes are released from the beads by photocutting and ligated to the nucleosomes. The contents of the droplets are combined and immunoprecipitated to sequence the enriched DNA. Deconvolution of barcode-related beads reconstructs the chromatin profile of a single cell by attributing all sequences to the cells from which they originated. [Figure 2]This shows the synchronization and quiescence of MNase activity between droplets. In particular, Figure 2a shows gel electrophoresis of DNA fragments derived from human Jurkat T cells at different time points during MNase incubation. At t=0 min, the DNA is not yet fragmented, and it is confirmed that MNase activity is synchronized when the droplets are collected on ice. At t=12 min + 1 hour (on ice), a digestion profile similar to that at 12 mins after incubation is shown, confirming that MNase activity is quiesced when the droplets are stored on ice. Figure 2b shows the quiescence of MNase activity in droplets starting from fixed nuclei. MNase is quiesced for 3 hours without overdigesting chromatin. [Figure 3] This shows the EGTA concentration required to completely inactivate MNase activity within the droplet. The fraction of oligonucleotides remaining after incubation was measured by TapeStation and standardized against a digestion negative control (i.e., a droplet without MNase). A final concentration of 26 mM EGTA completely inactivated the MNase within the droplet. The bar plot shows the mean fraction of undigested oligonucleotides for each duplicate, and the error bars represent the standard deviation. [Figure 4] The nucleic acid sequence according to the present invention is shown. The novel structure (v2) enables digestion of the barcode concatemer by adding half of the Pac1 restriction site to both ends of the barcode. A C3 spacer is also added to the 3' end of the lower strand to enforce the direction of ligation. Non-full-length barcodes are completed with protecting groups, including modifications with modified bases to prevent unwanted ligation. The table shows the percentage of correct barcodes identified after sequencing for each barcode structure. [Figure 5]This demonstrates quality control of barcoded hydrogel beads. (a) Tapestation profile of DNA barcodes after photocleavage from hydrogel beads shows the presence of full-length barcodes (larger peak at 146 bp) as well as incomplete intermediates (peaks at 72 bp, 94 bp, and 119 bp). (b) Imaging of hydrogel beads in droplets using epifluorescence microscopy. From left to right: (i) bright-field image; (ii) imaging of barcodes after hybridization of complementary DNA probes to the barcodes in the Illumina sequencing adapter; (iii) similar to (ii) after release of barcodes in droplets by photocleavage. Scale bar is 35 μm. (c) Results of deep sequencing of a single bead, showing the first and second most abundant barcode fractions of 16 beads. On average, 97.7% of the barcodes present on the beads fit into the same sequence, while the second most abundant barcodes represent only 0.17% of the total sequencing reads. [Figure 6] This report evaluates the ligation efficiency in droplets. Ligation was performed under 2-hour, 4-hour, and overnight incubations. Both measurement methods provided similar evaluations of the ligation product fraction, with a significant increase in efficiency observed after overnight incubation (approximately 10%). The bar plot shows the average fraction of ligated oligonucleotides for duplicates (experiments conducted on two different days), and the error bars represent the standard deviation. [Figure 7]This study presents a conceptual study and evidence of expected results. (a) Human B and T lymphocytes were separately encapsulated in droplets and specifically indexed. The indexed chromatins from the two emulsions were then combined for ChIP-seq targeting three distinct histone modifications (H3K4me3, H3K27ac, and H3K27me3). (b) Data analysis involved grouping cells into two clusters using an unsupervised clustering method. The correct clustering of B and T cells, based on their chromatin profiles, was then confirmed by cell type-specific barcode sequences. [Figure 8] This shows the monitoring of cell and hydrogel inclusion in droplets. (a) Experimental time trace recorded for droplets analyzed at 1.8 kHz. Orange fluorescence is present in all droplets and is used to control for drop size (drop-code). Green fluorescence indicates the presence of cells in the droplet (cell-code). Plot of cell-code intensity (green) against drop-code intensity (orange) for each droplet. Droplets containing cells have high green cell-code intensity, allowing for the counting of inclusion cells by defining a cell gate above the noise level. A cell density adjusted to λ=0.1 results in approximately 9% of droplets containing a single cell. (b) Experimental time trace recorded for 100 pl droplets analyzed at 650 Hz. Orange fluorescence is present in all droplets and is used as a drop-code to control for droplet size. Red fluorescence indicates the presence of hydrogel beads in the droplet (bead-code). Plot of bead-code strength (red) against drop-code strength (orange) for each droplet. By using close-packed ordering of beads, 65% to 75% of droplets contain a single bead. [Figure 9]This shows live monitoring of droplet fusion. (a) Experimental time trace recorded for fused droplets at 150 Hz. Orange fluorescence is present in all droplets and is used as a drop-code to control the size of the fused drops. Green fluorescence indicates the presence of cells, and red fluorescence indicates the presence of hydrogel beads. Blue fluorescence is a drop-code specific to the cell emulsion. (b) Plot of the intensity of the drop-code "cell" against the drop-code intensity in each droplet, defining four main droplet populations after fusion. The central main population represents correctly paired and fused droplets (70%~80%). Unpaired droplets from the bead-emulsion are in the lower population, and unpaired droplets from the cell-emulsion with high blue fluorescence intensity are in the upper left population. The last population (upper right) is associated with improperly paired droplets, including two cell-droplets fused with one bead-droplet. (c) A plot of cell-coding intensity against bead-coding intensity in each droplet allows for an accurate count of available drops (those containing one cell and one bead). Droplets derived from a parallel time trace in (a) are shown as examples of different populations. [Figure 10] This shows the total number of cells and the number of cells encapsulated with barcoded hydrogel beads, as detected by fluorescence on the microfluidic station, in single-cell ChIP-seq experiments with H3K4me3 and H3K27me3. Analysis of the sequencing data revealed a number of identified barcodes closely correlated with the number of droplets containing both cells and beads counted on the microfluidic station, demonstrating the high overall effectiveness of this system. [Figure 11]This shows the sensitivity of the scChIP-seq procedure for identifying subpopulations. The t-SNE plot shows the H3K27me3 scChIP-seq dataset in an in silico simulation of the detection limit, where the ratio of spiked-in B cells within the T cell population is varied (from top to bottom), and the threshold for uniquely mapped reads per barcode is varied (from left to right). Points are colored according to cell type-specific barcode sequences. [Figure 12] The performance of sequencing in Drop-ChIP compared to the inventors' procedure is shown. Table 1 compares the expected number of cells per sequencing library, the number of live sequencing reads, and the average number of live reads per cell in DropChIP and the inventors' scChIP-seq system. Table 2 compares the number of cells identified after sequencing, the final number of cells used in post-QC analysis, and the average number of usable reads per cell after QC in DropChIP and the inventors' scChIP-seq system. [Figure 13] The mixing of human and mouse cells confirms single-cell resolution. (a) A scatter plot of the number of reads per barcode aligning to the mouse versus human reference genome shows that 96.5% of the barcodes are specific to one species (at least 95% of reads with the same barcode map to one of the two species). The species percentages for mouse (26.4%), human (70.1%), and mixed (3.5%) were close to the expected values based on the Poisson distribution of cells in a droplet (32.6%, 65.2%, and 2.2%, respectively) with an average number of cells per droplet λ=0.1. (b) Bar plot showing the number of barcodes identified for each type (light gray to dark gray-black bars) (corresponding to the number of cells) compared to the expected number of cells counted on the microfluidic station (gray bars; 3,000 in total, derived from a mixture containing 1 / 3 mouse cells and 2 / 3 human cells). [Figure 14]Clustering single-cell ChIP-seq data reveals cell type-specific biological similarities. (a) Histogram of the distribution of raw and intrinsic sequencing reads per barcode (i.e., per cell) for scChIP-seq in single-cell ChIP-seq datasets of H3K4me3 and H3K27me3. (b) Density scatter plot showing log2 cumulative counts for three independent fractions of identical emulsions of B cells recovered and processed in parallel to generate the H3K4me3 single-cell ChIP-seq dataset. Correlations between replicates are calculated based on cumulative counts per million reads in a 5kb genome bin across single cells. Pearson correlation scores and p-values are calculated genome-wide. (c) Density scatter plot for two biological replicates corresponding to two emulsions of B cells recovered from different cell culture flasks and processed with different batches of barcoded hydrogel beads to generate the H3K27me3 single-cell ChIP-seq dataset. Correlation between replicates is calculated based on the cumulative count per million reads in a 50kb genome bin across a single cell. Pearson correlation scores and p-values are calculated genome-wide. (d) A t-SNE plot showing H3K27me3 scChIP-seq data from two biological replicates, colored according to the originating batch (left) or consensus clustering result (right), to discuss batch effects of cell population clustering. (e) Left panel: Hierarchical clustering of intercellular Pearson correlation scores and corresponding heatmap for H3K27me3 scChIP-seq data from a 1:1 mixed population of human B cells and T cells together with separately barcoded B cells and T cells. Unique read counts, originating batches, and consensus clustering results are shown above the heatmap. Right panel: Corresponding t-SNE plot, with points from the mixed population colored gray and points from separately barcoded B cells and T cells colored with different gray shades according to cell type-specific barcode sequences.(f) A Venn diagram comparing H3K4me3 peaks detected by single-cell and bulk methods for T cell and B cell datasets. [Figure 15] This shows the reconstruction of cell type-specific chromatin states from single-cell ChIP-seq profiles. (a) A t-SNE plot showing scChIP-seq datasets of H3K4me3 and H3K27me3 from human B lymphocytes and T lymphocytes, separately indexed in droplets using hydrogel beads containing both single-cell and cell type-specific barcodes and then mixed for immunoprecipitation. Points are colored according to the cell type-specific barcode sequence. Accuracy is shown as the consensus clustering of scChIP-seq data (Figure 16a) and agreement between classifications by known cell identity, assessed by cell type-specific barcodes. (b) Snapshots of differentially enriched loci (Figure 16b) using cumulative single-cell and bulk profiles for each cell type. Differentially bound regions, identified by the Wilcoxon signed-rank test, are shown in gray along with the corresponding adjusted p-values and log2 factor changes (fold changes). (c) Scatter plot showing enrichment of log2 RPM (read count per million mapped reads) in cumulative single-cell vs. bulk ChIP-seq data, calculated within a 5kb genome bin for H3K4me3 and within a 50kb bin for H3K27me3. Pearson correlation scores and p-values are calculated at the genome level. [Figure 16]Single-cell ChIP-seq data demonstrate the distinction between human T cells (Jurkat) and human B cells (Ramos). (a) Consensus clustering matrix for the scChIP-seq datasets of H3K4me3 (upper panel) and H3K27me3 (lower panel). Consensus scores range from 0 (white: never clustered together) to 1 (dark gray: always clustered together). (b) Volcano plot showing adjusted p-values (Wilcoxon rank test) for scaling changes (thresholds of 0.01 for q-value and 1 for |log2FC|) for differential analysis comparing chromatin features between B cells and T cells in the scChIP-seq datasets of H3K4me3 (upper panel) and H3K27me3 (lower panel). (c) Bar plot showing -log10 adjusted p-values from pathway analysis in the H3K4me3 scChIP-seq dataset. The top 10 most important gene sets are shown below the bar plot. [Modes for carrying out the invention]
[0021] The inventors have developed an improved single-cell ChIP method based on droplet microfluidics, which results in a 5- to 10-fold increase in the number of enriched loci per individual cell compared to the Drop-ChIP technique disclosed in Rotem (see Figure 7). This method enables the evaluation of histone modifications, DNA-modified bases (including modified nucleotides for identifying ongoing DNA replication events at the single-cell or arbitrary biological element level), and chromatin / DNA-related factors with high sensitivity and accuracy at the single-cell level. This method is applicable to the identification of cell populations or arbitrary biological elements with distinguishing features, which include histone and / or DNA modifications and the presence of factors. The presence or absence of these elements may consequently indicate variability in gene expression and can therefore be used as biomarkers to revert the changes, therapeutic targets.
[0022] It is well understood that a cell may mean the nucleus as a compartment of chromatin structure. A cell or nucleus or any biological element may be an immobilized biological element. Examples of fixatives include aldehydes (not limited to formaldehyde and paraformaldehyde), alcohols (not limited to ethanol and methanol), oxidizing agents, mercury, picric acid, and the Hepes-glutamic acid buffer-mediated organic solvent protection effect (HOPE) fixative.
[0023] As in Drop-ChIP (Rotem et al. 2015, Nature Biotechnol. 33(11):1165-1172), droplets containing cells and droplets containing barcodes are produced separately before being reinjected into a dedicated microfluidic fusion device and merged one-to-one (see Figure 1). However, this method differs from Rotem in at least two ways that characterize the barcode strategy. First, the inventors replaced the soluble barcode emulsified from a microtiter plate containing oligonucleotides with hydrogel beads (or any solid support) containing millions of unique or most abundant DNA sequences. Second, the novel design of the barcode structure according to one aspect of the invention allows for linear amplification of all barcoded nucleosomes, not just nucleosomes barcoded symmetrically at both ends, as in Rotem. Third, the barcode design includes further features that increase the efficiency of barcoding the target nucleic acid. These features include the addition of protective portions against “incomplete barcodes” that prevent adhesion to the target nucleic acid. In another embodiment, the barcode may contain protective bases on the full-length oligonucleotide (protective bases include, but are not limited to, phosphorothioates, LNA / BNA, nucleotide phosphoramidites, synthetic cycles, non-3'OH or 5'P bases, and 2'-O-methyl-DNA / RNA), where these protective bases protect the full-length barcode, but the non-full-length barcode may be digested by exonucleases.
[0024] An additional set of barcodes, called "experimental barcodes," can be added to multiplex different experiments within a single immunoprecipitation reaction. Subsequent bioinformatics analysis allows for the demultiplexing of experimental states based on the sequence of "experimental barcodes."
[0025] It is well understood that barcodes are nucleic acid sequences capable of distinguishing feature details from nucleic acids originating from one compartment to another. The production of these barcodes is known to those skilled in the art and can exhibit random sequences (with or without known sequences on both sides) or be generated by split-pool synthesis (Klein et al., Cell, 2015).
[0026] Furthermore, a method according to one aspect of the present invention is characterized by a synchronization / pause step, thereby limiting intercellular variability in chromatin digestion between droplets.
[0027] The aforementioned advantages are disclosed hereafter in this specification in embodiments and features that characterize one aspect of the present invention. An embodiment of the present invention is provided in the examples and drawings.
[0028] In one aspect of the present invention, a method is provided for identifying one or more target genomic regions using a microfluidic system, the method comprising the following steps: a. A step of providing at least one droplet of a first type, Here, the first type of droplet is, i. Biological elements, ii. Dissolution buffer, and iii. Nuclease, including b. A step of collecting the first type of droplet under conditions that temporarily inactivate the nuclease, c. A step of incubating the first type of droplet described above, thereby reactivating the nuclease, d. Providing at least one droplet of a second type, wherein the droplet of the second type comprises a nucleic acid sequence. e. A step of merging the first and second types of droplets to generate a third type of droplet, f. Incubating the third type of droplet described above to thereby bind the nucleic acid sequence to at least one(s) desired genomic region. g. Sequencing one or more of the target genomic regions mentioned above. Includes.
[0029] A method according to one aspect of the present invention is carried out in a microfluidic system. In relation to one aspect of the present invention, the term “microfluidic system” generally refers to a system or device having one or more channels and / or chambers manufactured on a micron or submicron scale.
[0030] A method according to one aspect of the present invention is characterized by the presence of first, second, and third types of droplets. The terms “first,” “second,” and “third,” as used herein in relation to droplets, are used to distinguish droplets according to their contents. Since the method is carried out in a microfluidic system, the term “droplet” also refers to a “microfluidic droplet.” Thus, in the context of a microfluidic system, the term “droplet” also refers to an isolated portion of a first fluid surrounded by a second fluid, where the first and second fluids are immiscible.
[0031] According to a step or stage of a method in one aspect of the present invention, the droplet may be contained within a microfluidic system (on-chip) or within a recovery device located away from the microfluidic system (off-chip). The droplet may have a spherical or non-spherical shape.
[0032] In one embodiment of one aspect of the present invention, the droplet has a volume in the range of about 20 pl to about 100 pl. Preferably, the droplet has a volume in the range of about 30 pl to about 70 pl. More preferably, the droplet has a volume in the range of about 40 pl to about 50 pl. Ideally, the droplet has a volume of about 45 pl. As used herein, the term "about" refers to a range of ±10% of a specified value.
[0033] As used herein, the term "lysis buffer" refers to a buffer capable of lysing biological cells. The meaning of the term "lysis buffer" is within the scope of common general knowledge for those skilled in the art.
[0034] As used herein, the term "genome region" refers to a nucleic acid sequence encoded by DNA or RNA.
[0035] As used herein, the term "nuclease" refers to an enzymatic reagent capable of cleaving phosphodiester bonds connecting nucleotide residues within a nucleic acid molecule. Nucleases can digest double-stranded, single-stranded, cyclic, and linear nucleic acid molecules. In connection with one aspect of the present invention, the nuclease may be an endonuclease that cleaves phosphodiester bonds within a polynucleotide chain, or an exonuclease that cleaves phosphodiester bonds at the terminus of a polynucleotide chain, or a transposase. The nuclease may also be a site-specific nuclease that cleaves specific phosphodiester bonds within a specific nucleotide sequence, such as a recognition sequence. A non-limiting example of a nuclease is a micrococcal nuclease (MNase). In certain embodiments, the nuclease is a micrococcal nuclease (MNase).
[0036] As used herein, the term “biological element” may refer to a single cell, nucleus, or nucleic acid-containing organelle (e.g., mitochondria), and may be obtained from living organisms, humans, or non-human subjects. In the latter case, the non-human subject is not limited to mammals.
[0037] Performing enzymatic assays on individual biological elements within droplets presents a challenge because cells are processed sequentially on different timescales. For example, the encapsulation step of a cell or any biological element lasts approximately 20 minutes, which is within the same order of magnitude as the incubation step. Thus, cells or any biological elements encapsulated first in the droplet will have longer contact with the nuclease than cells or any biological elements encapsulated last in the production. Similar observations can be made with respect to the reinjection of droplets within a fusion device (see general scheme in Figure 1). In fact, the fusion of two emulsions can last from 1 to 4 hours depending on the experimental design, meaning that some droplets containing fragmented DNA will "wait" for hours before fusion and inactivation of their MNases by EGTA. Consequently, synchronizing and pausing enzymatic activity is crucial to avoid introducing variations in chromatin digestion between individual cells or any biological elements.
[0038] In particular, in conventional bulk ChIP-seq assays, nuclease inactivation occurs immediately after nuclease incubation by adding EGTA. In contrast, in single-cell ChIP-seq assays, EGTA cannot be added immediately to the droplet, and the nuclease is inactivated only after fusion with the droplet containing the barcode.
[0039] To control and limit variations in chromatin digestion between cells or any biological element, the inventors introduced a step of recovering a first type of droplet under conditions that temporarily inactivate the nuclease. The recovery step, aimed at synchronizing / quitting nuclease activity within the droplet, is performed before each incubation step. The inventors identified that the droplet compartment can make the MNase enzyme sensitive to temperature changes, allowing for selective inhibition / reactivation and reinhibition of enzyme activity. Such precise control of nuclease activity is not possible in bulk. This effect is thought to depend not only on MNase activity but also on any enzyme.
[0040] Therefore, according to another embodiment, the method further includes, prior to step (e), a step of collecting the first type of droplet under conditions that temporarily inactivate the nuclease.
[0041] In yet another embodiment, the condition of step (b) includes the step of selecting a temperature in the range of -20°C to 10°C, and the condition of step (c) includes the step of selecting a temperature in the range of 20°C to 40°C.
[0042] The droplets can be incubated outside the microfluidic system (off-chip) for fragmentation of single-cell chromatin. Since lysis occurs within the droplet, the nuclear DNA derived from the lysed cells is accessible to nuclease enzymes. Therefore, the kinetics of digestion are particularly important to preferentially yield mononucleosomes, which are retained within the droplet.
[0043] In certain embodiments, incubation step (c) is timed to combine the nuclear DNA fragment into a mononucleosome.
[0044] In yet another embodiment, one or more target genomic regions (one or more) include one or more modified genomic regions (one or more).
[0045] In yet another embodiment, one or more target genomic regions are modified genomic regions.
[0046] According to the present invention, a modified genomic region comprises a nucleic acid sequence-associated protein complex and / or a nucleic acid sequence. In certain embodiments, the modified genomic region is a modified mononucleosome. In other embodiments, the modified genomic region is a transcription factor binding site, a chromatin modifier binding site, a chromatin remodeler site, or a histone chaperone binding site.
[0047] According to the present invention, the modified genomic region may include post-translational modifications selected from the group including acetylation, amidation, deamidation, carboxylation, disulfide bonding, formylation, glycosylation, hydroxylation, methylation, myristoylation, nitrosylation, succinylation, butyrylation, phosphorylation, prenylation, ribosylation, sulfation, SUMOylation, ubiquitination, and derivatives thereof.
[0048] According to the present invention, the modified genomic region may include histone variants selected from the group including CENP-A / CID / cse4 (centromere epigenetic marker), H3.3 (transcription), H2A.Z / H2AV (transcription / double-strand break repair), H2A.X (double-strand break repair / sex chromosome meiotic rearrangement), macroH2A (gene silencing / X chromosome inactivation), H2A.Bbd (active chromatin epigenetic marker), H3.Z (regulation of cellular response to external stimuli), and H3.Y (regulation of cellular response to external stimuli).
[0049] According to the present invention, the modified genomic region may include a modified DNA sequence selected from the group comprising methylation and its derivatives, modified nucleotides such as EdU, BrdU, IdU, CldU, and others. The most common method of modifying a base is the addition of a methyl mark, and methylation has been found on cytosine and adenine across various types, resulting in 5mC, N4-methylcytosine (N4mC), or 6-methyladenine (6mA), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC).
[0050] As mentioned above, a notable limitation of the method disclosed in Rotem is that only symmetrically indexed nucleosomes can be amplified and become part of the sequencing library. This requirement dramatically increases the rigor of the system, imposing a strong selection on nucleosomes, thereby limiting them to those ligated with barcodes at both ends. In contrast, with the Drop-ChIP method, the inventors have surprisingly found that by indexing nucleosomes at only one end, single-cell coverage is increased, ultimately increasing the potential of the system to distinguish more subtle variations between single-cell chromatin profiles.
[0051] In yet another embodiment, the nucleic acid sequence is asymmetrically bound to at least one or more of the target genomic regions.
[0052] As used herein, the term "asymmetrically bound" means the presence of at least one barcode bound to the genomic region of interest, such that the binding is to only one of the two ends of the genomic region of interest.
[0053] In a further aspect of the present invention, a. At least one index array, b. Sequencing adapter, and c. At least one protective feature located at the 3' end and / or 5' end. A nucleic acid sequence containing is provided.
[0054] As used herein, the term “nucleic acid sequence” refers to a single-stranded or double-stranded nucleic acid. In yet another embodiment, the “nucleic acid sequence” may be DNA or RNA. In a preferred embodiment, the “nucleic acid sequence” is double-stranded DNA. In some embodiments, the “nucleic acid sequence” includes double-stranded DNA containing a first-strand barcode and a second-strand barcode. In some embodiments, the first-strand barcode and the second-strand barcode include complementary sequences. In some embodiments, the first-strand barcode and the second-strand barcode include non-complementary sequences.
[0055] As used herein, the term “index sequence” refers to a unique nucleotide sequence that is distinguishable from any other index sequences and any other nucleotide sequences within the nucleic acid sequence in which it is contained. An “index sequence” may be a random or specially designed nucleotide sequence. An “index sequence” may have any sequence length. Nucleic acid sequences according to further embodiments of the present invention may be appended to a target genomic region to tag a species, requiring the identification and / or distinction of different members of the tagged species within that population. Therefore, the terms “index sequence” and “barcode” in relation to one embodiment of the present invention may be used interchangeably.
[0056] As used herein, the term “sequencing adapter” refers to an oligonucleotide of a known sequence, whose ligation or incorporation into a polynucleotide or polynucleotide chain of interest enables the generation of a product of the said polynucleotide or polynucleotide chain of interest that is ready for amplification.
[0057] In one embodiment of a further aspect of the present invention, the nucleic acid sequence further comprises at least one cleavage site.
[0058] As used herein, the term “cleavage site” refers to a target region of a nucleic acid sequence that can be cleaved by any means including an enzyme capable of cleaving single-stranded or double-stranded nucleic acid sequences, but is not limited to those specified. In connection with one aspect of the present invention, the “cleavage site” may play a role in cleaving or otherwise releasing a portion of the nucleic acid sequence. The “cleavage site” is recognized by a cleavage reagent and may be natural, synthetic, unmodified, or modified.
[0059] In one embodiment of the present invention, the protective function is selected from the group comprising a 3'-terminated space element and a 5'-terminated dideoxy-modified base. In connection with one aspect of the present invention, a suitable non-limiting space element is a three-carbon spacer (C3 spacer).
[0060] In yet another embodiment of the present invention, at least one cutting region is a limiting region that includes a palindromic structure region.
[0061] As used herein, the term "restriction site" refers to a site recognized by a restriction enzyme, such as an endonuclease. Those skilled in the art are familiar with restriction endonucleases and their restriction sites. Non-limiting examples of restriction sites include BamHI, BsrI, NotI, XmaI, PspAI, DpnI, MboI, MnlI, Eco57I, Ksp632I, DraIII, AhaII, SmaI, MluI, HpaI, ApaI, BclI, BstEII, TaqI, EcoRI, SacI, HindII, HaeII, DraII, Tsp509I, Sau3AI, and PacI.
[0062] In yet another embodiment of the present invention, the nucleic acid sequence is suitable for use in the methods according to the first aspect and embodiments of the present invention.
[0063] In another embodiment of the present invention, nucleic acid sequences are suitable for use in profiling the epigenetic state in a sample obtained from a subject.
[0064] As used herein, the term "sample" refers to a biological sample.
[0065] As used herein, the term "subject" refers to either a human or a non-human subject. In the latter case, the non-human subject is not limited to mammals.
[0066] Methods according to one aspect of the present invention can find different applications in the identification of genes and factors involved in the diagnosis and / or prediction of pathological conditions in a subject, as well as in methods for diagnosing and / or predicting pathological conditions in a subject and methods for controlling the effects of therapeutic molecules on chromatin.
[0067] In connection with the present invention, pathological conditions may refer to any modifications including those of nucleosomes or nucleic acid sequences, and the arrangement of proteins that affect the structure, regulation, and function of chromatin. As used herein, the term “pathological conditions” also encompasses abnormal rates of cell proliferation such that disease treatment requires regulation of the cell cycle. Examples of proliferative disorders include, but are not limited to, cancer.
[0068] A method according to one aspect of the present invention can be used for in vitro diagnosis and / or prediction of drug resistance, wherein the chromatin state of a single cell is profiled in a cell obtained from a subject using a microfluidic system, and the method is as follows: a. A step of providing at least one droplet of a first type, Here, the first type of droplet is, a. Biological elements, b. Dissolution buffer, and c. Nuclease, including b. A step of collecting the first type of droplet under conditions that temporarily inactivate the nuclease, c. A step of incubating the first type of droplet described above, thereby reactivating the nuclease, d. Providing at least one droplet of a second type, wherein the droplet of the second type comprises a nucleic acid sequence. e. A step of merging the first and second types of droplets to generate a third type of droplet, f. Incubating the third type of droplet described above to thereby bind the nucleic acid sequence to one or more target genomic regions, g. Sequencing one or more of the target genomic regions mentioned above. Includes.
[0069] Drug resistance may be emerging drug resistance and / or existing drug resistance. The emergence of drug resistance may be due to epigenetic heterogeneity.
[0070] Single-cell chromatin profiling is considered a unique tool for investigating the heterogeneity and dynamics of chromatin states within any complex biological system and can be applied to other diseases, particularly autoimmune diseases, infectious diseases, metabolic diseases, and healthy systems, in addition to cancer, to study cell differentiation and development, and immune surveillance.
[0071] A method according to one aspect of the present invention can be used to determine patient stratification, where drug resistance is associated with a distinct chromatin state and requires the use of single-cell epigenetic profiling in microfluidic droplets.
[0072] Embodiments of the present invention provide a method for diagnosing and / or predicting drug resistance in subjects who are in a pathological condition and / or suspected to be in a pathological condition.
[0073] Embodiments of the present invention provide a method for diagnosing and / or predicting drug resistance in healthy subjects.
[0074] According to the present invention, the diagnosis and / or prediction of drug resistance in a subject may be made before, during, or after the subject receives a procedure or treatment. The diagnosis and / or prediction may be made at any other point in time. The procedure or treatment may be a chemotherapeutic drug, a chemical drug, or a biologic drug, such as antibodies (and their derivatives or fragments) including, anti-immune checkpoint therapy, chemokines, hormones, cytokines (and their derivatives), or cell therapy consisting of TILs (tumor-infiltrating T cells) injection, CAR T cells (chimeric-associated antigens), CAR NK cells, TCR therapy (soluble or cellular therapeutic forms), vaccination (cancer vaccines, viral vaccines, dendritic cell therapy to induce vaccination), oncolytic viruses, nanoparticles, etc.
[0075] Embodiments of the present invention provide a method for diagnosing and / or predicting subjects exhibiting drug resistance and / or subjects suspected of having drug resistance.
[0076] The subjects may be subjects in a pathological condition and / or suspected of having a pathological condition, or healthy subjects.
[0077] As used herein, the term “diagnosis” refers to determining whether a subject exhibits or is likely to develop drug resistance. As used herein, the term “diagnosis” refers to a method by which a person skilled in the art can evaluate and / or determine the probability ("likelihood") of a subject having drug resistance, and / or progressing further, to therapeutic agents, chemotherapeutic agents, chemical agents, or biologic agents, such as antibodies (and their derivatives or fragments), including, anti-immune checkpoint therapy, chemokines, hormones, cytokines (and their derivatives), or cell therapies such as TILs (tumor-infiltrating T cell) injection, CAR T cell (chimeric-associated antigen), CAR NK cell, TCR therapy (soluble or cellular therapeutic forms), vaccines (cancer vaccines, viral vaccines, dendritic cell therapy that induces vaccination), oncolytic viruses, nanoparticles, etc. In the present invention, “diagnosis” includes the use of assays, most preferably scChIP results.
[0078] As used herein, the term “prediction” refers to the prediction of the likelihood of death or progression of a disease such as cancer due to drug resistance, and includes recurrent and metastatic spread, inflammation, infectious diseases, autoimmune diseases, metabolic diseases, genetic diseases, and non-genetic diseases.
[0079] The method according to the present invention may use single cells derived from a body sample.
[0080] In embodiments of the present invention, the body sample is fluid and / or solid. The body sample used herein may be derived from tissue, blood, serum, plasma, saliva, excrement, urine, breast, lung, colon, gastrointestinal tract, brain, kidney, or any other body sample.
[0081] According to the present invention, one or more target genomic regions are modified genomic regions. A modified genomic region comprises a nucleic acid sequence and / or a protein complex associated with the nucleic acid sequence. A modified genomic region comprises post-translational modifications selected from the group including acetylation, amidation, deamidation, carboxylation, disulfide bonding, formylation, glycosylation, hydroxylation, methylation, myristoylation, nitrosylation, phosphorylation, prenylation, ribosylation, sulfation, SUMOylation, ubiquitination, and derivatives thereof.
[0082] Furthermore, the cells are derived from subjects in a pathological condition and / or suspected of being in a pathological condition, or from healthy subjects. In one embodiment of the present invention, the cells are untreated and / or treated, or the cells are derived from untreated or treated subjects.
[0083] In one embodiment, treatment cells derived from the target of treatment are treated with a chemotherapy drug, a chemical drug, or a biological drug.
[0084] In a preferred embodiment, cells are treated with chemotherapeutic drugs, chemical drugs, or biopharmaceutical drugs, such as antibodies (and their derivatives or fragments), anti-immune checkpoint therapy, chemokines, hormones, cytokines (and their derivatives), or cell therapies such as TILs (tumor-infiltrating T cells) injection, CAR T cells (chimeric-associated antigens), CAR NK cells, TCR therapy (a therapeutic form of soluble or cellular therapy), vaccines (cancer vaccines, viral vaccines, dendritic cell therapy that induces vaccination), oncolytic viruses, nanoparticles, etc. In a more preferred embodiment, cells are treated with chemotherapeutic drugs.
[0085] In one embodiment of the present invention, the cells are untreated or derived from an untreated subject.
[0086] As used herein, "pathological condition" refers to diseases such as cancer, or infectious diseases, autoimmune diseases, metabolic diseases, inflammatory diseases, genetic diseases, and non-genetic diseases.
[0087] In one embodiment of the present invention, the target pathological conditions include cancer, infectious diseases, autoimmune diseases, inflammatory diseases, metabolic diseases, genetic diseases, and non-genetic diseases.
[0088] In one embodiment of the present invention, the pathological condition includes cancer at any stage from no detection to stage IV. In one embodiment of the present invention, cancer includes any type of cancer, such as solid tumors and / or humoral tumors.
[0089] In a preferred embodiment of the present invention, the target pathological condition is breast cancer.
[0090] According to the present invention, the subject may be male or female; in a preferred embodiment of the present invention, the subject is female.
[0091] In one aspect of the present invention, the chromatin state of a single cell is devoid of chromatin marks relating to genes that promote drug resistance. The chromatin marks include distinct histone modifications H3K4me3, H3K27ac, and H3K27me3. The H3K4me3 mark is expected to allow the gene to be activated when needed rather than permanently silenced. The H3K27me3 mark is expected to silence the gene. The H3K27ac mark in the gene enhancer is expected to promote gene activation.
[0092] In one embodiment, the chromatin state of a single cell is devoid of chromatin marks related to genes that may lead to drug resistance. These chromatin marks include histone modifications H3K4me3 and H3K27me3.
[0093] As used herein, “chromatin marks,” “genetic marks,” or “marks” refer to DNA and histone modifications and / or variants that contribute to defining the epigenomic signature within distinct chromatin states, which highly indicate cell type and tissue identity. Genome-wide profiling of these marks can be utilized to understand the overall landscape of genomic regulation and, further, to distinguish epigenomic differences in relation to normal and diseased cellular states, for example. Those skilled in the art are familiar with the various chromatin marks or other studies disclosed, for example (Consortium Epigenomics 2015, Nature 518(7539):317-329).
[0094] In one embodiment, the chromatin state of a single cell acquires a mark, which has a desilencing effect.
[0095] As used herein, the terms “cancer” and “tumor” are interchangeable and refer to malignant neoplasms. Examples of malignant neoplasms include solid and hematological tumors. Solid tumors are exemplified by tumors of the breast, bladder, bone, brain, central and peripheral nervous systems, colon, endocrine glands (e.g., thyroid and adrenal cortex), esophagus, endometrium, germ cells, head and neck, kidneys, liver, lungs, larynx and hypopharynx, mesothelioma, ovaries, pancreas, prostate, rectum, kidneys, small intestine, soft tissue, testes, stomach, skin, ureters, vagina and vulva. Malignant neoplasms include hereditary cancers, exemplified by retinoblastoma and Wilms' tumor. Furthermore, malignant neoplasms include primary tumors in the organs mentioned above and corresponding secondary tumors in distant organs ("tumor metastases"). Hematological malignancies are exemplified by invasive and slow-chronic forms of leukemia and lymphoma, namely non-Hodgkin's disease, chronic and acute myeloid leukemia (CML / AML), acute lymphoblastic leukemia (ALL), Hodgkin's disease, multiple myeloma, and T-cell lymphoma. Myelodysplastic syndromes, plasma cell neoplasms, paraneoplastic syndromes, cancers of unknown primary sites, and AIDS-related malignancies are also included.
[0096] Cancer is typically labeled into stages I through IV to determine the extent of its progression in terms of proliferation and spread at the time of diagnosis. In stage I, the cancer is localized to a part of the body and can be removed surgically. In stages II and III, the cancer is locally advanced and can be treated with chemotherapy, radiation, or surgery. In stage IV, the cancer has metastasized or spread to other organs and can be treated with chemotherapy, radiation, or surgery (stage V is used only in patients affected by Wilms' tumor where both kidneys are affected).
[0097] As used herein, “metabolic disease” includes, but is not limited to, metabolic syndrome X, congenital metabolic anomalies, mitochondrial diseases, phosphate metabolism disorders, porphyria, proteostatic defects, metabolic skin diseases, wasting syndromes, water and electrolyte imbalances, metabolic brain diseases, calcium metabolism disorders, DNA repair defects, iron metabolism disorders, lipid metabolism disorders, and malabsorption syndromes.
[0098] The term "autoimmune disease" as used herein is not limited to, but includes, multiple sclerosis, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid antibody syndrome, autoimmune angioedema, autoimmune autonomic neuropathy, autoimmune encephalomyelitis, autoimmune hepatitis, polyarteritis nodosa, polyglandular syndrome type I, II, III, polymyalgia rheumatica, polymyositis, sleep attacks, pyoderma gangrenosum, Raynaud's phenomenon, interstitial cystitis (IC), juvenile arthritis, and juvenile diabetes mellitus (type 1 diabetes). This includes diseases such as juvenile myositis (JM), reactive arthritis, neonatal lupus, neuromyelitis optica, celiac disease, Chagas disease, primary biliary cirrhosis, primary sclerosing cholangitis, progesterone-induced dermatitis, psoriasis, psoriatic chronic inflammatory demyelinating polyneuropathy (CIDP), transverse myelitis, type 1 diabetes mellitus, ulcerative colitis (UC), chronic relapsing multifocal osteomyelitis (CRMO), Churg-Strauss syndrome (CSS) or eosinophilic granulomatosis (EGPA), neutropenia, ocular scarring pemphigoid, optic neuritis, and relapsing rheumatoid arthritis.
[0099] As used herein, “hereditary disorders” include, but are not limited to, severe combined immunodeficiency (SCID), sickle cell disease, skin cancer, Wilson’s disease, Turner syndrome, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, myotonic dystrophy, neurofibromatosis, Noonan syndrome, Darkham disease, Down syndrome, Duane’s retraction syndrome, Duchenne muscular dystrophy, factor V Leiden thrombosis, autism, autosomal dominant polycystic kidney disease, breast cancer, and Charcot-Marie-Tooth disease.
[0100] As used herein, “inflammatory disease” includes, but is not limited to, allergies, asthma, autoimmune diseases, celiac disease, glomerulonephritis, hepatitis, and inflammatory bowel disease. [Examples]
[0101] Microfluidic workflow for single-cell ChIP-seq procedures Figure 1 shows the overall scheme of the microfluidic method according to the present invention. (a) Cells are compartmentalized in a 45 pl droplet with reagents necessary for cell lysis and chromatin fragmentation. In parallel, hydrogel beads containing DNA barcodes are encapsulated in a 100 pl droplet with ligation reagents and EGTA to inactivate MNases. The two emulsions are reinjected into a fusion device to asymmetrically pair the droplet containing the barcodes (100 pl) and the droplet containing nucleosomes (45 pl), and merged by electrocoalescence triggered by an electric field. The fusion droplets are scanned one by one with a laser beam to analyze the composition of each droplet in real time. (b) The emulsions of the fusion droplets are collected and incubated off-chip for barcoding of nucleosomes in the droplets. The barcodes are released from the beads by photocutting and ligated to the nucleosomes. The contents of the droplets are combined and immunoprecipitated to sequence the enriched DNA. Deconvolution of barcode-related reads assigns all sequences to their originating cells in order to reconstruct a single-cell chromatin profile.
[0102] Synchronization and cessation of chromatin fragmentation in droplets Cells are compartmentalized in 45 pl droplets with a digestion mix containing lysis buffer and MNase (see Figure 1). After complete lysis, chromatin is released into the droplet to facilitate access for cleavage by MNase. This section demonstrates a typical calibration of MNase activity in the droplet to preferentially produce nucleosome-sized fragments. However, performing enzyme assays in droplets is challenging because cells are processed individually at different timescales. Consequently, fine-tuning of enzyme activity is necessary to avoid discrepancies in chromatin digestion between droplets and between single cells.
[0103] Cellular compartmentalization in droplets The number of cells per droplet follows a Poisson distribution that explains the probability of finding the average number of x cells λ per droplet (Howard Shapiro, Practical Flow Cytometry, 4th edition, Wiley-Liss, 2003). In single-cell ChIP-seq experiments, the cell density was adjusted to encapsulate λ=0.1 cells in 45 pl droplets, with 90.5% being empty, 9% containing one single cell, 0.5% containing two cells, and 0.015% containing more than two cells. Real-time monitoring of cell compartmentation within droplets is performed by pre-labeling cells with calcein AM (a non-fluorescent derivative of calcein). Upon entering the cell, the acetomethoxy group (AM) is cleaved by intracellular esterases, emitting strong green fluorescence (excitation / emission: 495 / 515 nm). The fluorescence is obtained when the droplet crosses the laser beam at the detection point, allowing for the counting of the number of encapsulated cells.
[0104] Droplet collection The droplets are collected in a recovery tube on ice until the end of the mounting process (10-20 minutes, depending on the starting number of cells). After mounting, the drops are incubated at 37°C for MNase digestion.
[0105] Calibration of MNase in a droplet At the end of mounting, the droplets are incubated off-tip for single-cell chromatin fragmentation. The cells are lysed within the droplets to allow the MNase enzyme to access their nuclear DNA. The kinetics of digestion are particularly important to preferentially yield mononucleosomes, which are retained within the droplets. The ideal incubation time is defined as the time required for 100% of the nuclear DNA to be fragmented into mononucleosomes. Digestion conditions, including lysis buffer composition, MNase concentration, and incubation time, are precisely calibrated for each sample by performing time-course tests. Calibration is performed as follows: 45 pl droplets containing cells, lysis buffer, and MNase are produced, collected in a recovery tube, and left at 37°C for different incubation times. At each time point, a fraction of droplets is broken, and the MNase is immediately inactivated by the addition of EGTA (see Figure 3). The DNA fragments are then purified and analyzed by electrophoresis. The incubation time selection is balanced to have the highest ratio of mono-nucleosomes while simultaneously preventing over-digestion of nucleosomal DNA. In fact, it was assumed that the DNA protruding from the nucleosomes should be long enough to allow for efficient ligation of the barcode in subsequent steps of the procedure.
[0106] Control of MNase activity within droplets MNase activity within droplets is controlled by retrieving the droplets onto ice during cell encapsulation (see Figure 2). Indeed, time t=0 in Figure 2 corresponds to the fraction of droplets obtained at the end of droplet production, but it is just before incubation, indicating that nuclear DNA has not yet been digested by MNase. This evidence confirms that chromatin digestion does not occur during droplet production but is immediately initiated by incubation at 37°C (see Figure 2).
[0107] Placing droplets on ice after incubation and during reinjection within the fusion device can "quit" MNase activity, thereby limiting intercellular variability in chromatin digestion. For this purpose, two droplet fractions are taken 12 minutes after MNase incubation: one fraction is processed immediately to control digestion, while the second fraction is stored on ice for 1 hour before being processed similarly. As expected, at time points t=12 minutes and t=12 minutes + 1 hour (on ice) in Figure 2, it is confirmed that MNase is no longer active in the fraction stored on ice (compared to time point t=20 minutes). Consequently, storing droplets on ice "quits" MNase activity and thus prevents intercellular variability in chromatin digestion.
[0108] DNA barcoding strategies DNA barcodes are grafted onto hydrogel beads via streptavidin-biotin binding and photocleavable moieties that allow release from the beads upon exposure to UV light (Klein et al. 2015, Cell 161(5):1187-1201). Barcode synthesis is performed by dispersing the beads into a microtiter well plate containing a ligation reagent and 96 combinations of 20bp oligonucleotides (later referred to as Index 1). Index 1 is ligated onto the beads, and the latter is pooled until it is again dispersed into a second microtiter plate containing 96 new combinations of 20bp oligonucleotides (later referred to as Index 2). This split-pool method is repeated three times to produce 96 3 A library of 100,000 possible barcode combinations can be easily produced (i.e., 884,736 combinations).
[0109] Quality control of barcoded hydrogel beads Barcoded beads are one of the core reagents in scChIP-seq technology, and their quality is systematically controlled to confirm intercellular variability that stems from genuine biological differences in their histone modification patterns rather than technical artifacts.
[0110] Tapestation profiles of DNA barcodes released from beads revealed that >75% were full-length (larger peak 146 bp), as well as the presence of incomplete intermediates (Figure 5). On average, the number of full-length barcodes was 5 × 10⁶ per barcoded hydrogel bead. 7 It is estimated to be a copy.
[0111] To evaluate the release of barcodes from hydrogel beads, a DNA probe is hybridized to the barcodes on the beads. The latter is then encapsulated in a 100 pl droplet and harvested off-tip, similar to scChIP-seq experiments. A portion of the droplet is reinjected in a single file into a micrometer chamber, as reported by Eyer (Eyer et al. 2017, Nature Biotechnology 35(10):977-982), and the beads are imaged by epifluorescence microscopy while the fluorescent barcodes are bound to the beads. As expected, the fluorescence is localized on the beads (see Figure 5). A second fraction of the droplet containing the beads is exposed to UV light to initiate barcode release. As described above, epifluorescence microscopy of the droplet containing the beads after photosection reveals a uniform distribution of fluorescence within the droplet, suggesting complete barcode release (see Figure 5). Finally, single-bead sequencing is performed for all new batches of barcoded beads. The beads are separated by limiting dilution in a 384-well plate. Only the wells containing one bead are selected by imaging for barcode amplification and sequencing. Analysis of the sequencing data shows the first and second most abundant barcode fractions for 16 beads. While 100,000 different barcodes are identified per bead, on average, the most abundant barcodes account for 97.7% of the sequencing reads. The second most abundant barcodes account for only 0.17% of the reads on average, suggesting that all other identified barcodes are negligible (see Figure 5).
[0112] Barcode Design The barcodes are bound to beads via streptavidin-biotin bonds, which are separated from the 5' end of the oligonucleotide by a photocleavable substance. The latter consists of a photocleavable group and an alkyl spacer to minimize steric interactions (all substances are called PC linkers, see Figure 4). The first biotinylated and PC linker oligonucleotides are common to all barcodes and contain a T7 promoter sequence and an Illumina sequencing adapter (SBS12 sequence). The T7 promoter sequence acts as a recognition site for T7 RNA polymerase to initiate linear amplification of enriched barcoded nucleosomes after immunoprecipitation in in vitro transcription (IVT). This amplification strategy is widely applicable to unbiased, highly sensitive, and reproducible amplification of cDNA after reverse transcription in single-cell RNA-seq protocols (Hashimshony et al. 2012, Cell Reports 2(3):666-673). In the second step, the Illumina sequencing adapter acts as a PCR handle to complete the preparation of the sequencing library. This adapter is also required for next-generation sequencing of the sample, acting as a primer to initiate read #2 and barcode sequence reading. These first common oligonucleotide-grafted beads are then used for barcode synthesis via sequential ligation of the three indices. Unfortunately, analysis of the first single-cell ChIP-seq dataset shows that only a small percentage (approximately 38%) of the reads have complete and correct barcode structures.
[0113] An optimized barcode structure that enables digestion of barcode concatemers and reduces ligation of non-full-length barcodes is shown in Figure 3. The barcode is framed with half of the Pac1 restriction site, and it is reconstructed only in the case of concatemer formation. These are digested after immunoprecipitation but before linear amplification to clean up the library. The photocleaved side of the barcode is modified by the incorporation of a 3'C3-spacer. This modification introduces a spacer arm to the 3'-hydroxyl group of the 3' base, inhibiting ligation. The addition of the spacer group forces the direction of ligation from the 3' end of the barcode to the nucleosome. A non-full-length barcode is completed by a “block” oligonucleotide sequence containing a 3'C3-spacer and a 5' inverse dideoxy-T base. Again, both modifications are intended to limit unwanted ligation events.
[0114] Encapsulation of barcoded beads within a droplet Loading discrete objects, such as hydrogel beads, into a droplet can be estimated using a Poisson distribution. Similar to cell inclusion, bead loading is monitored in real time as the droplet crosses the laser beam at the detection point. In single-cell ChIP-seq experiments, droplets containing barcoded hydrogel beads are typically acquired in 65%–75% of cases.
[0115] Merging droplets containing nucleosomes with droplets containing barcodes. Cells and DNA barcodes are encapsulated separately to prevent barcode digestion by MNases. To index chromatin at the single-cell level, the DNA barcodes must be delivered in a second step into droplets containing nucleosomes. This is achieved by active fusion of two droplet clusters within a dedicated microfluidic device using a triggered electric field.
[0116] Droplets derived from the "cell emulsion" and the "barcode emulsion" are reinjected into a microfluidic fusion device in a single file. A one-to-one pairing of the two emulsion-derived droplets is required to achieve proper electrocoalescence. Since contact is necessary for the two droplets to fuse, hydrodynamic forces enable the faster and smaller 45 pl droplet ("cell emulsion") to catch up with and contact the 100 pl droplet ("barcode emulsion") (Mazutis et al. 2009, Lab on a Chip 9(18):2665). Similar to droplet production, the fluorescence intensity of the fused droplets is obtained when they cross the laser beam at the detection point (see Figure 1).
[0117] Reconstruction of cell type-specific chromatin states from single-cell ChIP-seq profiles As shown in Figure 7, human T lymphocytes and human B lymphocytes are encapsulated separately and indexed with two separate barcode sets. After barcoding of nucleosomes in droplets, the indexed chromatins from the two cell types are combined, and the library is sequenced by chromatin immunoprecipitation.
[0118] The introduction of bias (batch effect) in immunoprecipitation or preparation of sequencing libraries is avoided by combining it with indexed chromatin. Each sequencing read contains dual information: (1) a single-cell barcode sequence (assigning the read to its cell of origin); and (2) a "cell type specific sequence" (assigning the read to a single cell type (B lymphocyte or T lymphocyte)).
[0119] To confirm that the barcodes were unique to single cells, experiments were conducted using a mixture of mouse and human cell lines, as shown in Figure 13, which demonstrated that 97% of the barcodes were clearly assigned to a single species, consistent with the percentage of occupied droplets (95%) containing single cells.
[0120] The efficiency and accuracy of the scChIP-seq procedure were evaluated to reconstruct cell identity from the distribution of H3K4me3 and H3K27me3-modified single cells. Human Ramos (B cells) and Jurkat (T cells) were processed separately using two unique sets of barcoding adapters (as shown in Figure 1a-b), and after ligation of the adapters in droplets, the barcoding nucleosomes were pooled and immunoprecipitated. For the histone marks of H3K4me3 and H3K27me3, mean coverage of 1,630 and 1,633 unique reads per cell, respectively, and high correlation across technical and biological replicates were achieved (Figure 14a-c, r=0.96 and 0.98 with p<10⁻¹⁵, respectively).
[0121] For both single-cell chromatin profiling experiments, consensus clustering identified two stable clusters corresponding to each cell line (Figures 15a and 16a), matching cell identity with specificity exceeding 99.7% and 99.5% for H3K4me3 and H3K27me3 profiles, respectively. The collected single-cell profiles accurately reproduced bulk ChIP-seq profiles (Figures 15b-c, r=0.93 and 0.97 with p<10⁻¹5 for H3K4me3 and H3K27me3, respectively, Figure 14f). Through differential analysis, permissive and repressive chromatin features specific to Ramos and Jurkat cells were identified (Figure 16b). Focusing on H3K4me3 accumulating around transcription start sites, we identified concordant, lineage-specific gene sets as enriched chromatin features specific to each cell line (Figure 16c). These results confirm that the scChIP-seq procedure is a robust method for detecting chromatin landscapes at the single-cell level, enabling high-precision classification of single cells according to their chromatin state and identification of discriminative chromatin features across cell populations.
[0122] Deconvolution of single-cell barcodes and cell type-specific sequences In each linker, allowing up to one mismatch, the barcode was extracted from read #2 by a first search for immutable 4bp linkers found between the 20-mer indices of the barcode. If the correct linker was identified, three interspersed 20-mer indices were extracted and concatenated together to form a 60bp non-redundant barcode sequence. Three sets of 96 indices (96 3 Using a library of all 884,736 combinations, barcode sequences were mapped using the highly sensitive read mapper Cushaw3. Each set of indices is error-correcting, as it takes more than an edit-distance of 3 to convert one index to another. Therefore, we set a total mismatch threshold of 3 across the entire barcode, with no more than 2 mismatches per index, to avoid misassigning sequences to incorrect barcode IDs. In a second, slower step, sequences that could not be mapped to the Cushaw3 index-library were split into their individual indices, and each index was compared against a set of 96 possible indices, allowing up to 2 mismatches for each individual index. Any sequences that could not be assigned to a barcode ID by these two steps were discarded.
[0123] Profiling histone modifications at the single-cell level with high coverage, averaging up to 10,000 loci per cell, provides a means to reveal the presence of relatively rare chromatin states within tumor samples. Single-cell chromatin profiling is expected to be a unique tool for investigating the role of heterogeneity and dynamics of chromatin states in any complex biological system: in addition to cancer, it can be applied to other diseases and healthy systems, and, in particular, to study cell differentiation and development for patient stratification.
[0124] The method according to the present invention demonstrates the existence of rare cells possessing chromatin characteristics specific to resistant cancer cells, and that these cells can be selected for cancer treatment. before It can be used to clarify.
Claims
1. A method for profiling the chromatin state of single cells obtained from a target using a microfluidic system, The above method involves the following steps: a. A step of providing at least two droplets of the first type, Here, the first type of droplet is, i. Biological elements, ii. Dissolution buffer, and iii. Nuclease, including, b. A step of recovering the first type of droplet under conditions that temporarily inactivate the nuclease during encapsulation, wherein the conditions include selecting a temperature in the range of -20°C to 10°C, thereby suspending the activity of the nuclease. c. A step of incubating the recovered first type droplets in a controlled and synchronized manner and simultaneously reactivating the nuclease, wherein the incubation includes selecting a temperature in the range of 20°C to 40°C. d. Providing at least two droplets of a second type, wherein the droplets of the second type contain nucleic acids, e. A step of merging the first and second types of droplets to generate a third type of droplet, f. Incubating the third type of droplet described above to thereby bind the nucleic acid to one or more target genomic regions, g. The step of sequencing one or more of the target genomic regions. including, method.
2. The method according to claim 1, The aforementioned one or more target genome regions (one or more) include one or more modified genome regions (one or more), method.
3. The method according to claim 2, The one or more modified genomic regions include post-translational modifications and / or histone variants and / or modified DNA sequences. method.
4. A method according to any one of claims 1 to 3, Cells are obtained from subjects in a pathological state and / or suspected pathological state or from healthy subjects. method.
5. The method according to any one of claims 1 to 4, The subject is a pathological condition, and said pathological condition includes cancer, infectious diseases, autoimmune diseases, metabolic diseases, inflammatory diseases, genetic diseases, and non-genetic diseases. method.
6. A method according to any one of claims 1 to 5, The chromatin state of a single cell shows a loss of chromatin markers related to genes that promote drug resistance. method.
7. The method according to claim 6, The chromatin marks are histone modifications H3K4me3 and H3K27me3. method.
8. A method for identifying one or more target genomic regions using a microfluidic system, The one or more target genome regions include one or more modified genome regions. The above method involves the following steps: a. Providing at least one droplet of a first type, Here, the first type of droplet is, i. Biological elements, ii. Dissolution buffer, and iii. Nuclease, including b. A step of collecting the first type of droplet under conditions that temporarily inactivate the nuclease, wherein the conditions include selecting a temperature in the range of -20°C to 10°C. c. A step of incubating the recovered first type droplets in a controlled and synchronized manner and simultaneously reactivating the nuclease, wherein the incubation includes selecting a temperature in the range of 20°C to 40°C. d. Providing at least one droplet of a second type, wherein the droplet of the second type comprises a nucleic acid. e. A step of merging the first and second types of droplets to generate a third type of droplet, f. A step of incubating the third type of droplet so that the nucleic acid can be bound to at least one(s) desired genomic region. g. The step of sequencing one or more of the target genomic regions. including, method.
9. The method according to claim 8, The modified genomic region includes post-translational modifications selected from the group including acetylation, amidation, deamidation, carboxylation, disulfide bond formation, formylation, glycosylation, hydroxylation, methylation, myristoylation, nitrosylation, phosphorylation, prenylation, ribosylation, sulfation, SUMOation, ubiquitination, and derivatives thereof. method.
10. The method according to claim 8 or 9, The nucleic acid is asymmetrically bound to at least one or more target genomic regions as described above. method.
11. Use of nucleic acid in the method of claim 1 or claim 8, wherein the nucleic acid is a. At least one index array, b. Sequencing adapter, and c. At least one protective feature located at the 3' end and / or 5' end. including, use.
12. The use of nucleic acid according to claim 11, The nucleic acid further comprises at least one cleavage site. use.
13. Use of nucleic acid according to claim 11 or 12, The protective function is selected from the group including a space element at the 3' end and a dideoxy-modified base at the 5' end, or vice versa. use.
14. The use of nucleic acid according to claim 12, The aforementioned at least one cutting region is a limiting region that includes a palindromic structure region. use.
Citation Information
Patent Citations
Systems and methods for epigenetic sequencing
WO2013134261A1