Conformation capture methods
The new conformation capture method addresses the limitations of existing technologies by using combinatorial indexing to specifically identify DNA regulatory interactions within accessible chromatin regions in each cell of a heterogeneous sample, achieving enhanced sensitivity and resolution.
Patent Information
- Application Number
- PCT/GB2024/053113
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Current conformation capture methods for identifying DNA regulatory interactions are limited by high background levels, complexity, and low resolution, particularly when applied to heterogeneous cell populations, which makes it difficult to detect specific interactions between accessible chromatin regions.
A new conformation capture method that uses combinatorial indexing to identify DNA regulatory interactions in each cell of a heterogeneous cell sample, reducing background noise and increasing resolution by focusing on interactions within open or accessible chromatin regions.
This method enables the analysis of DNA regulatory interactions with improved sensitivity and resolution, allowing for the identification of specific interactions in each cell type within a heterogeneous sample without the need for cell isolation.
Smart Images

Figure GB2024053113_19062025_PF_FP_ABST
Abstract
Description
[0001] CONFORMATION CAPTURE METHODS
[0002] TECHNICAL FIELD
[0003] The present invention relates to conformation capture methods to identify DNA regulatory interactions between accessible chromatin regions. Specifically, the methods relate to identifying DNA regulatory interactions comprising steps of capturing and sequencing to separate and analyse interactions in open chromatin, for example, between enhancers and promoters, at individual loci at high resolution and to kits for performing said methods. In particular, the present invention relates to a conformation capture method using combinatorial indexing to identify DNA regulatory interactions of each single cell in a heterogeneous cell population e.g. a tissue or blood sample. The method enables assessment of each cell type or state in a cell population of interest in one assay, without the need to isolate individual cells or cell types.
[0004] BACKGROUND
[0005] The identity of each cell type in the human body relates to its specialized function and the establishment and maintenance of a cell’s identity is largely driven by chromatin structure. The spatial organisation of chromatin plays a central role in gene expression, DNA replication and repair. Physical access to DNA is a highly dynamic property of chromatin vital for the establishment and maintenance of cellular identity. Accessible chromatin across the genome occurs at regulatory elements such as enhancers, promoters, and insulators which engage in specific physical interactions through chromatin folding to regulate gene expression.
[0006] The assay for transposase-accessible chromatin with sequencing (ATAC-seq) provides a method for assaying chromatin accessibility genome wide, enabling a library preparation of chromatin accessible regions across the genome followed by the sequencing and identification of these regions. ATAC-seq employs tagmentation, a “cut & paste” method involving a hyperactive TN5 transposase, to insert sequencing adapters into open chromatin regions followed by next-generation sequencing to uncover how chromatin packaging and other factors affect gene expression (see Buenrostro J.D. et al., 2013, Nature Methods, 10, 1213-1218). Adey A.C., 2021 , Genome Research, 31 :1693-1705, sets out developments in the use of tagmentation, in which a hyperactive transposase is used to simultaneously fragment target DNA and append universal adapter sequences, in single-cell genomics. Adey A.C. describes various ways of using TN5 mediated insertion of a tag into accessible chromatin regions (ATAC-seq) to map accessible regions throughout the genome. Accessibility of chromatin is a hallmark of gene and genome regulatory elements and ATAC- seq is currently used instead of DNAase sensitivity to map these elements across the genome. Other methods described are for RNAseq or RNAseq combined with chromatin accessibility, or simply to create DNA libraries for whole genome sequencing. None of the methods described refer to the use of tagmentation to map interactions between regulatory sequences.
[0007] Regulatory elements, such as promoters and enhancers, are the primary genomic regulatory components of gene expression and have been shown to contribute to health and disease, for example, in cancer and auto-immune disorders. Enhancers are composed of clusters of transcription factor binding sites and can activate gene expression over large genomic distances due to the three dimensional (3D) organisation of chromatin bringing the promoters of the genes they control into close proximity. Approaches for capture of the DNA interactions between regulatory elements, such as promoters and enhancers, have been developed and are widely applied to study the impact of regulatory landscape dynamics on gene expression and phenotype establishment, as well as to study the role of genetic modifications in disease development.
[0008] Technologies that have been directed to identifying DNA interactions include, restriction fragment-based technologies, such as, chromosome conformation capture (3C) techniques which analyse the spatial organisation of chromatin in a cell. The principle of 3C technology is that interactions between distant regulatory regions that come close together in the 3D space will be more frequently detected than random interactions. Developments in the technology coupled with next-generation sequencing have led to Hi-C (see WO2010 / 036323 and Lieberman-Aiden E. et al. 2009, Science, 326(5950) :289-93) which uses junction markers to isolate all of the ligated interacting sequences in the cell to provide information on all interactions occurring within the nuclear composition at a particular time point. However, the resulting libraries are extremely complex making it difficult to identify interactions between distant regulatory regions that come close together in the 3D space more frequently than random interactions unless sequenced to a depth of 10 billion reads. To overcome this limitation high resolution methods such as 4C, Capture-C and Promoter Capture Hi-C (PCHi- C) (see WO2014 / 168575, WO2015 / 033134 and WO2021 / 064430) have been developed to capture chromosomal contacts between genes and their regulatory regions with statistical significance. Furthermore, these methods are time-consuming and are restricted to a limited number of pre-defined loci.
[0009] WO2022 / 094474A1 discloses a method of Hi-C on accessible regulatory DNA (HiCAR) which involves using a population of 30,000 to 100,000 cells to generate DNA for analyzing chromatin structure. The HiCAR method uses tagmentation to insert an adapter sequence into accessible regions, followed by steps of restriction enzyme digestion and in situ proximity ligation. The HiCAR method captures interactions between accessible regions and all genomic fragments making it more sensitive than Hi-C, but still provides complex results with low resolution.
[0010] Conformation capture, such as, Capture Hi-C-based assays are currently used on purified cell types, to provide DNA regulatory interaction profiles for each cell type, but isolating individual cell types is difficult and time consuming. If such conformation capture technology is used on heterogeneous tissues only an average signal is obtained across the multiple cell types. There is therefore a need for a conformation capture method to identify DNA regulatory interaction profiles and, in particular, for a single cell method to identify DNA regulatory interaction profiles for each cell in a heterogeneous sample.
[0011] SUMMARY OF INVENTION
[0012] The present inventors have developed new conformation capture methods, that capture the three-dimensional structure of accessible chromatin in whole genome scale, enabling the analysis of DNA regulatory interactions that occur between regions of regulatory DNA within accessible chromatin regions in a cell. In particular, the present inventors have developed a new single cell conformation capture method, for identifying the interactions between accessible regulatory DNA in each cell of a heterogeneous cell sample. The present method overcomes limitations of all current methods. Importantly, the present method overcomes the high background levels generated by current methods such as Hi-C (which measures all-to-all), HiCAR (which measures accessible regions-to-all) and PCHi-C (which measures promoters-to-all), which limits their sensitivity to detect regulatory interactions. Specifically, Hi-C, PCHi-C and HiCAR methods are restricted to interactions between particular restriction fragments, which is dependent on the choice of restriction enzyme and the resulting DNA fragments come from both open and closed chromatin. In contrast, the present method identifies all interactions that occur in open or accessible chromatin and, therefore, measures only interactions between accessible regions, thus greatly reducing the huge background of non-regulatory interactions found in the other methods. Further, when combinatorial barcoding is incorporated in the single cell version of the present method, the method can provide single cell measurements from hundreds of thousands of cells in a single experiment. It overcomes the issue of multi-cell heterogeneity of primary tissues or samples enabling the mapping of specific DNA regulatory interactions to each cell type in a highly heterogenous sample. This could include groups of cells at different cell states, or stages of development or disease progression.
[0013] The method steps of the present invention enrich for DNA regulatory interactions that come into close proximity due to the 3D organisation of the genome to provide DNA regulatory interaction profiles with reduced complexity and improved resolution, compared to the more complex results provided by previous methods, such as Hi-C, HiCAR, and PCHi-C. The methods of the present invention can be used to identify DNA regulatory interactions with statistical significance. The identification of DNA regulatory interactions may provide a measure of the function of a gene in the form of the inducement of gene regulation or expression. The present method can be applied with or without combinatorial indexing. The bulk processing method of the present invention can be applied with no indexes to provide an average DNA regulatory interaction profile of the cells in the sample. Applying combinatorial indexing to intact cells or nuclei in the single cell process of the present invention makes it possible to attain data from thousands of single cells from a heterogeneous sample such as a dissociated tissue sample, a blood sample, or a mixture of cell types and / or cell states without needing to individually process them or separate them first. Further, where combinatorial indexing is not applied but data from single cells is required, the present method may involve a step of cell sorting of each cell into a compartment, where a unique index can be assigned to the DNA from each cell. Processes according to the present methods may include the addition of one index, two indexes or three indexes. Preferably, three indexes are added. Four or more indexes could be considered where larger cell samples are processed. For samples containing a single cell type such as an isolated primary cell type or a cell line, the same index / barcode may be assigned to all the cells from that sample and all the data from that sample can be analysed in bulk. Alternatively, a single indexing step may be used to mark multiple bulk samples before processing them together in a single reaction to increase throughput of the bulk method.
[0014] In the present invention, distributing or physical separation of the cells into compartments may be used to uniquely label each cell with an index / barcode so that the interaction data from each cell can be identified. Traditionally, cell sorting is used to purify a cell type of interest, for example, by FACS, column-based cell isolation, or microdissection. However, in the present invention when combinatorial indexing is used or physical separation of each cell into a compartment, a step of cell sorting is not required. Compartments may be any means of separating cells, for example, tubes or wells in a plate. In a first aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in cells, the method comprising the steps:
[0015] (a) treating the cells with a fixative to cross-link the DNA regulatory interactions;
[0016] (b) processing the cells to fragment the DNA and introduce first adapters into accessible regulatory DNA;
[0017] (c) proximal ligation of the DNA fragments to join the proximal first adapters of the DNA fragments, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;
[0018] (d) reversing the cross-linking;
[0019] (e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and
[0020] (f) sequencing and sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions.
[0021] In a second aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in each cell in a heterogenous cell population, the method comprising the steps:
[0022] (a) treating the heterogenous cell population with a fixative to cross-link the DNA regulatory interactions;
[0023] (a1) distributing the cells from the heterogenous cell population into a first plurality of compartments, each compartment comprising a subset of cells;
[0024] (b) processing the cells to fragment the DNA and introduce first adapters comprising a first index into accessible regulatory DNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index;
[0025] (b1) pooling of first indexed DNA fragments and distributing into a plurality of second compartments;
[0026] (c) proximal bridge ligation of the first indexed DNA fragments to introduce and join one or more second adapter comprising a second index to proximal first adapters of the first indexed DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;
[0027] (d) reversing the cross-linking; (e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and
[0028] (f) sequencing and multiplex sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions; wherein sequence reads bearing the same combination of indexes will be derived from a single cell.
[0029] In a third aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in each cell in a population of cells, the method comprising the steps:
[0030] (a) cross-linking the DNA regulatory interactions of the cells;
[0031] (a1) distributing the cells into a first plurality of compartments, each compartment comprising a subset of cells;
[0032] (b) processing the cells to fragment the DNA and introduce first adapters, optionally comprising a first index into accessible regulatory DNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index;
[0033] (b1) pooling of first indexed DNA fragments and distribution into a plurality of second compartments;
[0034] (c) proximal bridge ligation of the first indexed DNA fragments to introduce and join one or more second adapter, optionally comprising a second index to proximal first adapters of the first indexed DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes;
[0035] (c1) optionally, pooling of second indexed DNA fragments and distribution into a plurality of third compartments;
[0036] (d) reversing the cross-linking;
[0037] (e) enriching for DNA fragments comprising a first half of a binding pair moiety using a second half of the binding pair moiety; and
[0038] (f) sequencing and multiplex sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions, optionally, wherein the sequencing comprises paired-end sequencing with sequencing primers comprising a third index to generate third indexed DNA fragments; wherein sequence reads bearing the same combination of indexes will be derived from a single cell; and wherein the DNA fragments comprise one or more nucleotides comprising the first half of the binding pair moiety, preferably the second indexed DNA fragments comprise one or more nucleotides comprising the first half of the binding pair moiety.
[0039] According to the present invention, one, two, three or more indexes may be added. When two indexes are used, the indexes may be added to the DNA fragments at step (b) with one or more of the first adapters and at step (c) with one or more of the second adapters. Alternatively, at step (b) and at step (d) or step (f) with sequencing adapters. Again alternatively, at step (c) and step (d) or step (f). It is envisioned that further indexes could be added as required, however, addition of the indexes alongside adapters is advantageous when carrying out the protocol.
[0040] In a fourth aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in cells, the method comprising the steps:
[0041] (a) cross-linking the DNA regulatory interactions in the cells;
[0042] (a1) distributing multiple samples across compartments;
[0043] (b) performing tagmentation with a recombinase enzyme to fragment the DNA in the cells and introduce first adapters with compatible sticky ends or blunt ends and a unique index for each sample into accessible regulatory DNA;
[0044] (c) proximal ligation of the DNA fragments, wherein the DNA fragments further comprise a first half of a binding pair moiety;
[0045] (d) reversing the cross-linking;
[0046] (e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and
[0047] (f) sequencing and sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions.
[0048] In a fifth aspect, the invention provides a method of identifying DNA regulatory interactions in a cell population that are indicative of a particular disease state, comprising:
[0049] (a) performing the method according the first or fourth aspect on a cell sample obtained from an individual with a particular disease;
[0050] (b) quantifying a frequency of DNA interactions within the cell sample; and
[0051] (c) comparing the frequency of DNA interactions from the individual with said disease state with the frequency of interaction in a normal control cell sample from a healthy subject, such that a difference in the frequency of DNA interactions is indicative of a particular disease.
[0052] In a sixth aspect, the invention provides a method of identifying DNA regulatory interactions in a heterogenous cell population that are indicative of a particular disease state, comprising:
[0053] (a) performing the method according to the second or third aspect of the present invention on a heterogenous cell sample obtained from an individual with a particular disease;
[0054] (b) quantifying a frequency of DNA interactions within an identified single cell or a group of single cells from the same cell type or cell state; and
[0055] (c) comparing the frequency of DNA interactions from the individual with said disease state with the frequency of interactions in a normal control heterogenous cell sample from a healthy subject, such that a difference in the frequency of DNA interactions is indicative of a particular disease.
[0056] In a seventh aspect, the invention provides a method of identifying DNA regulatory interactions in a cell population that are indicative of the effect of a medicament of interest, comprising:
[0057] (a) performing the method according to the first or fourth aspect of the present invention on a cell sample obtained from an individual treated with said medicament;
[0058] (b) quantifying a frequency of DNA interactions within the cell sample; and
[0059] (c) comparing the frequency of DNA interactions from the individual treated with said medicament with the frequency of DNA interactions in a normal control cell sample from a non-treated subject, such that a difference in the frequency of DNA interactions is indicative of the effect of treatment with said medicament.
[0060] In an eighth aspect, the invention provides a method of identifying DNA regulatory interactions in a heterogenous cell population that are indicative of the effect of a medicament of interest, comprising:
[0061] (a) performing the method according to the second or third aspect of the present invention on a heterogenous cell sample obtained from an individual treated with said medicament;
[0062] (b) quantifying a frequency of DNA interactions within an identified single cell or a group of single cells from the same cell type or cell state; and (c) comparing the frequency of DNA interactions from the individual treated with said medicament with the frequency of DNA interactions in a normal control heterogenous cell sample from a non-treated subject, such that a difference in the frequency of DNA interactions is indicative of the effect of treatment with said medicament.
[0063] In a ninth aspect, the invention provides a kit for identifying DNA regulatory interactions between accessible chromatin regions in a cell population, preferably in each cell in a heterogenous cell population, comprising buffers and reagents capable of performing any one of the methods disclosed herein. In some embodiments, kits may include the buffers and reagents necessary to carry out the methods disclosed herein. In other particular embodiments, the kit includes equipment, reagents and instructions for the methods disclosed herein. In further particular embodiments, the kit includes equipment, reagents, computer software program and instructions for the methods.
[0064] BRIEF DESCRIPTION OF THE FIGURES
[0065] Figure 1. Schematic overview of the single cell conformation capture protocol using combinatorial barcoding is presented herein.
[0066] Figures 2A, 2B and 2C. Figure 2A shows a schematic for a cross-linked DNA fragment. Figure 2B shows a schematic identifying the step for addition of the first index into accessible chromatin, for example, transposon insertion by tagmentation. Figure 2C shows a schematic identifying the step for addition of the second index by proximal bridge ligation of the two interacting regions.
[0067] Figure 3. Shows a schematic of one possible method for proximal bridge ligation according to the present invention. Tagmentation introduces the Tn5 adapter into accessible regulatory DNA. The Tn5 adapter contains the first barcode / index and has an A overhang at its free 3’ end. 5’ phosphates (P) are present to allow ligation of fragments. The bridging adapter contains barcode / index 2, biotin (Bio) for purification, 5’ phosphates (P) for ligation and 3’ T overhangs. During bridge ligation, the T overhangs in the bridging adapter anneal to the A overhangs of the Tn5 adapters and are ligated, bringing the two interacting regions of interest together in one DNA fragment. Finally, the gaps in the DNA generated during the tagmentation by Tn5 are filled in with polymerase and ligated. Figure 4. Shows an Integrative Genomics Viewer (IGV) plot for genome resolution comparing two human CD4+ T cell single cell libraries with a human CD4+ T cell ATAC-seq library. The figure shows a snapshot in Integrative Genomics Viewer (IGV) of alignment of the three CD4+ libraries to the human genome. The plot shows chromosome 8 and the region around 127,500 kb to 130,000 kb in the first track The second track is an ATAC-seq library generated from CD4+ T cells (scale: 0 to 5.5) and the two tracks below are of libraries independently prepared using the new protocol using purified CD4+ T cells and a single index combination (scale: 0 to 18 and 0 to 41 , respectively). The fifth track provides an illustrative list of relevant genes at the locations of the reads. The data demonstrates that the new method has a similar profile in IGV to the ATAC-seq library, showing that the interaction data generated from the new method is from hyper-accessible regulatory sites without requiring specific capture of promoter sequences.
[0068] Figure 5. Shows chromatin contact profiles of enhancers for two library types. The profiles shows chromosome 7 and the region around 50,200 kb to 50,400 kb in the first track. The second track provides marked PCHi-C baits. The fourth track is according to the current method generated using purified CD4+ T cells and a single index combination. The third track is using the same cells in a PCHi-C method (WO2021 / 064430). The PCHi-C method follows a 5 day protocol which includes a 2 day protocol of hybridising promoter designed specific probes to specifically pull out the promoter regions. The method according to the present invention follows a 3 day protocol and does not include this enrichment step, since the enrichment in the present case comes from the tagmentation in open chromatin which selects for any region which is in open chromatin, including DNA regulatory regions, such as enhancer-promoter interactions and also other regulatory sequences. The fifth track shows comparative ATAC-seq data and the sixth track provides a relevant gene at the locations of the reads. The data demonstrates that the present method is identifying new enhancer interactions not detected by PCHi-C. PCHi-C specifically enriches for promoter contacts, whereas the new method enriches for any DNA regulatory interactions in open active chromatin.
[0069] Figure 6. Shows an interaction profile comparing the single cell dev. experiment method according to the present invention with promoter capture Hi-C (PCHi-C). The interaction profile shows chromosome 19 and the region 41 ,200 kb to 41 ,400 kb in the first track. A Chicago map with marked Hind III baits is provided for PCHi-C in the second track. The third and fourth tracks compare PCHi-C interaction contacts and single cell interaction contacts, according to the present method. The fifth, sixth and seventh tracks show, single cell coverage in the form of reads according to the present invention, ATAC-seq CD4+ coverage in the form of reads and PCHi-C coverage in the form of reads, respectively. The data shows that the method according to the present invention provides much greater sensitivity and resolution compared to promoter capture Hi-C (PCHi-C). Because PCHi-C is restriction fragment based, short range interactions cannot be detected with confidence due to the high background that is found 0-50kb from the promoter fragment. PCHi-C favours detection of long-range interactions and fails to detect short range interactions. The method according to the present invention identifies numerous short and long-range interactions between the promoter of HNRNPIL1 and ATACseq regions (an indicator of regulatory elements), and does so with fewer sequence reads than PCHi-C, making the present method 10-100 fold more sensitive than PCHi-C and with at least 10 fold improved resolution. The eighth track provides an illustrative list of relevant genes at the locations of the reads.
[0070] Figure 7. Shows an interaction profile providing an expanded view of the HNRNPUL1 locus from the profile of Figure 6, showing an approximately 2Mb region. The interaction profile shows chromosome 19 and the approximate region of 40,000 kb to 42,500 kb in the first track. A Chicago map with marked Hind III baits is provided for PCHi-C in the second track. The third and fourth tracks compare PCHi-C interaction contacts and single cell interaction contacts, according to the present method. The fifth, sixth and seventh tracks show, single cell coverage in the form of reads according to the present invention, ATAC-seq CD4+ coverage in the form of reads and PCHi-C coverage in the form of reads, respectively. The method according to the present invention is shown on the fourth track and detects both short and long-range interactions. ATACseq peaks are shown on the sixth track, it is noted that not all ATAC-seq peaks interact with the HNRNPUL1 promoter demonstrating specificity of contact between regulatory elements and promoter according to the present invention. The third track shows PCHi-C results. The eighth track provides an illustrative list of relevant genes at the locations of the reads.
[0071] Figure 8. Shows a further interaction profile comparing the method according to the present invention (single cell dev experiment 1) to PCHi-C. The interaction profile shows chromosome 19 and the region of 5,750 kb to 6,050 kb in the first track. A Chicago map with marked Hind III baits is provided for PCHi-C in the second track and for the present method, the single cell development experiment 1 , in the third track. The fourth and fifth tracks compare PCHi-C interaction contacts and single cell interaction contacts, according to the present method. The sixth, seventh and eighth tracks show, single cell coverage in the form of reads according to the present invention, ATAC-seq CD4+ coverage in the form of reads and PCHi-C coverage in the form of reads, respectively. The present method outperforms PCHi-C here detecting numerous short-range interactions between the VMAC promoter and ATAC-seq regions whereas PCHi-C does not detect these interactions due to high local background and preference for long-range interactions. The ninth track provides an illustrative list of relevant genes at the locations of the reads.
[0072] Figure 9. Shows the species alignment per barcode wherein each datapoint on the graph can be assigned to a species. The results from Example 3 are dictated in Figure 9, showcasing that human and mouse cells can be identified from a mixed sample using the methods disclosed herein. High quality barcodes were filtered (i.e. those with more than 20 unique, mappable location pairs which can be unambiguously assigned as mouse or human derived) and tested to see whether all the pairs within a cell can be uniquely assigned to either the human or mouse species. Figure 9 shows that 96% of the high quality barcodes describe a single species (i.e. fall on the axes of the graph in this figure). The points along the X axis shows all the cells determined to be mouse, while all points along the Y axis show all the cells determined to be human. Any points found between the ones falling along each axis are described as mixed species where doublet barcodes occurred. This demonstrates that the doublet rate of the technique is approximately 8%, an improvement over current ATAC-seq protocols which report doublet rates in excess of 10%, and do not measure chromatin contacts. The doublet rate is significant as doublet “cells” must be excluded from downstream analysis and interpretation to avoid formation of spurious clusters where the cell type is not well defined.
[0073] Figure 10. Shows the cell type assignment of two differing human cell types (Jurkat cells, shown as squares on the graph, and IMR90 cells, represented by circles on the graph). The squares along the X axis of the graph identify human Jurkat cells, while the circles along the Y axis identify human IMR90 cells. Any unassigned cells are represented as a plus symbol (+) in Figure 10. This Figure, and the results of Example 4, showcase the methods disclosed herein can successfully determine single cells of a particular cell type from an input sample comprising a mixture of human cell types.
[0074] DETAILED DESCRIPTION
[0075] In a first aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in cells, the method comprising the steps:
[0076] (a) treating the cells with a fixative to cross-link the DNA regulatory interactions; (b) processing the cells to fragment the DNA and introduce first adapters into accessible regulatory DNA;
[0077] (c) proximal ligation of the DNA fragments to join the proximal first adapters of the DNA fragments, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;
[0078] (d) reversing the cross-linking;
[0079] (e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and
[0080] (f) sequencing and sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions.
[0081] In a first embodiment, the proximal ligation step (c) comprises proximal bridge ligation, wherein one or more second adapters are introduced and join the proximal first adapters.
[0082] Preferably, step (c) is performed prior to step (e). Again preferably, the one or more nucleotides comprising a first half of a binding pair moiety are introduced to the DNA fragments before step (e).
[0083] In a second embodiment, the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (b), preferably one or more of the first adapters comprises the one or more nucleotides comprising the first half of the binding pair moiety. Again preferably, the first half of the binding pair moiety may be introduced on incorporated nucleotides during any fill in steps required.
[0084] The first half of the binding pair moiety may comprise biotin and the second half of the binding pair moiety may comprise streptavidin.
[0085] In a third embodiment, the method further comprises introducing one or more indexes to the DNA fragments. Preferably, each index comprises a barcode or Unique Molecular Identifier (UMI) which differs for each index.
[0086] In a fourth embodiment, the cells comprise a homogenous cell population or a heterogenous cell population. The heterogenous cell population may comprise cells of different cell types and / or different cell states, for example, stages in the cell cycle or development. Preferably, the heterogenous cell population comprises different cell types from a sample, for example a human blood sample or a human tissue sample. Again preferably, the heterogenous cell population comprises a sample of cells in different cell states. Samples for use according to the present invention may be sourced from, for example, a human or animal or an immortalized cell line.
[0087] In a fifth embodiment, after step (a) the cells are distributed into a first plurality of compartments, each compartment comprising a subset of cells (step (a1)).
[0088] In a sixth embodiment, processing the cells in step (b) further comprises introducing the first adapters comprising a first index into accessible regulatory DNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index that is different from the first index sequence in the other compartments. Preferably, after processing the cells in step (b) the first indexed DNA fragments are pooled and distributed into a plurality of second compartments (step (b1)).
[0089] In a seventh embodiment, the proximal ligation step (c) further comprises proximal bridge ligation of the DNA fragments to introduce and join one or more second adapter comprising an index to the proximal first adapters of the DNA fragments, to generate indexed DNA fragments. Preferably, the index in the proximal bridge ligation step (c) comprises a second index introduced to the DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes.
[0090] In an eighth embodiment, step (d), step (e) or step (f) further comprise introducing a sequencing adapter or sequencing index. The sequencing index comprises a sequencing adapter with an index.
[0091] In a second aspect, the invention provides a method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in each cell in a heterogenous cell population, the method comprising the steps:
[0092] (a) treating the heterogenous cell population with a fixative to cross-link the DNA regulatory interactions;
[0093] (a1) distributing the cells from the heterogenous cell population into a first plurality of compartments, each compartment comprising a subset of cells;
[0094] (b) processing the cells to fragment the DNA and introduce first adapters comprising a first index into accessible regulatory DNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index;
[0095] (b1) pooling of first indexed DNA fragments and distributing into a plurality of second compartments;
[0096] (c) proximal bridge ligation of the first indexed DNA fragments to introduce and join one or more second adapter comprising a second index to proximal first adapters of the first indexed DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;
[0097] (d) reversing the cross-linking;
[0098] (e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and
[0099] (f) sequencing and multiplex sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions; wherein sequence reads bearing the same combination of indexes will be derived from a single cell.
[0100] According to the second aspect, in an embodiment, the second indexed DNA fragments may be pooled and distributed into a plurality of third compartments (step (c1)).
[0101] Pooling and distribution between steps ensures that each compartment in step (b1) comprises a subset of DNA fragments with the first index sequence that is different from the first index sequence in the other compartments, and each compartment in step (c1) comprises a subset of DNA fragments with the second index sequence that is different from the second index sequence in the other compartments.
[0102] In a further embodiment, step (d) or step (f) comprises the introduction of sequencing adapters comprising a third index to generate third indexed DNA fragments for multiplex sequence analysis. Sequencing adapters are generally required for sequencing and are preferably added after step (d) or after step (e), and could include an index if required. An index many be added by sequencing adapter or by PCR using PCR primers containing the index. If the index is added by adapter this can occur at step (d) or (f). If added by PCR this has to occur after step (e), otherwise the biotin moiety is lost during the PCR and enrichment is not possible after PCR. According to the first and second aspects, in a first embodiment, the processing step
[0103] (b) is performed by tagmentation using a recombinase enzyme, preferably the recombinase enzyme is a retroviral integrase, such as mutant transposase, such as hyperactive Tn5 transposase. Alternatively, the processing step (b) is performed using an endonuclease, for example DNase I, or a restriction enzyme.
[0104] In a second embodiment, the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (b) with the first adapters. Preferably, the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (c) with the one or more second adapters. Where step
[0105] (c) introduces a bridge adapter as a second adapter, a single bridge adapter may incorporate the one or more nucleotides comprising the first half of the binding pair. Alternatively, two second adapters may be used and then ligated together to form a single bridge adapter incorporating one or more nucleotides comprising the first half of the binding pair moiety.
[0106] In a third embodiment, step (c) comprises proximal bridge ligation. Optionally, proximal bridge ligation is performed by two-step ligation in which an adapter is ligated to each end of the Tn5 adapter and then the two adapters are ligated together.
[0107] In a fourth embodiment, the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (c) when the gaps in the DNA fragments generated are filled. Alternatively, the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (c) with the second adapters.
[0108] The methods of the present invention may use combinatorial barcoding to add unique indexes to subsets of DNA fragments, the distribution of DNA fragments into compartments resulting in unique combinations of first, second and / or third indexes. An initial step of cell sorting to separate single cells into individual compartments may be used to profile chromatin interactions in single cells without the use of combinatorial barcoding. Compartmentalisation and barcoding of chromatin interactions in single cells could also be performed after the initial chromatin digestion and ligation steps, using microfluidics, for example on the 10x platform with the single cell ATAC or multiome kits.
[0109] Only a single index may be used where the aim is to multiplex multiple purified cell types or cell lines for high-throughput profiling, without the need to analyse single cells. In this case, a single index may be introduced in accessible chromatin during the initial digestion step, after which samples may be pooled to be processed in a single reaction. Similarly, no indexing is required to profile the interaction landscape between accessible chromatin regions in individual bulk samples. Where combinatorial barcoding of the DNA fragments is not used, the results obtained will provide an indication of the average interactions over the entire population of cells studied.
[0110] The methods of the present invention provide a means for identifying interactions that occur between regulatory sequences (regulatory DNA) that are brought into proximity due to the 3D folding of the chromatin, only regulatory sequences within accessible chromatin regions will be identified, so called accessible regulatory DNA. These interactions, herein referred to as DNA regulatory interactions, are identified within nucleic acid sequences within a sample of cells by targeting, or targeting and indexing two or more regions of these sequences. Such methods have the advantage of focussing the data on particular interactions within enormously complex libraries.
[0111] The cell sample may be homogenous or heterogenous according to different aspects of the invention. Homogenous cell samples are generally considered to consist of samples that are composed of a single cell type. Heterogenous cell samples are samples composed of multiple cell types, such as found in tissues of the body, for example, brain, lung, blood etc. or they can be different cell samples that are mixed in the laboratory so that they can be processed as a single sample in a single assay. Heterogenous cell samples may also comprise cells of the same or different types but in different cell states or developmental stages. The cell sample may contain a heterogenous cell population as the cells do not need to be isolated from the sample, and therefore, multiple cell types may be present which can be resolved during analysis. The cell sample may comprise a heterogenous cell suspension derived from a tissue or the blood, a heterogenous mixture of isolated cell types to be analysed together, or a single isolated cell type. The sample of cells in biological samples may include cerebrospinal fluid (CSF), whole blood, blood serum, plasma, or an extract or purification therefrom, or dilution thereof. Biological samples may also include tissue homogenates, tissue sections and biopsy specimens from a live subject or taken post-mortem. The cell sample may contain a homogenous cell population such as an isolated cell type or a cell line when the bulk-version of the method is applied without combinatorial indexing. The samples can be prepared, for example, where appropriate diluted or concentrated, and stored in the usual manner. In one embodiment, the heterogenous cell population may be a cell sample from a tissue, a plasma sample or a blood sample.
[0112] The cell sample may contain one or more types of cells. As the different cell types do not need to be isolated from the sample, all cells within the sample of interest can be assessed in a single assay. As there is no longer a need to isolate the cells or sort the cells into cell types when conducting the method incorporating combinatorial barcoding disclosed herein, the time it takes to perform such methods and the costs associated are greatly reduced.
[0113] The cell sample may comprise multiple cell types, as well as cells of differing cell subtypes and cells within various developmental stages. For example, a blood sample, may include lymphocytes (T cells, B cells and NK cells), neutrophils and monocytes / macrophages. Subsets of T cells that may be identified include cytotoxic T (CD8+) cells and CD4+ T cells, that can be further identified as T helper (Th) cells, regulatory T cells (Treg) cells and follicular helper T (Tfh) cells. Such cell types are characterised by their different cytokine profiles and can be identified according to their DNA interaction profiles.
[0114] In further embodiments, the cell sample may be derived from 1 million or fewer cells, 750,000 or fewer cells, 500,000 or fewer cells, 250,000 or fewer cells, 200,000 or fewer cells, 100,000 or fewer cells, 50,000 or fewer cells, or 10,000 or fewer cells.
[0115] The cell sample may be obtained from an individual with a particular disease state. The disease state can be selected from, but not limited to: cancer, autoimmune disease, a developmental disorder, a genetic disorder, diabetes, cardiovascular disease, kidney disease, lung disease, liver disease, neurological disease, viral infection or bacterial infection. In a further embodiment, the disease state is cancer, for example, breast, bowel, bladder, bone, brain, cervical, colon, endometrial, oesophageal, kidney, liver, lung, ovarian, pancreatic, prostate, skin, stomach, testicular, thyroid or uterine cancer, leukaemia, lymphoma, myeloma, or melanoma.
[0116] The cell sample may also be obtained after drug treatment, for example, to compare the effects of different drugs on the regulatory DNA interactions within the sample. Samples may be taken before and after treatment and / or at different stages of treatment for comparison.
[0117] Each cell comprises genetic material within a nucleus. The packaging of the genome into the cell nucleus is accomplished through chromatin. Chromatin is the complex of genomic DNA with proteins called histones, where each histone-bound DNA molecule is referred to as a chromosome. Accessible chromatin may also be known as accessible regions, relaxed or open chromatin, open chromatin regions (OCRs) or euchromatin. Chromatin accessibility reflects a network of permissible physical interactions through which enhancers, promoters, insulators and regulatory elements cooperatively regulate gene expression. Regions of accessible chromatin allow regulatory elements more access to interaction sites thereby facilitating transcriptional activation, whilst more densely packed non-accessible regions, or heterochromatin, conceal such sites, such as DNA promoter sequences, from transcription factors. As such, accessible chromatin within the genome indicate regions where gene expression and function are regulated.
[0118] Regulatory DNA regions or sequences may include promoters, enhancers, silencers and insulators. In particular, there are many enhancer sequences within a cell’s DNA structure. Enhancer sequences are non-coding gene regulatory elements that are composed of clusters of transcription factor binding sites. Transcription factor binding opens enhancer chromatin rendering it hyper-accessible, for example, to tagmentation, and more specifically transposase-mediated tagmentation. Enhancers are often separated from their target genes by large genomic distances of up to a megabase or more. Enhancers regulate gene transcription, determining both the timing and the level of transcription of their target genes by physically interacting or coming into close proximity to the promoters of the genes they control i.e. cross-linking can occur between promoters and enhancers within accessible regions of a DNA sequence, accessible regulatory DNA. Such long-range interactions often bypass proximal genes to exert control over specific distal target genes. Determining which enhancers control which genes in which cell types is critical to understanding how human genetic variation controls health and disease.
[0119] According to the present invention, the cells may be treated with a fixing agent or fixative. The fixing agent may also be known as a cross-linking agent. This fixing agent crosslinks the DNA molecules so that genomic loci that are in close spatial proximity become linked, specifically cross-linking DNA regulatory interactions. Cross-linking may refer to any stable chemical association between two regions within the nucleic acid sequence, such that they may be further processed as a unit. A cross-linked / fixed cell, therefore, accurately maintains the spatial relationships between components within the nucleic acid sequence at the time of fixation. Cross-linking can, but does not necessarily, lead to DNA strands coming into contact; however, cross-linking ‘freezes’ the 3D genome organisation and preserves DNA interactions. The fixative may be selected from, but not limited to, the following: paraformaldehyde, formaldehyde, glutaraldehyde, mercuric chloride, disuccinimidyl glutarate (DSG), osmium tetroxide, glyoxal, picric acid, Bouin’s (e.g. 25% of 37% formaldehyde solution, 70% picric acid, 5% acetic acid), methacarn (e.g. 60% methanol, 30% chloroform, 10% Glacial acetic acid), neutral buffered formalin (NBF) (e.g. 10% of 37% formaldehyde solution, in a neutral pH), or a mixture thereof. Preferably, formaldehyde or disuccinimidyl glutarate (DSG) are used as discussed in Denis L. et al., Hi-C 3.0: Improved Protocol for Genome-Wide Chromosome Conformation Capture, Current Protocols, 2021.
[0120] The cells may then be permeabilized in order to gain access to the nuclei within such cells i.e. permeablization results in removal of some, or all, cellular membrane lipids from the cell walls making them more permeable. The permeablization step may be done using surfactants. Once the nucleus is accessible, it can be isolated from the cell. The isolated nuclei can then be distributed into a first set of compartments i.e. chambers, wells, or droplets. For example, distributed across a 24-well plate, or distributed across a 48-well plate, or distributed across a 96-well plate, or distributed across a 384-well plate, or distributed across a 1536-well plate.
[0121] In some embodiments, the nucleic acid sequences can be fragmented at all accessible regulatory DNA regions along such sequences. Oligonucleotide sequences can then be introduced at each accessible regulatory DNA sites.
[0122] Fragmentation is the separation or breaking of DNA or RNA strands into smaller pieces i.e. fragments. Nucleic acid fragmentation (e.g. DNA fragmentation) can be achieved by using a variety of physical, enzymatic and chemical reagents. The two most common methods of fragmentation for generating random double-stranded breaks without base bias within a nucleic acid sequence include: (i) sonication or acoustic shearing (using instrumentation), and (ii) enzyme-based fragmentation.
[0123] The oligonucleotide sequence to be introduced may be an adapter. An adapter may be a short, chemically synthesized, single-stranded or double-stranded sequence which can be ligated to the end of a DNA or RNA fragment. The adapter can be ligated to all accessible regulatory DNA regions of the nucleic acid sequence. These adapter sequences may be diverse in their sequence and comprise additional elements which enable further processing of the nucleic acid sequence or fragment into which they are inserted. In preferred embodiments, both fragmentation and adapter insertion can be performed within a single step or as part of a multi-step of the reaction. Such single-step fragmentation and adapter insertion has the advantage of providing a method with significantly fewer overall steps and reduced manipulation of the nucleic acid composition, leading to a method which is more efficient, quicker and cheaper. Once both the adapter and index sequence have bound to the DNA fragment, a first indexed DNA fragment is formed.
[0124] If performed as a single step, both fragmentation and adapter insertion can be achieved via a tagmentation process. Tagmentation is a transposome-mediated reaction that combines tagging and DNA fragmentation into a single, rapid reaction, by a so called “cut and paste” mechanism. Specifically, tagmentation will only occur at accessible regions of the DNA, so-called accessible regulatory DNA, where the chromatin is open to allow for access to the DNA sequence. Transposomes are prepared with DNA that is afterwards cut so that the transposition events result in fragmented DNA with adapters (taking away the need for separate adapter insertion). Such methods utilise a recombinase enzyme which binds to the adapter sequences and inserts these onto the fragments.
[0125] Recombinase enzymes suitable for use in the present methods include any enzyme capable of removing (or cutting) and inserting sequence into an oligonucleotide or DNA fragment. Examples of such recombinase enzymes include retroviral integrase and transposase enzymes such as MuA, Tn5, Tn7 and Tc1 / mariner-type transposases. Thus, in one embodiment of the present method, the recombinase enzyme is a retroviral integrase. In a further embodiment, the recombinase enzyme is a transposase enzyme, such as Tn5 transposase. In order for the recombinase, integrase or transposase enzyme to be active in the method presented herein, the enzyme may be mutated to overcome the naturally occurring low level of activity of such enzymes. Thus, in a yet further embodiment, the recombinase enzyme is a mutant transposase, such as a hyperactive transposase. Such a hyperactive transposase may be a mutant Tn5 transposase. In one embodiment, the recombinase is Tn5 transposase, such as hyperactive Tn5 transposase.
[0126] In another embodiment, fragmentation is performed using an endonuclease enzyme. Examples of suitable endonucleases include, but are not limited to, non-sequence specific endonucleases, such as MNase or DNase, and sequence specific endonucleases, such as restriction enzymes.
[0127] In a further embodiment, fragmentation is performed using sonication. In a preferred embodiment, fragmentation is performed using a transposase enzyme, for example, a Tn5 transposase.
[0128] In some embodiments, single-step fragmentation and adapter insertion comprises inserting an index sequence into a ligated DNA fragment. In a preferred embodiment, adapter sequences comprise index sequences.
[0129] The index sequence, also referred to as a “barcode sequence” or “barcode”, may be a short single stranded sequence comprising of between 4 and 20 bases, or comprising of between 4 and 10 bases, preferably, comprising of between 6 and 10 bases, and more preferably, the index sequence may comprise 8 bases. The index sequence may be present at any point within the adapter sequence.
[0130] Once the first indexed DNA fragments are generated, each compartment will comprise a subset of DNA fragments with the same first index. All first indexed DNA fragments may then be pooled together and randomly redistributed into a second set of compartments e.g. into wells of a plate. The redistributed first indexed DNA fragments may be distributed across a 24-well plate, or distributed across a 48-well plate, or distributed across a 96-well plate, or distributed across a 384-well plate, or distributed across a 1536-well plate. The number of wells required in each plate may be easily determined by consideration of the number of cells in the population to be distributed and the availability of plates, in order that each cell or isolated nucleus is individually indexed.
[0131] A second adapter with a second index sequence can then be added to the redistributed first indexed DNA fragments. This may be done via proximal bridge ligation, wherein proximal first adapters of the first indexed DNA fragments are joined by second adapters comprising a second index to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes.
[0132] Proximal bridge ligation involves a second adapter binding between two first adapters within the DNA fragment creating a ‘bridge’ between the two first adapters that are located in close proximity. Bridge adapters add the second index and also ligate the two interacting regions together, through the first adapters. The bridge adapter can have sticky ends (of varying lengths) that ligate specifically to the sticky ends of the first adapters or they can be blunt ligated. In sticky ends, one strand is longer than the other (typically by at least a few nucleotides), such that the longer strand has bases which are left unpaired. In blunt ends, both strands are of equal length - i.e. they end at the same base position, leaving no unpaired bases on either strand. Bridge ligation creates a bridging sequence between two genomic regions of interest, in the present method the bridging sequence may comprise in order: a first adapter incorporating a first index - a second adapter incorporating a second index - a first adapter incorporating a first index. The addition of this second adapter during the bridge ligation step allows for incorporating a second index in conjunction with capturing structural and regulatory chromatin interactions. However, if combinatorial indexing is not required, proximal ligation of the first adapter is sufficient to capture the structural and regulatory chromatin interactions, improving the sensitivity and specificity of active chromatin detection in order to reveal important DNA regulatory interactions.
[0133] References to “ligation” refer to any linkage of two DNA fragments usually comprising a phosphodiester bond. The linkage is normally facilitated by the presence of a catalytic enzyme (i.e. for example, a ligase such as T4 DNA ligase) in the presence of co-factor reagents and an energy source (i.e. for example, adenosine triphosphate (ATP)). In the methods described herein, the DNA is cross-linked then fragmented with first adaptors added to all accessible regulatory DNA. Proximal ligation or proximal bridge ligation then selects for first adapters that are in close proximity and ligates these first adapters together to form a single DNA fragment. In proximal bridge ligation one or more second adapter, or bridging adapter, is added to ligate the first adapters together. Cross-linking of DNA regulatory interactions ensures that these interacting regions are fixed and stay in close proximity for proximal ligation of the two interacting regions.
[0134] In certain embodiments, the second adapter may comprise a marker, such as biotin, histidine (i.e. 6H is), or FLAG. The marker may be the first half of a binding pair moiety. This marker may be introduced as part of the first or second adapter. In a preferred embodiment, the second adapter comprises biotin as the first half of a binding pair moiety. The addition of biotin within the second adapter sequence allows for the subsequent selection, or enrichment, of DNA fragments which have been bridge ligated according to step (e) and the methods as defined herein. The marking or addition of the first half of a binding pair allows bridge ligated fragments to be purified according to enrichment step (e), therefore ensuring that only bridge ligated fragments are enriched, rather than non-ligated (i.e. non-interacting) fragments. For example, where the first half of the binding pair moiety is biotin, the second half of the binding pair moiety may be streptavidin and the marked DNA fragments may be pulled down or enriched by streptavidin bead purification. In certain embodiments, the DNA fragments have their ends filled by the addition of nucleotides to the 3’ end, such filling comprises the addition of dATP, dCTP, dGTP and / or dTTP nucleotides to the 3’ end of the nucleic acid composition or segments. In one embodiment one or more of the nucleotides used for filling the ends may comprise a marker, for example, the first half of a binding pair moiety, such as biotin. Thus, in one embodiment, filling the ends of cross-linked DNA fragments comprises the addition of a marker to the crosslinked DNA fragments.
[0135] The generated indexed or second indexed DNA fragments can then be pooled together a third time and randomly redistributed into a third set of compartments e.g. into wells of a plate. The redistributed second indexed DNA fragments may be distributed across a 24-well plate, or distributed across a 48-well plate, or distributed across a 96-well plate, or distributed across a 384-well plate, or distributed across a 1536-well plate; all depending on the number of second indexed DNA fragments within the sample.
[0136] In a further embodiment, a different index sequence will be added to the DNA fragments redistributed into each of the first and second compartments (as defined in steps (b) and (c) of the method disclosed). The addition of more than one index sequence into the DNA fragment, as well as the pooling and redistribution of said DNA fragments between binding each index sequence greatly lowers the chance of two cells having the same indexing combination. This will result in a more reliable and efficient identification of differing cell types during the multiplex sequence analysis of the disclosed methods. This aids in the dissection of interactions in a multi-cell sample without the need to purify specific cell types. Any number of indexes can be used depending on the size of cell sample and the decrease in complexity required.
[0137] Once all adapters and index sequences have been added and bound to the DNA fragments, the uniquely indexed DNA fragments can then be purified. This purification involves reversal of the cross-linking within the DNA. Once the cross-links have been reversed, the DNA can be amplified and sequenced. Preferably, the cross-linking is reversed before the DNA fragments are enriched, preferably by biotin selection. It will be understood that there are several ways known in the art to reverse crosslinks and it will depend upon the way in which the crosslinks are originally formed. For example, crosslinks may be reversed by subjecting the cross-linked nucleic acid composition to high heat, such as above 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, or greater. Furthermore, the cross-linked nucleic acid composition may need to be subjected to high heat for longer than 1 hour, for example, at least 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours or 12 hours or longer. In one embodiment, reversing the cross-linking comprises incubating the cross-linked nucleic acid composition at 65°C for at least 8 hours (i.e. overnight) in the presence of Proteinase K. Purification may also refer to the removal of various other components, for example, by fractionation, while retaining the desired DNA fragments.
[0138] Gaps, or nicks, in the DNA that are generated during the tagmentation process, for example when using T n5, may be filled in with a DNA polymerase and ligase to seal the gaps. A preferred method for introducing one or more nucleotides comprising a first half of a binding pair moiety is by inclusion when the gaps are filled.
[0139] Amplification of the uniquely indexed DNA fragments may be performed using polymerase chain reaction (PCR). PCR is a technique involving enzymatic amplification of nucleic acid sequences via repeated cycles of denaturation, oligonucleotide annealing, and DNA polymerase extension. The DNA fragments which are indexed (following the methods disclosed) may be amplified using PCR prior at analysis. PCR is a very sensitive technique that allows rapid amplification of such DNA fragments, resulting in billions of copies of the same DNA fragment allowing for easier and more efficient detection and identification.
[0140] In some embodiments, sequencing adapters may be added to the purified DNA (eg. by tagmentation or fragmentation and covalent attachment of adapters) prior to amplification. In a preferred embodiment, the sequencing adapters are paired end adapters, a primer pair set that allows automated high throughput sequencing to read from both ends. For example, such high throughput sequencing devices that are compatible with these adapters include, but are not limited to Solexa (Illumina), the 454 System, and / or the ABI SOLiD. A third adapter and third index sequence can then be added to the purified DNA in the form of sequencing adapters and / or through amplification with indexed PCR primers. High throughput sequencing is also known as next generation sequencing (NGS).
[0141] In some embodiments the adapter sequences compatible for sequencing can be incorporated using a microfluidics device such as the 10x, where each cell / nuclei has been compartmentalised and unique index / barcodes can be introduced to the DNA of each isolated cell / nuclei, for example using the 10x ATAC or multiome kits. This may be done in combination with the first and / or second indexing steps to increase indexing complexity, or this may be done as the sole indexing step, depending on the size of the cell sample. The purified DNA fragments can then be pooled together for multiplex sequence analysis and for generating datasets. These datasets can then be used to identify the DNA regulatory interactions relating to a single cell from the cell population. Any sequence bearing the same combination of indexes can then be deduced to have derived from a single cell. Sequencing of the DNA fragments provides sequence reads where the two DNA sequences that were originally cross-linked as DNA regulatory interactions, for example, a promoterenhancer interaction, are provided in a single read. These hybrid reads can be decomposed into the two genomic sequences either side of the first adapters or bridge adapter sequence and aligned against a reference genome to produce a map of pairwise contacts between accessible regulatory DNA or open chromatin regions. This produces a high-dimensional map of DNA regulatory interactions or chromatin contacts across the genome for each cell in the sample. Dimensionality reduction by techniques such as uniform manifold approximation and projection (LIMAP) or t-distributed stochastic neighbour embedding (t-SNE) allows clustering of cells by population, for example, using methods developed for single cell ATAC-seq (for example, SnapATAC by Fang et al., 2021).
[0142] Once cell type clusters are identified, data from multiple cells in the cluster can be combined and significant contacts can be identified by identifying pairs of open chromatin regions which are in contact with high frequency after normalisation, for example by the techniques used in model-based analysis of PLAC-seq and HiChIP (Juric et al., 2019), or random forest classification techniques developed for a broad range of contact mapping assays (Salameh et al., 2020). Once the significant contacts are identified, it is possible to use a cell type contact map Atlas to identify the specific cell types within the sample. This would include using known profiles from cell Atlas and comparing to the interaction profiles measured from the cell sample obtained.
[0143] In certain embodiments, the DNA regulatory interactions within the heterogenous cell populations can be identified which are indicative of a particular disease state. This may be performed by quantifying a frequency of certain DNA regulatory interactions within an identified cell and then comparing the frequency of those DNA regulatory interactions from an individual with said disease state with the frequency of those DNA regulatory interaction in a control cell sample from a healthy subject. The difference in the frequency of the particular DNA regulatory interactions in the identified single cell compared to the healthy subject’s cell sample may be indicative of a particular disease. For example, high throughput sequencing results can also enable examination of the frequency of a particular interaction. In some instances, a lower frequency of particular DNA regulatory interactions in the identified single cells, compared to the healthy subject’s cell sample, is indicative of a particular disease state (i.e. because the DNA fragments are interacting less frequently). Alternatively, a higher frequency of interaction in the identified single cells, compared to the healthy subject’s cell sample, is indicative of a particular disease state (i.e. because the DNA fragments are interacting more frequently). In some instances, the difference will be represented by at least a 0.5-fold difference, such as a 1-fold, 1.5-fold, 2-fold, 2.5-fold, 3-fold, 4-fold, 5-fold, 7-fold or 10-fold difference. The set of interactions which are different between healthy and disease may constitute a disease profile. The frequency of DNA regulatory interactions between the identified single cells in a sample may be compared to a disease profile. Samples may be indicative of disease if a certain proportion of interactions have contact frequencies matching the disease profile. In some instances, a sample may be indicative of disease if at least 50% of profile interactions are similar to the disease profile. In a preferred embodiment, the frequency of DNA regulatory interactions in the sample would be compared to a disease profile through the action of a classifier, for example by logistic regression.
[0144] Sequences
[0145] * denotes the index sequence placement within the adapter sequence. The index sequence may be placed anywhere within the adapter sequence.
[0146] Bold nucleic acid bases (e.g. T) denotes that such base can be biotinylated. Certain aspects and embodiments of the invention will now be illustrated by way of example and with reference to the Figures and Examples.
[0147] Definitions
[0148] References to “DNA”, “deoxyribonucleic acid”, “DNA sequences” or “DNA fragments” generally refer to any polymer of nucleotides (i.e. for example, adenine (A), thymidine (T), cytosine (C), guanosine (G), and / or uracil (II)), comprised in a nucleic acid sequence.
[0149] References to “restriction enzyme” as used herein, refers to any protein that cleaves nucleic acid at a specific base pair sequence. Cleavage can result in a blunt or sticky end, depending on the type of restriction enzyme chosen. Examples of restriction enzymes include, but are not limited to, Eco Rl, Eco RII, Bam HI, Hind III, Dpn II, Bgl II, Neo I, Taq I, Not I, Hinf I, Sau 3A, Pvu II, Sma I, Hae III, Hga I, Alu I, Eco RV, Kpn I, Pst I, Sac I, Sal I, Sea I, Spe I, Sph I, Stu I, Xba I. The resulting DNA sequences are referred to as “DNA fragments”.
[0150] References to “cross-linking” or “cross-link” as used herein, refer to any stable chemical association between two regions within the nucleic acid sequence, such that they may be further processed as a unit. Such stability may be based upon covalent and / or non- covalent bonding (e.g. ionic). For example, nucleic acids and / or proteins may be cross-linked by chemical agents (i.e. for example, a fixative, fixing agent or cross-linking agent), heat, pressure, change in pH, or radiation, such that they maintain their spatial relationships during routine laboratory procedures (i.e. for example, extracting, washing, centrifugation etc.). Cross-linking as used herein is equivalent to the terms “fixing” or “fixation”, which applies to any method or process that immobilises any and all cellular processes. A cross-linked / fixed cell, therefore, accurately maintains the spatial relationships between components within the nucleic acid sequence at the time of fixation.
[0151] References to the term “fragments” or “DNA fragments” as used herein, refers to any nucleic acid sequence that is shorter than the sequence from which it is derived. Fragments can be of any size, ranging from several megabases and / or kilobases to only a few nucleotides long. Fragments are suitably greater than 5 nucleotide bases in length, for example 10, 15, 20, 25, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 5000 or 10000 nucleotide bases in in length. Fragments may be even longer, for example 1 , 5, 10, 20, 25, 50, 75, 100, 200, 300, 400 or 500 nucleotide kilobases in length. Methods such as tagmentation, restriction enzyme digestion, sonication, acid incubation, base incubation, microfluidization etc., can all be used to fragment a nucleic acid composition. As referenced herein, an “index sequence” can also be referred to as a “barcode sequence”, or as a “barcode” or an “index” or a unique molecular identifier (UMI). An index sequence is a short nucleotide sequence. Such index sequences may allow for the identification of a particular nucleic acid sequence in subsequent analysis and processing. The index sequence may be comprised within an adapter sequence or a PCR primer. The index sequence may be formed of between 4 and 10 nucleotide bases, preferably 8 nucleotide bases.
[0152] References to “accessible chromatin regions” can also be referred to as accessible or open DNA regions or sites, for example, “hyper accessible regions”, “hyper accessible DNA regions”, “hyper accessible chromatin”, “open chromatin”, “open chromatin regions (OCRs)”, “euchromatin” etc.. Chromatin accessibility is the degree to which nuclear macromolecules are able to physically contact chromatinized DNA and is determined by the occupancy and topological organisation of nucleosomes as well as other chromatin-binding factors that occlude access to the DNA. These accessible chromatin regions may be regions within the sequence of the DNA that can be contacted by DNA regulatory elements, for example, transcription factors. The regulatory elements may bind to the DNA at these accessible regions for regulation of gene expression. Accessible chromatin regions also form where regulatory sequences or regions, such as enhancers and promoters, interact. Therefore, the accessibility of chromatin affects the gene expression of tissue cells and these regions are considered to be sites of gene regulation.
[0153] References to “DNA regulatory interactions” refers to DNA-DNA interactions that occur between regulatory sequences or regions, such as between enhancers and other regulatory elements such as repressors or insulators; or between promoters and enhancers; or enhancers and enhancers; or promoters and promoters. The interactions may cause one interacting sequence to have an effect upon the other, for example, silencing or activating the sequence it binds to. The interaction may occur between two nucleic acid sequences that are located close together or far apart on the linear genome sequence. These DNA sequences are brought into close proximity by the 3D organisation of the genome, such as DNA-looping mechanisms. Chromatin loops are an example of DNA sequences in close proximity, where stretches of genomic sequence that lie on the same chromosome are in closer physical proximity to each other than to intervening sequences. In particular, DNA regulatory interactions may include promoter to enhancer interactions, enhancer to enhancer interactions, and promoter to promoter interactions. DNA regulatory interactions may be cross- linked in close proximity by use of a fixative, fixative agent or cross-linking agent. Sites or regions of DNA regulatory interaction form interacting regions of two DNA sequences which may be cross-linked due to their proximity. Such interacting regions may consist of accessible chromatin, referred to according to the present invention as “accessible regulatory DNA”.
[0154] Heterogenous cell population refers to a composition of cells comprising different cell types and / or same cell types at different stages of development, stages of cell cycle or different cell states.
[0155] References to “frequency of DNA regulatory interaction”, or “DNA regulatory interaction frequency”, “frequency of DNA interaction”, or “DNA interaction frequency” as used herein, refer to the number of times a specific interaction occurs within a nucleic acid sequence (i.e. within the cell sample).
[0156] References herein to an “autoimmune disease” include conditions which arise from an immune response targeted against a person’s own body, for example Acute disseminated encephalomyelitis (ADEM), Amyotrophic Lateral Sclerosis (ALS), Ankylosing Spondylitis, Behget's disease, Celiac disease, Crohn's disease, Diabetes mellitus type 1 , Graves' disease, Guillain-Barre syndrome (GBS), Multiple Sclerosis (MS), Psoriasis, Rheumatoid arthritis, Rheumatic fever, Sjogren's syndrome, Ulcerative colitis and Vasculitis.
[0157] References herein to a “developmental disorder” include conditions, usually originating from childhood, such as learning disabilities, communication disorders, Autism, Attentiondeficit hyperactivity disorder (ADHD) and Developmental coordination disorder.
[0158] References herein to a “genetic disorder” include conditions which result from one or more abnormalities in the genome, such as Angelman syndrome, Canavan disease, Charcot- Marie-Tooth disease, Colour blindness, Cri du chat syndrome, Cystic fibrosis, Down syndrome, Duchenne muscular dystrophy, Haemochromatosis, Haemophilia, Klinefelter syndrome, Neurofibromatosis, Phenylketonuria, Polycystic kidney disease, Prader-Willi syndrome, Sickle-cell disease, Tay-Sachs disease and Turner syndrome.
[0159] References to “sticky ends” or “non-blunt ends” used herein, refer to the creation of various overhangs of a nucleic acid sequence. An overhang is a stretch of unpaired nucleotides in the end of a DNA molecule. These unpaired nucleotides can be in either strand, creating either 3' or 5' overhangs. These overhangs are in most cases palindromic.
[0160] References to “blunt ends”, also known as “non-cohesive ends”, refer to a blunt-ended molecule, such as a blunt-ended DNA molecule, wherein both strands terminate in a base pair.
[0161] References herein to cell sorting include separating or isolation individual cells or cell types or cell sub-types from a heterogenous cell population. It will be appreciated that the cells, cell population or heterogenous cell population may be from a range of organisms, not just humans. For example, the present method may also be used to identify DNA regulatory interactions in plants and animals.
[0162] It will be appreciated that the terms “isolating”, “isolation”, “separating”, “removing”, “purifying” can be used interchangeably.
[0163] Therefore it will be appreciated that the present invention provides methods for identifying DNA regulatory interactions between accessible chromatin regions, of which some methods may identify DNA regulatory interactions indicative of a particular disease state or treatment state, and a kit for identifying DNA regulatory interactions.
[0164] All references, patents, and publications cited herein are hereby incorporated by reference in their entirety.
[0165] EXAMPLES
[0166] Example 1
[0167] This protocol was initially conducted without combinatorial indexing (multiplexing) as a bulk method, however the addition of indexing adapters in a single cell protocol is envisaged. According to this bulk processing method a consensus analysis is obtained for the cell population in the sample. For a single experiment a sample comprising 50,000 to 500,000 cells may be prepared.
[0168] 1) Cell fixation to crosslink the interactions
[0169] • For a single experiment 200,000 cells are fixed for 10 minutes at room temperature at a final formaldehyde concentration of 0.5 %.
[0170] • Quench the formaldehyde with glycine at a final concentration of 0.2 M for 10 minutes. • Collect the fixed cells by centrifugation and discard the supernatant and re-suspend the pellet in ice-cold PBS supplemented with 1 / 50th cell culture medium to wash the cells.
[0171] • Collect the washed cells by centrifugation and discard the supernatant.
[0172] • Snap-freeze the pellet by immersing the tube in liquid nitrogen for at least 5 minutes followed by storage at -80°C or proceed directly to the next step.
[0173] 2) Cell Permeabilisation
[0174] • To each 200K of fixed cells, add 0.4 mL of ice-cold 1X Lysis Buffer (1X PBS, 5% BSA, 1 mM DTT, 0.2 % IGEPAL) and incubate on ice for 15 minutes to permeabilise the cells.
[0175] • Collect the cells by centrifugation and remove the supernatant without disrupting the pellet, leaving ~20 uL behind.
[0176] 3) Tagmentation (Introduction of adapter in open chromatin, with the first barcode) For multiplexing the cells will be distributed into wells to add the different barcodes.
[0177] • Using the adapter oligos listed and Tn5 Transposase and buffers (Creative Enzymes, NATE1629), set up the tagmentation at 37°C for 1 hour with 800 rpm mixing. 5 uL of the transposome is required in 100 uL reaction volume.
[0178] • Stop the reaction by adding SDS to a final concentration of 0.33% for 1 hour at 37°C with 800 rpm mixing.
[0179] For multiplexing the cells will be repooled.
[0180] • Add Triton X-100 to a final concentration of 1.8% and incubate for 10 min at 37°C with 800 rpm mixing.
[0181] • Centrifuge to collect the pellet and remove the supernatant, leaving ~20 pL behind.
[0182] 4) Ligate the Bridge Adapter (Bridge adapter ligates to the transposome adapters bringing the interacting DNA regions together and adding a second barcode)
[0183] • Wash the pellets with 200 pL 1X ligase buffer.
[0184] • Remove the wash supernatant leaving -20 pL behind.
[0185] For multiplexing the cells will be distributed into wells to add the different barcodes.
[0186] » Set up a 1 mL ligation, in 1X ligase buffer with 10 uL of 50 uM annealed bridge adapters and 10 uL T4 ligase (ThermoFisher, EL0011).
[0187] • Incubate overnight at 16°C.
[0188] 5) Fill the gap using T4 DNA Polymerase and Ligase
[0189] • Centrifuge to collect the pellet and remove the supernatant leaving -100 pL behind.
[0190] • Add dNTPs to a final concentration of 200 uM each nucleotide and 4 uL of T4 DNA polymerase (ThermoFisher, EP0062) in 1X ligase buffer in a total volume of 180 uL
[0191] • Incubate for 30 minutes at 11 °C with 800 rpm mixing.
[0192] • Add 10 pL of 2X Ligase Buffer and 10 uL of T4 ligase (ThermoFisher, EL0011).
[0193] • Incubate at 16 °C for 2 hours.
[0194] • Stop the reaction by adding EDTA to 25 mM final concentration.
[0195] For multiplexing the cells will be repooled, mixed and distributed into wells to add the different barcodes.
[0196] 6) Purification of DNA
[0197] • Add 30 pL of 10 mg / mL Proteinase K and pipette up and down to mix.
[0198] • Incubate for 2 hrs at 65°C.
[0199] • Purify with 1X volume of AM Pure beads (Beckman Coulter, A63881) following the manufacturers instructions and elute the DNA from the beads with 31 pL of nuclease free water for 15 min at room temperature.
[0200] • Quantify the DNA by Qubit 1X dsDNA, high sensitivity assay (ThermoFisher Q33231). 7) Taqment (adds on the Illumina adapters for sequencing)
[0201] • Using a 1 : 1 mix of adapter A and B and the Tn5 T ransposase and buffers (Creative Enzymes, NATE1629), set up the tagmentation in 40 uL total volume at 55 °C for 15 minutes. Titration of the transposome batch is required before use, aiming at a final library size of ~650bp to 750bp.
[0202] • Stop the reaction by adding SDS to a final concentration of 0.04% at 55 °C for 7 min.
[0203] 8) Pull down of Biotinylated DNA Fragments
[0204] • Bind the tagmented DNA to 40 pL of MyOne C1 Dynabeads beads per sample (ThermoFisher, 65002), following the manufacturer’s instructions, incubating for 30 minutes at room temperature with 1300 rpm mixing.
[0205] • Wash the beads coated with biotinylated DNA according to manufacturer’s instructions, with 200 uL of wash buffers, resuspending the beads in the wash buffer by vortexing. After the final wash, resuspend the beads in 20 uL of nuclease free water.
[0206] 9) PCR (Amplifies the DNA for sequencing and adds a third barcode for multiplexing)
[0207] To the 20 uL of purified beads, add 0.6 uM of each i5 and i7 primers, 0.3 mM each dNTP and 2 uL of Kapa HiFi polymerase, in 1X Kapa HiFi fidelity buffer (Roche, 07958846001) in a total reaction volume of 100 uL. Thermal cycler programme
[0208] For multiplexing the cells will be repooled.
[0209] 10) Purification of DNA
[0210] * Separate the amplified sample on a magnet. Collect and transfer the supernatant to a fresh tube.
[0211] • Purify the supernatant with 0.7X or 1X volume of AMPure beads (Beckman Coulter, A63881) following the manufacturer’s instructions and elute the DNA from the beads with 20 pL of 10 mM Tris pH 8.0 for 5 minutes at room temperature.
[0212] 11) Library QC and Sequencing
[0213] • Quantify the DNA by Qubit 1X dsDNA, high sensitivity assay (ThermoFisher Q33231).
[0214] • Quantify the DNA and determine the average library size using the Bioanalyzer High Sensitivity DNA Kit (Agilent, 5067-4626).
[0215] • Sequence using NextSeq 2000 P2 flow cell (2X 300bp reads) (Illumina, 20075295).
[0216] Further examples conducted according to the present invention include combinatorial indexing (96x96x96 barcodes) on a mixture of Human cultured Jurkat cells and Mouse cultured 3T3 cells to identify single cell interaction profiles. From the sequencing data and analysis, the two cell types should be able to be distinguished from each other; the interactions from each cell grouping into two different populations. The number of cells that by chance have the same barcode combination can be found, to determine if the number of cells relative to barcodes is optimal. By mixing the cells at different ratios, the sensitivity of the assay to low abundant cell types can be determined. Combinatorial indexing on a blood sample will identify single cell interaction profiles in a heterogeneous sample and single cell interaction data will be used to identify populations of cells in the heterogeneous sample.
[0217] Example 2
[0218] Single Ceil Protocol Including Combinatorial Barcoding
[0219] This protocol was conducted with combinatorial indexing (multiplexing) (96x96x96 barcodes) on a 1 : 1 mixture of Human cultured Jurkat cells and Mouse cultured 3T3 cells to identify single cell interaction profiles. From the sequencing data and analysis, the two cell types can be distinguished from each other and the number of cells that by chance have the same barcode combination has been determined.
[0220] 1) Cell fixation to crosslink the interactions
[0221] • For a single experiment use 50,000 mouse NIH 3T3 cells and 50,000 human Jurkat cells. Fix the cells separately for 10 minutes at room temperature at a final formaldehyde concentration of 0.5 %.
[0222] • Quench the formaldehyde with glycine at a final concentration of 0.2 M for 10 minutes.
[0223] • Collect the fixed cells by centrifugation and discard the supernatant and re-suspend the pellet in ice-cold PBS supplemented with 1 / 50th cell culture medium to wash the cells.
[0224] • Collect the washed cells by centrifugation and discard the supernatant.
[0225] • Snap-freeze the pellet by immersing the tube in liquid nitrogen for at least 5 minutes followed by storage at -80°C or proceed directly to the next step.
[0226] 2) Cell Permeabilisation
[0227] • To 50K of 0.5% fixed cells, add 100 uL of ice-cold 1X Lysis Buffer (1X PBS, 5% BSA, 1 mM DTT, 0.2 % IGEPAL). Mix 50K of fixed human Jurkat cells and 50K of fixed mouse NIH 3T3 cells (1 :1 mix) and incubate on ice for 15 minutes to permeabilise the cells.
[0228] • Collect the cells by centrifugation and remove the supernatant without disrupting the pellet, leaving -100 uL behind.
[0229] 3) Tagmentation (Introduction of adapter in open chromatin, with the first barcode) For multiplexing the cells are distributed into wells to add the different barcodes.
[0230] 96 different transposomes were prepared, each with a different index.
[0231] • Using the adapter oligos listed and Tn5 Transposase and buffers (Creative Enzymes, NATE1629), 96 transposomes are prepared, according to manufacturer’s instructions.
[0232] • 100 uL of cells are diluted to 800 uL in transposome buffer (Creative Enzymes NATE1629) and dispense 8 uL to each well of a 96 well plate.
[0233] • 2 uL of the prepared transposomes is added per well (one unique index per well) and incubated at 37°C for 1 hour with 1000 rpm mixing.
[0234] • Stop the reaction by adding SDS to a final concentration of 0.33% for 1 hour at 37°C with 1000 rpm mixing.
[0235] • Pool together the 96 wells into a single tube.
[0236] • Add Triton X-100 to a final concentration of 1.8% and incubate for 10 min at 37°C with 1000 rpm mixing.
[0237] • Centrifuge to collect the pellet and remove the supernatant, leaving ~20 pL behind.
[0238] 4) Ligate the Bridge Adapter (Bridge adapter ligates to the transposome adapters bringing the interacting DNA regions together and adding a second barcode)
[0239] 96 different bridge adapters are prepared, each with a different index.
[0240] • Wash the pellets with 400 pL 1X ligase buffer.
[0241] • Remove the wash supernatant leaving ~20 pL behind.
[0242] » Set up a 1 mL ligation, in 1X ligase buffer and 50 units T4 DNA ligase (ThermoFisher, EL0011).
[0243] • Aliquot 10 uL of the cell suspension per well of a 96 well plate.
[0244] » Add 2 uL of 2.5 uM annealed bridge adapters (each well with a different index).
[0245] • Incubate overnight at 16°C.
[0246] 5) Fill the gap using T4 DNA Polymerase and Ligase • Set up a 200 uL mastermix adding dNTPs to a final concentration of 66.7 uM each nucleotide and 20 units of T4 DNA polymerase (ThermoFisher, EP0062) in 1X ligase buffer. Add 2 uL per well.
[0247] • Incubate for 30 minutes at 11 °C with 1000 rpm mixing.
[0248] • Add 0.5 units of T4 DNA ligase per well.
[0249] • Incubate at 16 °C for 2 hours.
[0250] • Stop the reaction by adding EDTA to 25 mM final concentration.
[0251] • Pool together the 96 wells into a single tube.
[0252] • Mix thoroughly and dispense equally into 96 wells of a fresh plate. ) Purification of DNA
[0253] • Add 4 ug of Proteinase K (Merck, 3115836001) per well.
[0254] • Incubate for 2 hrs at 65°C at 1000 rpm.
[0255] • Purify with 1X volume of AM Pure beads (Beckman Coulter, A63881) following the manufacturer’s instructions and elute the DNA from the beads with 15 pL of nuclease free water per well. ) Tagment (adds on the Illumina adapters for sequencing)
[0256] • Using a 1 : 1 mix of adapter A and B and the Tn5 T ransposase and buffers (Creative Enzymes, NATE1629), set up a tagmentation in 20 uL total volume at 55 °C for 15 minutes. Titration of the transposome batch is required before use, aiming at a final average library size of ~750bp.
[0257] • Stop the reaction by adding SDS to a final concentration of 0.04% at 55 °C for 7 min. ) Pull down of Biotinylated DNA Fragments
[0258] • Bind the tagmented DNA to MyOne C1 Dynabeads (ThermoFisher, 65002), following the manufacturer’s instructions. Wash the beads coated with biotinylated DNA according to manufacturer’s instructions. After the final wash, resuspend the beads in 5 uL of nuclease free water per well.
[0259] 9) PCR (Amplifies the DNA and adds the third barcode for multiplexing)
[0260] 96 different pairs of PCR primers were prepared, each with a different index.
[0261] • Set up a 10 uL PCR per well using 5 uL of purified beads, i5 and i7 PCR primers and the Kapa HiFi PCR Kit (Roche, 07958846001), according to manufacturer’s instructions. Each well of the 96 well plate is amplified with different indexed PCR primers pairs.
[0262] Thermal cycler programme
[0263] 10) Purification of DNA
[0264] • Pool all the wells of the PCR together into an Eppendorf tube.
[0265] • Separate the amplified sample on a magnet. Collect and transfer the supernatant to a fresh Eppendorf tube. • Purify the supernatant with 0.7X voiume of AM Pure beads (Beckman Coulter, A63881) following the manufacturer’s instructions and eluting the DNA from the beads with 50 pL of 10 mM Tris pH 8.0.
[0266] 11) Library QC and Sequencing
[0267] • Quantify the DNA by Qubit 1X dsDNA, high sensitivity assay (ThermoFisher Q33231).
[0268] • Quantify the DNA and determine the average library size using the Bioanalyzer High Sensitivity DNA Kit (Agilent, 5067-4626).
[0269] • Sequence using NextSeq 2000 P1 flow cell (2X 300bp reads) (Illumina, 20075294).
[0270] Example 3
[0271] Demonstration of Single Cell identification
[0272] To demonstrate that results describe individual cells the assay was conducted on a mixture of 50,000 Jurkat (Human) cells, and 50,000 NIH-3T3 (Mouse) cells. As each datapoint can be unambiguously assigned to a species this allows a measurement of the doublet rate (i.e. the rate at which more than one cell is assigned the same barcode).
[0273] Single Cell Protocol
[0274] The single cell protocol was conducted as per Example 2, with the protocol described for multiplexing (combinatorial barcoding).
[0275] Bioinformatic Processing
[0276] The results of library sequencing were analysed in the following steps:
[0277] 1) Cleanup a. Low quality reads were removed, and poly-G tails trimmed, using the tool FastP (Citation: Shifu Chen. 2023. Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp. iMeta 2: e107. https: / / doi.Org / 10.1002 / imt2.107)
[0278] 2) Read splitting at barcode a. The barcode sequence of each read was identified by the presence of the flanking Tn5 mosaic sequence. b. The sequences upstream and downstream of the barcode were extracted, and written to bam format with the barcode sequence appended to each read identifier. This produces two outputs: the set of reads upstream of the barcode, and the set of reads downstream of the barcode, which can be written into the bam file format expected by the HiCUP mapper tool.
[0279] 3) Alignment to reference a. The paired upstream and downstream sequences were aligned to the human and mouse reference genomes. Any reads which could not be unambiguously assigned (due to sequence similarity between species) were removed. b. This was performed using the tool HiCUP mapper (citation: Wingett S, et al. (2015) HiCUP: pipeline for mapping and processing Hi-C data FIOOOResearch, 4:1310 (doi: 10.12688 / f1000research.7334))
[0280] 4) Deduplication a. Duplicate reads (reads with the same index and barcode combination, and the same aligned location) were removed.
[0281] Analysis
[0282] Following alignment to the reference each pair of genomic locations is tagged by the inserted barcode sequence and Illumina sequencing adaptors. This combination is expected to uniquely identify a cell. In order to measure the doublet rate we filter for high quality cells (i.e. those with more than 20 unique, mappable location pairs which can be unambiguously assigned as mouse or human derived) and test whether all of the pairs within a cell can be uniquely assigned to either species. As the probability of a doublet barcode (i.e. a barcode with both a mouse and a human cell) reporting at least 20 reads being measured as a single species by chance is less than 1 / 1000thof 1 % this is a measurement of the doublet rate in the assay.
[0283] Of the high quality cells, 96% are reported to describe a single species (i.e. they fall on the axes of the graph of figure 9). Doublet barcodes are equally likely to describe the same as mixed species, so this leads to an estimated doublet rate of 8%. This is consistent with theoretical calculations, which suggest that multiplexing 884,736 barcodes over 100,000 cells should produce a doublet rate of 4.5% if sampling were perfectly random. This is a substantial improvement over common single cell ATAC-seq protocols, for which doublet rates of over 10% are reported in equivalent experiments (Fang, R., Preissl, S., Li, Y. et a / . Comprehensive analysis of single cell ATAC-seq data with SnapATAC. Nat Commun 12, 1337 (2021). https: / / doi.Org / 10.1038 / S41467-021-21583-9). In summary, this experiment demonstrates that the method according to the present invention can successfully resolve single cells in experiments with large cell numbers (100,000) and accuracy better than current ATAC-seq protocols.
[0284] Example 4
[0285] Demonstration that Human Cell Mixtures can be Deconvolved
[0286] In order to demonstrate that the single cell data described above can be used to deconvolve a mixture of human cells the assay was performed on 50,000 Jurkat (human lymphocyte cell line) and 50,000 IMR90 (human fibroblast cell line) cells.
[0287] S / nq / e Cell Protocol
[0288] The single cell protocol was conducted as per Example 2, with the protocol described for multiplexing (combinatorial barcoding).
[0289] Bioinformatic Processing
[0290] The results of library sequencing were analysed in the following steps:
[0291] 1) Cleanup a. Low quality reads were removed, and poly-G tails trimmed, using the tool FastP (Citation: Shifu Chen. 2023. Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp. iM eta 2: e107. https: / / doi.Org / 10.1002 / imt2.107)
[0292] 2) Read splitting at barcode a. The barcode sequence of each read was identified by the presence of the flanking Tn5 mosaic sequence. b. The sequences upstream and downstream of the barcode were extracted, and written to bam format with the barcode sequence appended to each read identifier.
[0293] 3) Alignment to reference a. Each component (upstream or downstream) sequence was aligned to the human reference. b. This was performed using the tool BWA-MEM (citation: Li H. (2013); arXiv:1303.3997v2) to align sequencing reads.
[0294] 4) Deduplication a. Duplicated reads were removed using the Picard toolkit (citation: http: / / broadinstitute.github.io / picard / )
[0295] 5) Identification of High Read Count Cells a. The number of unique, mappable reads for each cell was counted, and those cells with at least 200 reads retained for downstream analysis
[0296] Analysis
[0297] 1) Identification of cell-specific open chromatin regions a. A standard ATAC-seq library of Jurkat cells was prepared and processed. Peaks, corresponding to regions of open chromatin, were called using MACS2 (citation: Zhang, Y., Liu, T., Meyer, C.A. et a / . Model-based Analysis of ChlP-Seq (MACS). Genome Biol 9, R137 (2008). https: / / doi.Org / 10.1186 / gb-2008-9-9-r137) b. Similarly, published ATAC-seq data describing IMR90 cells was obtained (citation: Luo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K, Baymuradov UK, Graham K, Litton C, Miyasato SR, Strattan JS, Jolanki O, Lee JW, Tanaka FY, Adenekan P, O'Neill E, Cherry JM. New developments on the Encyclopedia of DNA Elements (ENCODE) data portal. Nucleic Acids Res. 2020 Jan 8;48(D1):D882-D889. doi: 10.1093 / nar / gkz1062. PM ID: 31713622; PMCID: PMC7061942.) c. An inverse intersection of the called open chromatin regions was performed using Bedtools (citation: Quinlan AR and Hall IM, 2010. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 26, 6, pp. 841-842.). This produced two sets of regions, corresponding to regions of chromatin which are open in IMR90 cells but not Jurkat, and vice- versa.
[0298] 2) Identification of single cell types a. For each cell, reads falling into Jurkat-associated and IMR90-associated regions were counted. Cells with more reads in Jurkat-associated regions were tentatively assigned to Jurkat type, and similarly for IMR90. b. The observed ratio of Jurkat-associated to IMR90-associated reads over all cells was used to generate a binomial model. This was used to calculate a p-value for the observed ratio of reads in each cell. c. The p-values were corrected for multiple testing using the Benjamini- Hochberg procedure (citation: Benjamini Y, Hochberg Y (1995). "Controlling the false discovery rate: a practical and powerful approach to multiple testing". Journal of the Royal Statistical Society, Series B. 57 (1): 289-300. doi:10.1111 / j.2517-6161. 1995.tb02031.x) d. Cell-type assignments were retained where the false discovery rate was shown to be below 1%.
[0299] This allowed the unambiguous assignment of the cell type of 61 IMR90 cells (shown as circles in Figure 10) and 45 Jurkat cells (shown as squares in Figure 10). Thus, our technique can successfully be used to produce data describing single cells of a particular cell type from an input sample comprising a mixture of human cell types, with high confidence.
Claims
CLAIMS1. A method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in cells, the method comprising the steps:(a) treating the cells with a fixative to cross-link the DNA regulatory interactions;(b) processing the cells to fragment the DNA and introduce first adapters into accessible regulatory DNA;(c) proximal ligation of the DNA fragments to join the proximal first adapters of the DNA fragments, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;(d) reversing the cross-linking;(e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and(f) sequencing and sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions.
2. The method of claim 1, wherein the proximal ligation step (c) comprises proximal bridge ligation, wherein one or more second adapters are introduced and join the proximal first adapters.
3. The method of claim 1 or 2, wherein the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (b) or step (c).
4. The method of any one of claims 1 to 3, wherein the method further comprises introducing one or more indexes to the DNA fragments.
5. The method of any one of claims 1 to 4, wherein the cells comprise a heterogenous cell population.
6. The method of any one of claims 1 to 5, wherein after step (a) the cells are distributed into a first plurality of compartments, each compartment comprising a single cell or a subset of cells.
7. The method of any one of claims 1 to 6, wherein processing the cells in step (b) further comprises introducing the first adapters comprising a first index into accessible regulatoryDNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index that is different from the first index sequence in the other compartments.
8. The method of claim 7, wherein after processing the cells in step (b) the first indexed DNA fragments are pooled and distributed into a plurality of second compartments.
9. The method of any one of claims 1 to 8, wherein the proximal ligation step (c) further comprises proximal bridge ligation of the DNA fragments to introduce and join one or more second adapter comprising an index to the proximal first adapters of the DNA fragments, to generate indexed DNA fragments.
10. The method of claim 9, wherein the index in the proximal bridge ligation step (c) comprises a second index introduced to the DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes.
11. A method of performing high-throughput chromosome conformation capture on DNA which identifies DNA regulatory interactions between accessible chromatin regions in each cell in a heterogenous cell population, the method comprising the steps:(a) treating the heterogenous cell population with a fixative to cross-link the DNA regulatory interactions;(a1) distributing the cells from the heterogenous cell population into a first plurality of compartments, each compartment comprising a subset of cells;(b) processing the cells to fragment the DNA and introduce first adapters comprising a first index into accessible regulatory DNA to generate first indexed DNA fragments, each compartment comprising a subset of DNA fragments with the same first index;(b1) pooling of first indexed DNA fragments and distributing into a plurality of second compartments;(c) proximal bridge ligation of the first indexed DNA fragments to introduce and join one or more second adapter comprising a second index to proximal first adapters of the first indexed DNA fragments to generate second indexed DNA fragments, each second compartment comprising a subset of DNA fragments with the first and second indexes, wherein the DNA fragments further comprise one or more nucleotides comprising a first half of a binding pair moiety;(d) reversing the cross-linking;(e) enriching for DNA fragments comprising the first half of the binding pair moiety using a second half of the binding pair moiety; and(f) sequencing and multiplex sequence analysis of the enriched DNA fragments obtained in (e) to identify the DNA regulatory interactions; wherein sequence reads bearing the same combination of indexes will be derived from a single cell.
12. The method of any one of claims 1 to 11 , wherein the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (b) with the first adapters.
13. The method of any one of claims 1 to 12, wherein the fixative is paraformaldehyde.
14. The method of any one of claims 1 to 13, wherein the processing step (b) is performed by tagmentation using a recombinase enzyme.
15. The method of claim 14, wherein the recombinase enzyme is a retroviral integrase, such as mutant transposase, such as hyperactive Tn5 transposase.
16. The method of claim 14 or 15, wherein gaps in the DNA fragments generated during tagmentation are filled in with a polymerase and ligated.
17. The method of claim 16, wherein the one or more nucleotides comprising the first half of the binding pair moiety are introduced to the DNA fragment in step (c) when filling the gaps in the DNA fragments.
18. The method of any one of claims 1 to 17, wherein the first half of the binding pair moiety comprises biotin and the second half of the binding pair moiety comprises streptavidin.
19. The method of any one of claims 1 to 18, wherein the DNA regulatory interactions comprise DNA-DNA interactions, preferably, promoter-enhancer interactions, enhancerenhancer interactions or promoter-promoter interactions.
20. The method of any one of claims 1 to 19, wherein sequencing in step (f) further comprises introducing a sequencing index.
21. The method of any one of claims 1 to 20, wherein the method uses combinatorial barcoding and wherein the distribution of DNA fragments into compartments results in unique combinations of indexes.
22. The method of any one of claims 1 to 21 , wherein each index comprises a barcode or Unique Molecular Identifier (UM I) which differs for each index.
23. A method of identifying DNA regulatory interactions in a cell population that are indicative of a particular disease state, comprising:(a) performing the method of any one of claims 1 to 22 on a cell sample obtained from an individual with a particular disease;(b) quantifying a frequency of DNA interactions within the cell sample; and(c) comparing the frequency of DNA interactions from the individual with said disease state with the frequency of interaction in a normal control cell sample from a healthy subject, such that a difference in the frequency of DNA interactions is indicative of a particular disease.
24. A method of identifying DNA regulatory interactions in a heterogenous cell population that are indicative of a particular disease state, comprising:(a) performing the method of any one of claims 1 to 22 on a heterogenous cell sample obtained from an individual with a particular disease;(b) quantifying a frequency of DNA interactions within an identified single cell or a group of single cells from the same cell type or cell state; and(c) comparing the frequency of DNA interactions from the individual with said disease state with the frequency of interaction in a normal control heterogenous cell sample from a healthy subject, such that a difference in the frequency of DNA interactions is indicative of a particular disease.
25. The method of claims 23 or 24, wherein the disease state is selected from: cancer, an autoimmune disease, a developmental disease, ageing or a genetic disorder.
26. A method of identifying DNA regulatory interactions in a cell population that are indicative of the effect of a medicament of interest, comprising:(a) performing the method of any one of claims 1 to 22 on a cell sample obtained from an individual treated with said medicament;(b) quantifying a frequency of DNA interactions within the cell sample; and(c) comparing the frequency of DNA interactions from the individual treated with said medicament with the frequency of DNA interactions in a normal control cell sample from a non-treated subject, such that a difference in the frequency of DNA interactions is indicative of the effect of treatment with said medicament.
27. A method of identifying DNA regulatory interactions in a heterogenous cell population that are indicative of the effect of a medicament of interest, comprising:(a) performing the method of any one of claims 1 to 22 on a heterogenous cell sample obtained from an individual treated with said medicament;(b) quantifying a frequency of DNA interactions within an identified single cell or a group of single cells from the same cell type or cell state; and(c) comparing the frequency of DNA interactions from the individual treated with said medicament with the frequency of DNA interactions in a normal control heterogenous cell sample from a non-treated subject, such that a difference in the frequency of DNA interactions is indicative of the effect of treatment with said medicament.
28. A kit for identifying DNA regulatory interactions between accessible chromatin regions in a cell population, comprising buffers and reagents capable of performing the method of any one of claims 1 to 27.
Citation Information
Patent Citations
Methods and compositions for scalable pooled RNA screens with single cell chromatin accessibility profiling
US20220267759A1
Method of identifing interactions between genomic loci
WO2010036323A1
Targeted chromosome conformation capture
WO2014168575A1
Chromosome conformation capture method including selection and enrichment steps
WO2015033134A1
Novel method
WO2021064430A1