Novel method

The method addresses the challenge of profiling multiple histone modifications and the transcriptome in a single cell by using pre-assembled antibody-transposase-adapter complexes for in situ single-step fragmentation, enhancing sensitivity and sample recovery in small-scale samples.

WO2026047343A1PCT designated stage Publication Date: 2026-03-05BABRAHAM INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing methods fail to simultaneously profile more than two different histone modifications and the transcriptome in a single cell, lacking in sample integrity, signal-to-noise ratio, and cell recovery, especially in small-scale samples.

Method used

A method involving pre-assembled antibody-transposase-adapter complexes for chromatin marks and genome binding proteins, with in situ single-step fragmentation and adapter insertion, followed by sequencing, and optional mRNA capture and reverse transcription.

Benefits of technology

Enables simultaneous profiling of multiple chromatin marks and the transcriptome in a single cell with high sensitivity and sample recovery, suitable for small samples, preserving sample integrity and reducing off-target signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051894_05032026_PF_FP_ABST
    Figure GB2025051894_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is directed to a method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising in situ single step fragmentation and adapter sequence insertion using individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, and mRNA transcript capture and in situ reverse transcription. Also provided are methods of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, as well as kits for performing the methods herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] BAB-C-P3666PCT

[0002] NOVEL METHOD

[0003] FIELD OF THE INVENTION

[0004] The present invention is directed to a method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising in situ single step fragmentation and adapter sequence insertion using individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, and mRNA transcript capture and in situ reverse transcription. Also provided are methods of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, as well as kits for performing the methods herein.

[0005] BACKGROUND OF THE INVENTION

[0006] Advances in single cell chromatin profiling technologies have provided new opportunities to investigate epigenome dynamics in individual cells, thereby accelerating the understanding of epigenetic and gene control processes in complex biological systems. Chromatin state maps can help to predict functional regions within a genome. To chart these maps, combinations of multiple, specific histone modifications can collectively infer regulatory sequences in different forms of repressed, poised and active configurations, which is information that can be used to better predict gene regulatory inputs particularly at enhancers and promoters. Such differences in regulatory activity are often indistinguishable when using single modality chromatin accessibility profiles alone. To more accurately decipher a broad range of epigenetic inputs, it is therefore necessary to capture information on multiple chromatin modalities within individual cells.

[0007] Towards this goal, several methods have been developed that can simultaneously profile a small number of histone modifications from the same single cell including scMulti-CUT&Tag, MuLTI-Tag, uCoTarget, nano-CT and NTT-seq. The methods use different strategies to achieve antibody-specific index tagmentation of genomic DNA loci using Tn5 transposase, whereby each antibody recognises a distinct histone modification. Following library sequencing, antibody-specific indexes are used to assign sequence reads to different antibodies and deconvolve multiple chromatin modality profiles in the same single cell. These incisive approaches provide a first step towards mapping select histone modifications in single cells and a means to directly investigate the complex interplay between different modalities. BAB-C-P3666PCT

[0008] Despite these advances however, methods that can profile more than two different histone modifications together with transcriptome in the same individual cell are lacking. There is therefore a need to overcome this limitation, enabling a more accurate definition of chromatin state maps where combinations of modifications are jointly required. Further, a need exists to develop a method which includes multiplexed antibody tagmentation in parallel rather than sequentially to preserve sample integrity, strong signal-to-noise with minimal off-target signals, and a high cell recovery rate which would open up the opportunity to profile small-scale or limited samples.

[0009] SUMMARY OF THE INVENTION

[0010] According to a first aspect of the invention, there is provided a method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising the steps of:

[0011] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0012] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0013] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0014] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0015] In certain embodiments, the method further comprises the steps of:

[0016] (iii b) capturing mRNA transcripts in the sample;

[0017] (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and

[0018] (iv b) sequencing the transcriptome library together with the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0019] In some embodiments, the sample is a single cell. BAB-C-P3666PCT

[0020] In particular embodiments, two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i), each complex differentially comprising one or more antibodies recognising a single chromatin mark, genome binding protein or genomic modification of interest and a single indexed transposase-adapter sequence, and the two or more antichromatin mark, anti-genome binding protein and / or anti genomic modification antibody- transposase-adapter complexes are used together simultaneously in step (iii). In further embodiments, three or more anti-chromatin mark, anti-genome binding protein and / or anti- genomic modification antibody-transposase-adapter complexes are individually preassembled in step (i) and used together simultaneously in step (iii), in particular four or more, more particularly five or more.

[0021] In a further aspect of the invention, there is provided a method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0022] (i) performing the method described herein on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0023] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with actively transcribed regions of the genome, such as actively expressed genes, wherein said actively transcribed regions are identified by sequencing the transcriptome library and wherein said actively transcribed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0024] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with actively transcribed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0025] In another aspect, there is provided a method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0026] (i) performing the method described herein on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ BAB-C-P3666PCT suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0027] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with transcriptionally repressed regions of the genome, wherein said transcriptionally repressed regions are identified by sequencing the transcriptome library and wherein said transcriptionally repressed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0028] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with transcriptionally repressed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0029] In a yet further aspect, there is provided a kit for identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said kit comprising individual pre-assembled and blocked antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, in addition to buffers and reagents capable of performing the method described herein. In some embodiments, the antibody- transposase-adapter complexes comprised in the kit are as defined herein.

[0030] BRIEF DESCRIPTION OF THE FIGURES

[0031] Figure 1 : A) Schematic overview of the present method. Ab, antibody; pA, proteinA; Tn5, Tn5 transposase; MEB, Mosaic End adapter sequence B; RT, reverse transcription. B) Genome browser tracks show H3K27me3 and H3K27ac histone modification signals in human pluripotent stem cells under various conditions tested. Tracks compare ENCODE ChlP-seq, single modality CUT&Tag and our new multi-modality method (H3K27me3 and H3K27ac, with and without joint transcriptome profiling). AS, adapter switching. Preincubation time refers to the duration when the antibody-proteinA-Tn5-adapter complexes were combined. IgG block refers to the amount of IgG antibody used to block any unreacted protein A-Tn5. RT, reverse transcription step, where “pre” means transcriptome profiling before the Tn5 tagmentation step, and “post” means after Tn5 tagmentation. C) Line plots and heatmaps show the same data as in B for all H3K27me3 peaks and all H3K27ac peaks defined by ENCODE ChlP-seq datasets. D) Scatter plots of histone modification signals (quantified as reads per million, RPM). X-axes show single-target with adapter switching; y- axes show different conditions that match the samples in B. Coefficients of determination BAB-C-P3666PCT

[0032] (adjusted R2) are shown. E) Bar charts show the fraction of reads in peaks (FRiP) as a measurement of on-target specificity. Peaks were defined using ENCODE ChlP-seq, singletarget CUT&Tag or ATAC-seq data. STB, single-target method; MTB, multi-target method.

[0033] F) Scatter plots of histone modification signals (quantified as reads per million, RPM) demonstrate that IgG blocking helps to reduce cross-contamination signals. X-axes show H3K27ac and y-axes show H3K27me3, which are mutually-exclusive signals in the genome. Coefficients of determination (adjusted R2) are shown.

[0034] Figure 2: A) Line plots show enrichment of histone modification signals and IgG at ENCODE ChlP-seq peaks, demonstrating accuracy of multi-target profiling in bulk and singlecell versions of the method. Plots on the right-hand side show the same data normalised with IgG signal, evidencing that this normalisation step further minimises off-target signals. B) Similar plots to A, except that the data shown are obtained from other histone modification profiling methods (CoTarget, NTT-seq and MulTI-Tag) and evidence the higher off-target and cross-contamination signals obtained with these other methods as compared to the present method as in A.

[0035] Figure 3: A) Schematic overview of the present method when applied to collecting data from single cells. Ab, antibody; pA, proteinA; Tn5, Tn5 transposase; MEB, Mosaic End adapter sequence B; RT, reverse transcription. B) Violin plot of unique reads per cell for the six targets profiled and of unique RNA counts plus the number of genes per cell for the transcriptome. Numbers above each dataset show mean values. Boxplots are median with interquartile range and minimum / maximum whiskers. C) Genome browser tracks show histone modification signals for ENCODE ChlP-seq, single-modality CUT&Tag, bulk sample multi-modality CUT&Tag (six targets; “Bulk-MT”) and computationally aggregated scMTR-seq (six targets; “Aggregated MT”). scMTR-seq profiles are also shown for individual cells (100 cells per histone modification and IgG; “Single cell MT”). D) Heatmaps show pairwise Pearson correlation between datasets from single-modality CUT&Tag, Bulk-MT and aggregated MT assays using genome-wide signals of 5 kb bins. The numbers represent the Pearson correlation coefficients. E) Violin plots show the fraction of reads in peaks (FRiP) for five histone modifications as a measurement of on-target specificity. Peaks were defined using scMTR-seq aggregated MT (upper), single-modality CUT&Tag (middle) or ENCODE ChlP- seq data (bottom). F) Heatmaps show the percentage overlap between peaks defined using ENCODE ChlP-seq or single-target CUT&Tag and scMTR-seq for five histone modifications.

[0036] G) Violin plots of Cramer’s V correlation show the relationship between each pair of histone modifications and with transcription. As expected, gene expression was positively associated with active histone marks, and H3K27me3 had low association with active histone marks and with gene expression. Boxplots are median with interquartile range and minimum / maximum whiskers. H) Plots show Pearson correlation coefficients of genome-wide histone modification BAB-C-P3666PCT signals of 5kb bin or 20kb bin between aggregated data of different cell numbers of down- sampled scMTR-seq with corresponding histone modification data from single-modality CUT&Tag (left) or total scMTR-seq aggregated data (right). Data from relatively few cells (-500 cells) can recapitulate the bulk CUT&Tag data. The red and blue dots indicate different bin sizes used for the correlation analysis. I) ChromHMM-defined chromatin states using computationally aggregated data from the single-cell multi-target method. Shown are different chromatin states (numbered 1 to 15). The numbers within the squares refer to emission probability. The numbers within the wider boxes refers to ratio of total genome. J) Alluvial plot shows the high overlap in regions defined in specific chromatin states as obtained from data generated by single-target CUT&Tag, bulk multi-target and scMTR-seq. MT, multi-target. K) Dimensional reduction analysis visualised on a Uniform Manifold Approximation and Projection (UMAP) for each histone modification, for transcriptome, and for joint projection with all modalities.

[0037] Figure 4: A) Genome browser tracks of histone modification and CTCF signals in human pluripotent stem cells. Tracks correspond to the single-target method, and to the multitarget method with increasing numbers of targets that were profiled at the same time (from 2 to 8). B) Line plots show the same data as in A, for peaks defined by ENCODE ChlP-seq data.

[0038] Figure 5: A) Genome browser tracks of histone modification signals in human pluripotent stem cells. Tracks correspond to single-target, multi-target (six targets) profiled simultaneously (T6-co), multi-targets (six targets) profiled sequentially in three steps (T6-seq3) and multi-targets (six targets) profiled sequentially in six steps (T6-seq6). The T6-co track is also shown normalised to IgG signal (T6-co*). B) Line plots show the same data as in A, over peaks defined by ENCODE ChlP-seq data. The number in the black square shows the order in the sequential tagmentation.

[0039] Figure 6: A) Dimensional reduction analysis visualised on a Uniform Manifold Approximation and Projection (UMAP) for each histone modification, for transcriptome, and for joint projection with all modalities, for human pluripotent stem cells undergoing differentiation to endoderm cells and sampled at four timepoints (days 0 to 3). B) Alluvial plot shows the different clustering results of single cells from four time points with different histone modification data, IgG signals, transcriptome (RNA) and all joint modality data. C) Genome browser tracks of transcripts, histone modifications and chromatin state annotations for human pluripotent stem cells undergoing differentiation to endoderm cells and sampled at four timepoints (days 0 to 3). D) Left, ChromHMM-defined chromatin states using computationally aggregated data from scMTR-seq. Shown are different chromatin states (numbered 1 to 15). The numbers within the squares refer to emission probability. The numbers within the wider boxes refers to the ratio of total genome. Right, alluvial plots show the changes in chromatin BAB-C-P3666PCT states of some regions as the cells transition over the differentiation timecourse. E) Scatter plots show the very low collision rates (the rate that multiple cells may be indexed with the same barcode by chance) using the single-cell multi-target method. This is visualised by comparing the number of unique fragments per cell for individual cells from different samples (days). F) Violin plot of unique reads per cell for the six targets profiled and of unique RNA counts plus the number of genes per cell for the transcriptome. Numbers above each dataset shows mean. Boxplots are median with interquartile range and minimum / maximum whiskers. G) Dimensional reduction analysis visualised on a Uniform Manifold Approximation and Projection (UMAP) for gene expression of cell type-specific markers. H) Top, boxplots show the length of regions (Iog10 bp) within each of the 15 ChromHMM-defined chromatin states. Numbers above each dataset show mean values. Boxplots are median with interquartile range and minimum / maximum whiskers. Bottom, same data as in the top panel showing the frequency of the regions with different lengths. I) Alluvial plot shows changes in the chromatin state of enhancer regions (left) or promoter regions (right) over the differentiation timecourse.

[0040] Figure 7: A) Schematic of mouse embryo at Day 4.5 of development, with the three main lineages highlighted. B) Genome browser tracks of transcripts, histone modifications and chromatin state annotations for day 4.5 mouse embryos, whereby the data have been split according to embryo lineage (TE, trophectoderm; PE, primitive endoderm; EPI, epiblast). C) Dimensional reduction analysis visualised on a Uniform Manifold Approximation and Projection (UMAP) for each histone modification, for transcriptome, and for joint projection with all modalities, for day 4.5 mouse embryos. D) Network graph of gene regulatory interactions in day 4.5 mouse embryos, as defined using the single-cell multi-target method. Each node represents an enhancer predicted to be bound by the indicated transcription factor and is labelled by the name of the predicted target gene of that enhancer. Lines are coloured according to the assigned embryo lineage based on transcriptional specificity of the transcription factors (green, trophectoderm; blue, primitive endoderm; orange, epiblast). E) Line plots showing lineage-specific differences in histone modifications signals at promoters (upper panels) and enhancers (lower panels). F) Dimensional reduction analysis visualised on a Uniform Manifold Approximation and Projection (UMAP) for three transcription factor-centred networks (Gata3, Gata4 and Nanog). Plots show i) transcription factor (TF) expression; ii) the activity of enhancers predicted to be bound by the indicated transcription factors (as inferred by H3K27ac signals); and iii) expression of the genes that are predicted to be targets of the enhancers.

[0041] Figure 8: A) Upper panel: Diagram shows JAK2 genomic region around V617 and location of the primer used for specific enrichment (amplification and selection; indicated by black arrow); Lower panel: example of sequenced results showing point mutation from G to T (indicated by arrow), causing Valine (V) to Phenylalanine (F) amino acid mutation. B) Bar BAB-C-P3666PCT plots show JAK2 mutation rate per cell (N=11 ,257) in bone marrow samples and peripheral blood samples from healthy donors (number of samples = 4) or MPN patients (number of samples = 5). Cells with valid genotyped reads <5 are not included in this analysis.

[0042] DETAILED DESCRIPTION OF THE INVENTION

[0043] The present invention is based on the development of a single cell multi-omics sequencing method termed scMTR-seq (single cell Multi-Targets and mRNA sequencing), a method of simultaneously identifying genomic regions associated with chromatin marks in cells and profiling the transcriptome. It is a multi-omics sequencing method allowing the profiling of over five (e.g. up to eight) histone modifications together simultaneously in the same step and sample, optionally together with the transcriptome (i.e. actively expressed regions / genes). The invention provides high sensitivity, high sample recovery and is highly scalable in terms of the amount of starting material required, particularly allowing its use in single cells and / or small samples, such as from patient samples / biopsy. Demonstrated herein is the successful use of the method of the invention to investigate the dynamics of human pluripotent stem cell to definitive endoderm differentiation and of mouse embryo lineages, revealing the gene regulatory networks changes that drive specification of these cell types.

[0044] In a first aspect of the invention, there is provided a method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising the steps of:

[0045] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0046] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0047] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0048] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0049] As used herein “genomic regions associated with chromatin marks” includes sequences and regions in the genome which are associated with cellular components capable of being recognised and bound by moieties for analysis and / or labelling, such as antibodies and fragments / derivatives thereof. In particular, the cellular components / chromatin marks are BAB-C-P3666PCT histones, including histone isoforms, modified histones and histone modifications. Also encompassed are genomic regions associated with any genome binding protein, transcription factors and the transcriptional machinery, as well as genomic regions / sequences harbouring direct modifications (e.g. DNA methylation). Thus in some embodiments, the genomic region is associated with genome binding proteins and / or genomic modifications. In particular embodiments, the genomic region is associated with a DNA-binding or a chromatin-binding protein, such as a transcription factor. In one embodiment, the transcription factor is CTCF. Further examples of DNA- and chromatin-binding proteins include, without limitation, chromatin remodelling enzymes. In further embodiments, the genomic region is associated with nucleic acid modifications, such as DNA modifications selected from methylation and / or hydroxymethylation, e.g. methylated adenine or cytosine. Thus, “genomic regions associated with genome binding proteins and / or genomic modifications” as used herein refers to any such DNA- or chromatin-binding protein (e.g. transcription factors) and / or DNA modification (e.g. methylation) which can be recognised and bound by antibodies and fragments / derivatives thereof. Any such DNA- or chromatin-binding protein and / or DNA modification can be used to profile the transcriptome according to the methods herein.

[0050] In preferred embodiments, the chromatin marks are histone modifications. In further particular embodiments, the histone modifications are any one or more of: phosphorylation, acetylation, ubiquitination, GIcNAcylation, citrullination, crotonylation, propionylation, sumoylation, isomerisation and / or methylation, including mono-methylation, di-methylation and trimethylation. Histone modifications are associated with the regulation of gene expression, such that they are involved in the ‘opening up’ of regions with actively transcribed genes or the ‘closing’ of regions not under active transcription or in which transcription is repressed. For example, euchromatin in which actively expressed genes and active transcription are commonly found is regulated by histone acetylation, phosphorylation and ADP ribosylation which add negative charges to the histone proteins, disrupting interactions with the DNA. By contrast, heterochromatin is associated with genomic regions not undergoing active expression or which are transcriptionally repressed (i.e. it is ‘closed’) and it is regulated by histone deacetylation and methylation, such as the di- and tri-methylation of H3K9. Thus in some embodiments, the chromatin mark, genome binding protein, genomic modification and / or histone modification is associated with euchromatin or heterochromatin. In one preferred embodiment, the chromatin mark, genome binding protein and / or genomic modification is associated with an actively transcribed genomic region. Thus in a certain preferred embodiment, the histone modification is associated with euchromatin. In a further preferred embodiment, the genome binding protein is involved in transcription, in particular active transcription, such as a transcription factor driving transcription of the genomic region BAB-C-P3666PCT and / or a component of the transcription machinery (e.g. an elongation factor, initiation factor or polymerase, such as RNA polymerase II). In another embodiment, the chromatin mark is associated with a transcriptionally repressed genomic region. Thus in one embodiment, the histone modification is associated with heterochromatin. In a further embodiment, the genomic modification is in the promoter or enhancer of a gene and for example represses transcription, such that the genomic modification is associated with transcriptional repression. Thus, according to these embodiments the transcriptome may be profiled (e.g. to determine whether a region / gene is actively expressed) using the identity and type of histone modification.

[0051] In yet further embodiments, the chromatin mark and / or histone modification is selected from any of: H3K4me1 , H3K4me2, H3K4me3, H3K36me3, H3K79me2, H3K9Ac, H3K27AC, H4K16AC, H3K27me1 , H3K27me2, H3K27me3, H2AK119ub, H3K122ac, H3K9me2, H3K9me3, H3R2me3, H3R8me3, H3R17me3, H3R26me3, H3R42me3, H3S10P, H3S28P, H4K5me3, H4K8me3, H4K12me3, H4K16me3, H4K20me3, H4R3me3 and gamma H2A.X. In particular embodiments, the chromatin mark is H3K27Ac, H3K27me3, H3K4me1 , H3K4me3 and / or H3K36me3. In other embodiments, the chromatin mark is a histone selected from a ‘core’ histone, such as any of: H2A, H2B, H3 and H4. In further embodiments, the chromatin mark is a histone isoform and / or variant. Various histone isoforms / variants are known in the art and will be known to the skilled person to be associated with / confer on chromatin various structural and functional features. Currently, the most comprehensive manually curated resource on histones and their variants is “HistoneDB 2.0” maintained by the NCBI (https: / / www.ncbi.nlm.nih.gOv / research / HistoneDB2.0 / ). In a further embodiment, the chromatin mark and / or histone modification is acetylation. In another embodiment, the chromatin mark and / or histone modification is methylation, including mono-methylation, dimethylation and tri-methylation.

[0052] Methods of identifying and profiling herein described chromatin marks, histone modifications, genome-associated proteins and DNA modifications, as well as the genomic regions associated with them, have been previously described in the art. These include scMulti- CUT&Tag, MuLTI-Tag, uCoTarget, nano-CT and NTT-seq. uCoTarget (Combined TAgmenting enRichment for multiple epiGEneTic proteins in the same cells) is a method of identifying genomic regions associated with histone modifications and proteins of the transcription machinery (e.g. transcription factors), and has been further combined with analysis of the transcriptome (dubbed uCoTargetX; Xiong et al. (2024) Set. Adv., 10(1): eadi3664, doi: https: / / oi.Org / 10.1 26 / sciadv.adi3664). While it is notable that Xiong et al. describe the analysis of up to five histone modifications “simultaneously” or up to two histone BAB-C-P3666PCT modifications with the transcriptome in cells, this appears to refer to the analysis of multiple parameters in the same sample and / or the same cell rather than to the labelling (corresponding to single step fragmentation and adapter sequence insertion step (iii) herein) of each histone modification or transcription factor which is clearly described therein as being performed sequentially. Sequential labelling without Fc blocking is described in Xiong et al. as providing the highest correlation scores with bulk in situ data in their hands. Thus, in contrast to the methods described herein when two or more anti-chromatin mark antibody- transposase-adapter complexes are used together simultaneously in single step fragmentation and adapter sequence insertion step (iii), uCoTargetX uses preassembled antihistone modification / transcription factor antibody-protein A-Tn5 transposase-adapter complexes recognising different histone modifications / transcription factors sequentially. The term “together simultaneously” herein may also be used interchangeably with “in the same step”, i.e. all two or more anti-chromatin mark antibody-transposase-adapter complexes are added in the single step fragmentation and adapter sequence insertion step (iii).

[0053] Related methods to uCoTarget for profiling chromatin mark-associated genomic regions without analysis of the transcriptome are: MAblD (Lochs et al. (2024) Nat. Methods, 21 :72-82, single step tagmentation, instead combining standard restriction digestion of the genome with ligation of barcoded sequences at the cut site, wherein the barcoded sequences are directed to regions associated with particular chromatin marks by conjugation to anti-chromatin antibodies. MulTI-Tag (Meers et al. (2023)) combines the sequential labelling / tagmentation of genomic regions associated with up to three histone modifications or up to two histone modifications with RNA polymerase II (RNA pol II) to indicate actively transcribed regions of the genome. NTT-seq (Stuart et al. (2023)) uses barcode adapter-loaded Tn5 transposase fused to nanobodies as secondary labels recognising primary anti-histone modification antibodies to tagment the adapters to genomic regions associated with said histone modifications. Up to three primary anti-histone modification antibodies or two anti-histone modification antibodies with an anti-RNA pol II antibody have been used in a single labelling step, followed by secondary addition of Tn5-nanobodies and transposase activation. Multi- CUT&Tag (Gopalan et al. (2021)) simultaneously tagments multiple chromatin mark- or RNA pol Il-associated genomic regions using complexes of anti-histone modification and anti-RNA pol II antibodies with adapter-loaded Tn5 transposase. Up to two anti-histone modification BAB-C-P3666PCT antibody-transposase-adapter complexes are used with an anti-RNA pol II antibody- transposase-adapter complex to identify genomic regions associated with said histone modifications and which are being actively transcribed. Thus, none of the previously described methods simultaneously identify genomic regions associated with chromatin marks, in particular multiple chromatin mark-associated genomic regions (such as five or more) together in a single step, and optionally the transcriptome by sequencing. As demonstrated herein, blocking step (ii) provides the simultaneous identification of multiple chromatin mark- associated genomic regions together in a single step, which has not previously been disclosed or appreciated in the art. For example, using Fc to block protein A after each sequential round of tagmentation in the uCoTarget method did not affect the resulting fraction of reads in peaks (FRiP) or correlation scores with bulk in situ data (Xiong et al. (2024)).

[0054] As demonstrated herein, the present methods are able to identify genomic regions and the transcriptome in samples with little / small amounts of starting material. In other words, the methods allow for a reduction in starting material compared to previously described methods. As such, a single cell may be used as the starting material / sample. Thus in some embodiments, the genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications are identified simultaneously in a single sample, such as in a single population of cells or in a single cell. In further embodiments, the transcriptome of the cell is identified simultaneously in the single sample, such as in the single population of cells or in the single cell. In yet further embodiments, the genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome of the cell are identified simultaneously in the single sample, such as in the single population of cells or in the single cell. The ability to use a reduced amount of starting material is achieved by increasing the recovery of genomic nucleic acid, i.e. increasing the amount of adapter-tagged library and transcriptome library generated from a given amount of starting material. This is achieved by the method e.g. utilising a reduced number of wash steps compared to as required with previously described methods which comprise sequential labelling of chromatin marks, histone modifications, genome binding proteins, genomic modifications, etc. In particular, sequential labelling of multiple genomic regions involving multiple sequential washes between each labelling step as described in previous methods leads to significant loss of material which is avoided herein. As will be appreciated, this loss due to the number of washes will increase with the sequential labelling of larger numbers of genomic regions. Further, the present method can utilise cells and / or intact nuclei immobilised onto beads (e.g. conA beads). BAB-C-P3666PCT

[0055] In some embodiments, the sample is a small sample, such as one or more or a small population of cells. Thus in one embodiment, the sample comprises one or more cells. In some embodiments, the sample is a bulk sample of cells or a population of cells, such as a bulk population of cells. The number of cells may be small, e.g. the sample is a small cell population. In particular embodiments, the sample is a single cell. In another particular embodiment, the sample comprises a single cell. Thus in further embodiments, the genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome of the cell are identified simultaneously in a single population of cells or in a single cell. In further embodiments, the cell is obtained from a subject, such as from a tissue or organ of a subject, e.g. the blood. In one embodiment, the cell is obtained from a biopsy, such as a tissue or organ biopsy from a subject. In a further embodiment, the cell is obtained from a diseased tissue or organ (e.g. from a biopsy of a diseased tissue or organ) or from a tissue / organ suspected as being diseased. In a yet further embodiment, the cell is obtained from a biopsy of a cancer. Examples of cancers, including tumours, solid and blood cancers, are well known. In another embodiment, the cell is obtained from a developing or developmental tissue or organ (e.g. from an embryonic or reproductive tissue or organ). As demonstrated herein, the method described herein finds utility in analysing developing tissues such as from the embryo, where the number of cells and numbers of each cell type may be particularly small. As will be appreciated from the disclosures herein, the method of the invention is performed in vitro such that the sample / cell may be obtained and subjected to the method ex vivo and / or the sample is an in vitro sample. Thus in certain embodiments, the methods described herein are performed in vitro. In further embodiments, the methods are performed ex vivo. In another embodiment, the sample is obtained from a cell culture, such as an in vitro cell culture. Examples of in vitro cell culture samples on which the method of the invention may be performed include, without limitation, differentiation and / or development cultures. Differentiation and / or development cultures comprise culturing a cell under conditions whereby the phenotype, properties and / or functions of the cell are changed in response to changes in the culture environment (e.g. in directed differentiation) or in response to factors introduced into or expressed in the cell (e.g. transcription factors, pluripotency reprogramming factors and / or differentiation factors).

[0056] In further embodiments, the sample comprises / is permeabilised cells and / or intact cell nuclei. Thus in some embodiments, step (iii) is performed on permeabilised cells and / or intact cell nuclei. Step (iii b) may also be performed on permeabilised and / or intact cell nuclei. As such, certain steps of the present method may be described as being performed in situ. “In situ" as used herein refers to the performing of said method step within the sample mixture, without any isolation or separation of the individual components (e.g. cells) or any performing of the BAB-C-P3666PCT step separately from the sample. In other words, the in situ step is performed within the cell / nuclei, such as within the permeabilised cells and / or intact nuclei. Thus, according to steps (iii) and (iii b) herein, single step fragmentation and adapter ligation (step (iii)) and optional reverse transcription of captured mRNA transcripts (step (iii b)) are performed in situ in the cells / intact nuclei without any purification of the starting material from the sample, i.e. they are performed in the sample per se. The cells / nuclei subjected to the method may be cross-linked to maintain the spatial arrangement of sub-cellular and sub-nuclear components, for example to maintain an interaction / association between a genomic region and histones (in particular modified histones) and / or transcription factors during the steps of the method. Thus in one embodiment, the cells and / or intact nuclei are cross-linked. In a further embodiment, the method additionally comprises reversing cross-linking, in particular reversing cross-linking following step (iii) and optionally step (iii c). Reversing cross-linking additionally comprises lysing the cells / nuclei. In another embodiment, the permeabilised cells and / or intact nuclei are not cross-linked.

[0057] References herein to “cross-linking” refer to any stable chemical association between two components, such that they may be processed as a unit. Such stability may be based on covalent and / or non-covalent binding (e.g. ionic). For example, nucleic acids and proteins may be cross-linked by chemical agents (i.e. for example, a fixative), heat, pressure, pH change or radiation, such that they maintain their spatial relationships during laboratory procedures (e.g. extracting, washing, etc.). Cross-linking as used herein is equivalent to the term “fixing” and the like, which applies to any method or process that immobilises any and all cellular processes. A cross-linked / fixed cell therefore accurately maintains the spatial relationships between components with the genome at the time of fixation. Many chemicals suitable for cross-linking / fixation are known in the art, including without limitation formaldehyde, formalin and glutaraldehyde.

[0058] A ntibody-Transposase-A dapter Complexes

[0059] The method comprises the step of: (i) pre-assembling individual anti-chromatin mark, antigenome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes. “Pre-assembling” as used herein refers to the assembling of the antibody- transposase-adapter complexes prior to their use in the method (i.e. prior to their use in step (iii)). Thus, pre-assembly is performed prior to any contacting of the sample. Such preassembly is performed for individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, such that each complex comprises antibodies specific for a single chromatin mark, genome binding protein or genomic modification of interest. Thus, each complex recognises and binds to a single chromatin mark, BAB-C-P3666PCT genome binding protein and / or genomic modification of interest. Each complex may comprise one or more antibody (including antigen binding fragments and the like) provided they are specific for the same single chromatin mark, genome binding protein and / or genomic modification of interest, for example wherein the antibodies are polyclonal. Thus in embodiments, the anti-chromatin mark antibody-transposase-adapter complexes comprise one or more antibodies, antigen binding fragments and the like as described herein, such as two or more antibodies wherein the antibodies are polyclonal. As such, as used in certain embodiments herein “a” may be interchangeable with “a single”.

[0060] The terms “antibody”, “antibodies”, “antigen binding fragments” and the like herein are used in their normal context in the art, namely to describe a moiety which recognises and binds to a target of interest, i.e. a particular chromatin mark, genome binding protein or genomic modification as described herein. Examples of such binding moieties will be readily recognised by and are known to the skilled person, and thus will be appreciated to include full- length antibodies (including monoclonal and polyclonal antibodies), antigen binding fragments (including antigen binding fragments of monoclonal and polyclonal antibodies), minibodies, Fab fragments, F(ab’)2 fragments, diabodies, scFvs, scFv-Fcs, VHHs and VNARs. Other binding aptamers may also be known in the art.

[0061] Wherein multiple (e.g. two or more) antibody-transposase-adapter complexes are used in the present method, each complex comprises antibodies for different chromatin marks, genome binding proteins and / or genomic modifications of interest, such that each complex recognises and binds to each of the different chromatin marks, genome binding proteins and / or genomic modifications. Pre-assembly is also performed such that each complex comprises a single indexed transposase-adapter sequence. “Single indexed transposase-adapter sequence” as used herein refers to the adapter sequence perse being a “single” or unique sequence to the complex, with the transposase being a standard enzyme known in the art and / or described herein. Therefore in certain embodiments, the indexed adapter sequence comprises a unique molecular identifier, such as an indexing and / or barcoding sequence. Other suitable molecular identifiers may be known in the art. The presence of antibodies specific for a single chromatin mark, genome binding protein or genomic modification together with a single indexed transposase-adapter sequence provides the recognition, binding (by the antibody) and subsequent identification (using the unique identifier of the adapter sequence, inserted by the transposase enzyme) of single chromatin marks, genome binding proteins and / or genomic modifications and the genomic regions associated therewith. Thus according to these embodiments, each anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complex comprises one or more antibodies BAB-C-P3666PCT recognising different chromatin marks, genome binding proteins or genomic modifications of interest and different indexed transposase-adapter sequences, such that each indexed transposase-adapter sequence is unique to each chromatin mark, genome binding protein and genomic modification of interest.

[0062] In some embodiments, two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually preassembled in step (i). As described hereinbefore, according to these embodiments each complex differentially comprises one or more antibodies recognising a single chromatin mark, genome binding protein or genomic modification of interest and a single indexed transposase- adapter sequence. Furthermore, the two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are used together simultaneously in step (iii). “Used together simultaneously” as used herein refers to the multiplex / parallel use of the two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, such that they are used in step (iii) for tagmentation in the sample at the same time / in a single reaction, i.e. they are added together in a single step. This is distinct from sequential addition / use in a single sample which has previously been erroneously described as “simultaneous” (e.g. in Xiong et al. (2024) as discussed hereinbefore). Thus, the present disclosure is the first description of true simultaneous tagmentation of genomic regions associated with two or more chromatin marks, genome binding proteins and / or genomic modifications, with analysis of the transcriptome within a cell. In further embodiments, three or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii). In a preferred embodiment, four or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii). In a particularly preferred embodiment, five or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii). In a further preferred embodiment, six or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually preassembled in step (i) and used together simultaneously in step (iii). In a yet further preferred embodiment, seven or more anti-chromatin mark, anti-genome binding protein and / or anti- genomic modification antibody-transposase-adapter complexes are individually preassembled in step (i) and used together simultaneously in step (iii). In a still further preferred embodiment, eight anti-chromatin mark, anti-genome binding protein and / or anti-genomic BAB-C-P3666PCT modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii). As described hereinbefore, none of the previously described methods (scMulti-CUT&Tag, MuLTI-Tag, uCoTarget, uCoTargetX, nano-CT and NTT-seq) comprise the simultaneous use of four or more antibody-transposase- adapter complexes, and thus the present disclosure is the first description of the simultaneous tagmentation of genomic regions associated with four or more chromatin marks, genome binding proteins and / or genomic modifications, i.e. the simultaneous tagmentation of four or more genomic regions associated with said features.

[0063] As will be readily appreciated, following the simultaneous use of two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase- adapter complexes together in step (iii), one or more further anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes may be used. This further use may be such that the one or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are used sequentially to the two or more complexes used together simultaneously in step (iii). Therefore, step (iii) may be repeated sequentially any number of times in the method herein with one or more further anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes following the simultaneous use together of two or more complexes in step (iii). In one embodiment, following step (iii), one or more further anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are used sequentially or together simultaneously for further in situ single step fragmentation and adapter sequence insertion in the sample. Such further in situ single step fragmentation and adapter sequence insertion may be combined with the use of only one anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complex in step (iii) in order to profile multiple genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in the sample.

[0064] Thus in an exemplary embodiment, the method comprises:

[0065] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0066] (ii) blocking the assembled antibody-transposase-adapter complexes; BAB-C-P3666PCT

[0067] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using two or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing further in situ single step fragmentation and adapter sequence insertion in the sample using one or more further assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0068] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0069] In another embodiment, the method comprises:

[0070] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0071] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0072] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using one or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing further in situ single step fragmentation and adapter sequence insertion in the sample using one or more further assembled antibody-transposase-adapter complexes, followed by performing yet further in situ single step fragmentation and adapter sequence insertion in the sample using one or more further assembled antibody-transposase-adapter complexes, followed by performing still further in situ single step fragmentation and adapter sequence insertion in the sample using one or more further assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0073] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0074] In a yet other embodiment, the method comprises:

[0075] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0076] (ii) blocking the assembled antibody-transposase-adapter complexes; BAB-C-P3666PCT

[0077] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using one or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing one or more sequential further in situ single step fragmentation and adapter sequence insertions in the sample using one or more further assembled antibody-transposase-adapter complexes to generate an adapter- tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0078] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0079] In a further embodiment, the method comprises:

[0080] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0081] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0082] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using one or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing three or more sequential further in situ single step fragmentation and adapter sequence insertions in the sample using one or more further assembled antibody-transposase-adapter complexes to generate an adapter- tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0083] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0084] In a yet further embodiment, the method comprises:

[0085] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0086] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0087] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using two or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing further in situ single step fragmentation and adapter sequence insertion in the sample using two or more further assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of BAB-C-P3666PCT sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0088] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0089] In a further embodiment, the method comprises:

[0090] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0091] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0092] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using one or more assembled antibody-transposase-adapter complexes together simultaneously as described herein, followed by performing four or more, in particular five or more, sequential further in situ single step fragmentation and adapter sequence insertions in the sample using one or more further assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0093] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0094] Blocking

[0095] The method comprises the step of: (ii) blocking the assembled antibody-transposase-adapter complexes. “Blocking” as used herein refers to a process which prevents or reduces potential off-target binding of the assembled antibody-transposase-adapter complexes, e.g. due to cross-reactivity between antibodies of the complexes (e.g. between different complexes) and / or between components of the complexes used to conjugate the antibodies with the indexed transposase-adapter sequence (e.g. between antibodies and the protein A, protein G, protein A / G or protein L or between the protein A, protein G, protein A / G or protein L of different complexes). As will be appreciated from the normal use of “blocking” in the art, only interactions which are undesirable are prevented / reduced. This is achieved by performing blocking step (ii) after assembling step (i), such that the interactions between anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies and the indexed transposase-adapter sequences within each individual antibody-transposase-adapter complex are not affected since these are pre-assembled before blocking, but any interactions between antibodies of one complex with another complex are prevented / reduced since these would otherwise occur after said blocking. In other words, intra-interactions within complexes BAB-C-P3666PCT are not blocked but inter-interactions between different complexes are blocked. Thus in certain embodiments, blocking step (ii) reduces off-target binding of the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes. Such reduction in possible off-target binding is particularly useful when two or more antibody-transposase-adapter complexes are used together simultaneously as described herein. Thus in a further embodiment, blocking step (ii) reduces off-target binding of the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes wherein two or more antibody-transposase-adapter complexes are used together simultaneously in step (iii). Blocking may also be used during steps (iii) and optionally step (iii b) to prevent or reduce off-target binding of antibody- transposase-adapter complexes to chromatin marks, genome binding proteins and / or genomic modifications not of interest, and / or to prevent / reduce the capturing of nucleic acid transcripts other than the mRNA transcripts of interest. Thus, blocking may comprise addition of free antibody (e.g. IgG antibody), a free half of a binding pair and / or sequences which are not complementary to the mRNA transcripts of interest but which bind other non-relevant transcripts. The particular identity of blockers and method of blocking will be readily recognised by the skilled person in light of the disclosures herein, as well as the specific e.g. conjugating moiety used or transcripts of interest.

[0096] In some embodiments, the anti-chromatin mark, anti-genome binding protein and / or anti- genomic modification antibodies are conjugated to the indexed transposase-adapter sequences using protein A, protein G, protein A / G or protein L. Protein A, protein G, protein A / G and protein L are bacterial proteins that bind to the Fab and Fc region of immunoglobulins, i.e. antibodies. Protein A and protein G are expressed in streptococcal bacteria, protein L is from the bacterial species Peptostreptococcus magnus and protein A / G is a recombinant fusion protein of the binding domains of protein A and protein G. The suitability of each of these will be recognised by the skilled person in light of the antibodies or fragments thereof used, for example if an anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification Fab-only fragment is used in the antibody-transposase-adapter complex then protein L which binds through light chain interactions (specifically with K light chains) will preferentially be selected over protein A, protein G or protein A / G which bind Fc regions. By contrast, if an antibody comprising only A light chains is used, then protein L would not be selected. Thus in one embodiment, the indexed transposase-adapter sequence comprises protein A, protein G, protein A / G or protein L, with a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence. In a preferred embodiment, the indexed transposase-adapter sequence comprises protein A, a transposases enzyme and a transposase adapter sequence. BAB-C-P3666PCT

[0097] As described hereinbefore, wherein the protein A, protein G, protein A / G or protein L is used to conjugate the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody to the indexed transposase-adapter sequence, blocking comprises blocking free protein A, protein G, protein A / G or protein L in the complexes. Free antibodies which do not recognise and bind to the chromatin mark, genome binding protein and / or genomic modification may be used, such as free IgG antibodies. Other suitable blocking antibodies or fragments may be used. In other words, blocking may comprise adding free antibody (e.g. free IgG). Thus in one embodiment, blocking step (ii) comprises blocking free protein A, protein G, protein A / G or protein L in the assembled antibody-transposase-adapter complexes. In a further embodiment, blocking step (ii) comprises blocking using free IgG antibodies.

[0098] In other embodiments, the anti-chromatin mark, anti-genome binding protein and / or anti- genomic modification antibodies are conjugated to the indexed transposase-adapter sequences using streptavidin and biotin, or another suitable binding pair. According to these embodiments, the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies comprise one half of the binding pair (e.g. biotin) and the indexed transposase-adapter sequences comprise the other half of the binding pair (e.g. streptavidin). Thus in one embodiment, the anti-chromatin mark, anti-genome binding protein and / or anti- genomic modification antibodies are biotinylated. In a further embodiment, the indexed transposase-adapter sequences comprise streptavidin. In a yet further embodiment the antichromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies comprise one half of a binding pair and the indexed transposase-adapter sequences comprise the other half of said binding pair. In a particular embodiment, the indexed transposase- adapter sequence comprises streptavidin, a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence, and wherein the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies are biotinylated.

[0099] Wherein a binding pair, such as streptavidin and biotin, are used to conjugate the antichromatin mark, anti-genome binding protein and / or anti-genomic modification antibody to the indexed transposase-adapter sequence, blocking comprises blocking free halves of the binding pair. For example, free biotin may be added in blocking step (ii). Thus in one embodiment, blocking step (ii) comprises blocking free streptavidin in the assembled antibody- transposase-adapter complexes. In a further embodiment, blocking step (ii) comprises blocking using free biotin. BAB-C-P3666PCT

[0100] In still other embodiments, the indexed transposase-adapter sequence may be directly conjugated to the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies. As will be readily appreciated, when direct conjugation of the antichromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies to the indexed transposase-adapter sequences is used, blocking is not necessary. Thus in a further embodiment, wherein the indexed transposase-adapter sequence is directly conjugated to the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies, blocking step (ii) may be omitted. In a yet further embodiment, blocking step (ii) is omitted with direct conjugation is used.

[0101] Single Step Fragmentation and Adapter Sequence Insertion / Tagmentation

[0102] The method further comprises the step of: (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification. “Single step fragmentation and adapter insertion” as used herein refers to the fragmentation of the genomic region associated with the chromatin mark, genome binding protein and / or genomic modification of interest and insertion of indexed adapter sequences in a single step, i.e. fragmentation of the genome and ligation of the indexed adapter sequences occurs concurrently. Such methods utilise a recombinase enzyme which binds to the adapter sequences and inserts these onto the fragmented genomic regions. This process is also known as “tagmentation”. Therefore in one embodiment, single step fragmentation and adapter sequence insertion comprises tagmentation. In a further embodiment, the method comprises the step of: (iii) performing in situ tagmentation in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification. As will be appreciated from the descriptions herein, such tagmentation is directed to the genomic regions of interest (i.e. the genomic regions associated with chromatin marks, genome binding proteins and / or genomic modifications of interest) due to the complexation of the transposase-adapter sequences with the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies, thus leading to the adapter-tagging of the genomic regions associated with said chromatin mark, genome binding protein and / or genomic modification.

[0103] Transposase enzymes suitable for use in the present methods will be appreciated to include any enzyme capable of removing (or cutting) and inserting sequence into a nucleic acid fragment (i.e. a genomic region). Examples of such transposases include retroviral integrase BAB-C-P3666PCT and transposase enzymes such as MuA, Tn5, Tn7 and Tc1 / mariner-type transposases. Thus in one embodiment, the transposase is a retroviral integrase. In a further embodiment, the transposase is Tn5 transposase. In order for the recombinase, integrase or transposase enzyme to be active, the enzyme may be mutated to overcome the naturally occurring low level of activity of such enzymes. Thus in a yet further embodiment, the transposase is a mutant transposase, such as a hyperactive transposase. Such a hyperactive transposase may be a mutant Tn5 transposase. In one embodiment, the transposase is a mutant Tn5 transposase, such as hyperactive Tn5 transposase.

[0104] Tn5 transposase is a member of the RNase superfamily of recombinase proteins which includes retroviral integrases and catalyses the movement of a portion of nucleic acid, known as a transposon, to another part of or another genome by a so called “cut and paste” mechanism. Recombinases, such as transposase enzymes, and transposon elements can be found in certain bacteria and are involved in the acquisition of antibiotic resistance. Transposase enzymes are commonly inactive and mutations in either the active site or elsewhere in the protein can lead to the generation of a hyperactive enzyme. Methods of producing Tn5 transposase enzyme are known in the art (Picelli et al. (2014) Genome Research 24:2033-2040, doi: .genome, grg / cgi / doj / 0. 0 / g lZ788;J.14).

[0105] However, these methods may be further adapted by utilising the indexed adapter sequences described herein when purifying the Tn5 transposase enzyme.

[0106] Adapter sequences used when purifying the transposase enzyme (e.g. the Tn5 transposase) incorporate with the enzyme and are subsequently inserted by said transposase into a nucleic acid fragments / genomic regions. Such sequences may be diverse in their sequence and comprise additional elements which enable further processing of the genomic region into which they are inserted. For example, adapter sequence incorporated with a purified transposase enzyme comprise an adapter sequence for sequencing and optionally a barcode sequence. It will be appreciated, however, that all such adapters comprise a transposon sequence or element which allows for incorporation with the enzyme. Examples of transposon sequences or elements include the Tn5 transposase-compatible Mosaic End (ME) sequence and sequences which are sterically compatible with the binding pocket of a recombinase and / or transposase enzyme. Thus according to one embodiment, the transposase-adapter sequence comprises Mosaic End Double-Stranded (MEDS) oligonucleotides, which comprise a half of paired end adapter sequences. In a further embodiment, the transposase-adapter sequences comprise one half of paired end adapter sequences for sequencing. In yet further embodiments, the transposase-adapter sequences may comprise paired end adapter sequences for sequencing which additionally comprise barcode sequences. In another BAB-C-P3666PCT embodiment, the transposase-adapter sequences comprise any sequence that enables subsequent library preparation and sequencing. Such sequences will be appreciated to enable the amplification and isolation of nucleic acid segments as well as the binding of said nucleic acid segments for analysis of sequence by high-throughout or next generation sequencing. Examples of next generation sequencing platforms include: Roche 454 (i.e. Roche 454 GS FLX), Applied Biosystems’ SOLiD system (i.e. SOLiDv4), Illumina’s GAIIx, HiSeq 2000 and MiSeq sequencers, Life Technologies’ Ion Torrent semiconductor-based sequencing instruments, Pacific Biosciences’ PacBio RS and Oxford Nanopore’s MinlON.

[0107] In further embodiments, the indexed transposase-adapter sequences comprise only Mosaic End B (MEB) adapter sequences. According to these embodiments, the method additionally comprises adapter switching, whereby Mosaic End A (MEA) adapter sequences are added to the one half MEB adapter sequences, resulting in genomic regions of interest that are pairend tagged for sequencing. Thus in one embodiment, the method additionally comprises adding Mosaic End A (MEA) adapter sequences to the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library following step (iii), i.e. adapter switching. In a further embodiment, the indexed transposase-adapter sequence comprises only Mosaic End B (MEB) adapter sequences, and the method additionally comprises adding Mosaic End A (MEA) adapter sequences to the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library following step (iii), i.e. adapter switching. Comprising only MEB adapter sequences and adapter switching overcomes the random incorporation of forward and reverse adapters inherent to tagmentation protocols; where 50% of molecules contain one forward and one reverse and are thus viable for sequencing, and the remaining 50% may contain two forward or two reverse adapters and are not viable for subsequent processing (reviewed in: Adey (2021) Genome Res., 31 (10):1693-1705, doi:

[0108] Thus, tagging with only MEB adapter sequences and subsequently adding MEA adapter sequences results in the tagging of all of the library with both adapter sequences required for processing and sequencing. Furthermore, comprising only MEB adapter sequences and adapter switching also prevents / reduces the generation of chimeric reads often seen after sequencing when two or more antibody-transposase-adapter complexes are used for tagmentation and a forward adapter is added to one genomic region and a reverse adapter added to another. Such chimeric reads appear as single reads following sequencing of the library and thus are either analysed incorrectly or are discarded, leading to a loss of data. Thus, tagging with only MEB adapter sequences and subsequently adding MEA adapter sequences prevents or reduces the generation of chimeric reads in the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence BAB-C-P3666PCT library. As such, in certain embodiments all of the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library in step (iv) and optionally step (iv b) is sequenced and / or the generation of chimeric sequence reads in the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification- associated sequence library is prevented. This is achieved by tagging with only MEB adapter sequences and subsequently adding MEA adapter sequences in the present method.

[0109] In yet further embodiments, the indexed transposase-adapter sequences comprise a unique molecular identifier. Examples of molecular identifiers include barcode sequences which can read out in sequencing step (iv) and optional step (iv b). Thus in some embodiments, the indexed transposase-adapter sequences comprise a barcode sequence. Such unique identifiers / barcodes allow the identification of the tagged genomic regions associated with the chromatin mark, genome binding protein and / or genomic modification of interest. As will be appreciated, the identifier / barcode is unique to each individual antibody-transposase-adapter complex such that, e.g. when two or more antibody-transposase-adapter complexes are used, each genomic region associated with each chromatin mark, genome binding protein or genomic modification of interest is uniquely identified / barcoded and can be separately identified, i.e. the identifier / barcode is unique to each chromatin mark-, genome binding protein- or genomic modification-associated genomic region. This allows the deconvolution of each tagged genomic region when two or more regions associated with chromatin marks, genome binding proteins and / or genomic modifications are tagged simultaneously / together in a single step herein. Such deconvolution or separation of data is performed after sequencing step (iv) or optional step (iv b).

[0110] The presence of a unique identifier / barcode or a further unique identifier / barcode allows for further identification and / or separation of tagged genomic regions when pooling. For example, the method described herein may additionally comprise splitting the sample following step (iii) and optionally step (iii c) into distinct samples followed by re-pooling. Between splitting and re-pooling, unique molecular identifiers / barcodes are added to each distinct sample to uniquely identify / barcode prior to re-pooling. This is known as split-pool combinatorial barcoding (also known simply as split-pool barcoding and combinatorial indexing), and allows the profiling of single cells in a sample by relying on the random distribution of individual cells in each split for barcoding. Split-pool barcoding and combinatorial barcoding are described further in Rosenberg et al. (2018) Science, 360(6385): 176-182, doi: Humphreys (2021) Kidney360,

[0111] 2(7):1196-1204, doi: https: / / dQs.Qrq / 10.340S7%2FKjD.0001S22021). Thus in one embodiment, wherein the method additionally comprises splitting the sample following BAB-C-P3666PCT step (iii) and / or optionally step (iii c) into distinct samples and adding unique molecular identifiers to each distinct sample, followed by re-pooling the uniquely identified distinct samples. In further embodiments, the splitting and re-pooling is repeated one or more times. According to these embodiments, the sample is split, barcoded and re-pooled two or more times. In a particular embodiment, the splitting and re-pooling is repeated two times. According to this particular embodiment, the sample is split, barcoded and re-pooled three times. In a still further embodiment, the splitting and re-pooling is repeated three times. According to this embodiment, the sample is split, barcoded and re-pooled four times. In one embodiment, addition of the first combinatorial barcode is by reverse transcription-PCR (RT- PCR) using barcoded RT primers. In a further embodiment, addition of the second combinatorial barcode is by ligation of barcoded sequences. In a yet further embodiment, addition of the third combinatorial barcode is also by ligation. In an alternative embodiment, addition of the third combinatorial barcode is by PCR. In a still further embodiment, addition of the fourth barcode is by PCR. The fourth split, barcode and re-pool step may be used to add a further combinatorial barcode or can instead add an adapter for sequencing rather than a combinatorial barcode. Such adapters are known and their identity / sequence will be readily recognised to be subject to the method of sequencing and / or machine used in step (iv) and optionally in step (iv b). Suitable alternative methods to split-pooling are known in the art (see Preissl, Gaulton & Ren (2023) Nat. Reviews Genetics, 24:21-43, doi: https : / / doi .org / 10.1038 / s41576-022-00609-1 , the barcoding methods discussed therein, in particular in Fig. 3, being specifically incorporated by reference) and will be appreciated to find utility in the present invention. Examples include, without limitation, plate-based systems (e.g. as used in Clark et al. (2018) Nature Comms., 9:781 , doi: https: / / doi.prg / 10 1038 / s41467-0 8-03149-4 and Takara’s ICELL8 nano-dispensing system as used in Kaya-Okur eta / . (2019) Nature Comms., 10:1930, doi: https: / / doi.org / 10.1038 / s41467-019-Q9982-5) including microwell and nanowell plates / chips, tube-based methods, microfluidic and integrated fluidic circuit methods, and droplet-based methods (e.g. as used in Zhang et al. (2022) Nature Biotechnol., 40:1220-1230, doi: https: / / dos.orQ / 1Q.1Q38 / s41587-022-01250-0: Xie et al. (2023) Nat. Struct. Mol. Biol., 30:1428-1433, doi: htps: / / dGi.orq / 1Q.1038 / s41594-023-01050-1 ; Macosko et al. (2015) Cell, 161 :1202-1214, doi: https: / / doi org / 10.101 S / i ceii.2015.05.0Q2; and Lareau et al. (2019) Nat. Biotechnol., 37:916-924, doi:

[0112] Standard steps and methods for the preparation of the libraries generated by the method herein may be used prior to or concurrently with sequencing step (iv) and optionally step (iv b). For example, the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and / or the transcriptome library may be amplified prior to sequencing. Thus in one embodiment, the method additionally comprises amplifying BAB-C-P3666PCT the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification- associated sequence library and the transcriptome library prior to step (iv) and optionally step (iv b). In certain embodiments, amplification is by PCR. As described hereinbefore, this amplification may be used to add adapter sequences for sequencing, the identity of which will be subject to the method of sequencing. Amplification may also be used to select for certain sequences, in particular within the transcriptome library, for example using primers specific and / or complementary for transcripts of interest (e.g. those encoding a particular subset of genes or non-coding transcripts).

[0113] Thus, in some embodiments amplification comprises one or more primer sequences complementary to a region surrounding a target sequence of interest. The target sequence may be genomic in the starting sample / material, within the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and / or within the transcriptome library. Accordingly, the present method comprising amplification (e.g. by PCR) may be used to determine the sequence of a target genomic region of interest or a target subset of transcriptomic sequences in the sample, e.g. genotyping said genomic or transcriptomic sequences. Such sequence determination (genotyping) may also use selection as well as or instead of amplification, for example using complementary sequences that are bound to a selection marker such as biotin or other unique molecular identifier. In certain embodiments, sequence determination (genotyping) comprises one or more primer sequences which are complementary to a region surrounding the target sequence and which further comprise a selection marker, such as biotin. The use of complementary primer sequences further comprising a selection marker yields amplified genomic regions or transcriptomic sequences comprising the target sequence of interest and the selection marker (e.g. biotin), allowing subsequent selection of the amplified sequences (e.g. using streptavidin beads and / or pulldown). Complementary primer sequences for amplification and / or selection are complementary to and thus anneal to a region surrounding the target sequence of interest within the genomic or transcriptomic sequences, in particular upstream (5’) of the target sequence in transcriptomic sequences and / or to both the upstream and downstream regions for genomic sequences. Complementarity and annealing to a region upstream of the target sequence in transcriptomic sequences provides the ability to selectively amplify and / or select said target with barcode sequences located at the 3’ end of the transcriptomic sequences, while complementarity and annealing to regions both upstream and downstream of the target sequence in genomic sequences provides the ability to selectively amplify and / or select the target with barcodes that may be located either side of the genomic fragment. Said primer sequences may also comprise a unique tag or further barcode for identification following sequencing (e.g. during data analysis). Following amplification and / or selection of the target BAB-C-P3666PCT sequences of interest, sequence information is determined by sequencing step (iv) and optionally step (iv b). Sequence information may be determined for the same target region of interest in both genomic sequences within the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and transcriptomic sequences within the transcriptome library.

[0114] The target sequence may be mutated compared to a wild type sequence located in the same genomic region or transcript. The target sequence may be a gene, an mRNA transcript or other coding sequence, including without limitation a protein coding region or region for a noncoding sequence (e.g. an miRNA, ncRNA, IncRNA or rRNA and the like). Thus, in a further aspect of the present invention there is provided a method of determining the sequence of, i.e. genotyping, a target sequence in a sample, said method comprising performing the method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome as described herein, and further amplifying or selecting genomic regions or transcriptomic sequences comprising a target sequence of interest prior to step (iv) and optionally step (iv b). Said amplifying or selecting comprises using primer sequences which are complementary and thus anneal to a region surrounding the target sequence of interest. The sequence of the target is then determined during sequencing step (iv) and optionally step (iv b) described herein. In one embodiment, the sample is obtained from a subject. The subject’s genotype may be of interest, such as for diagnostic and / or prognostic purposes.

[0115] Therefore and as will be readily appreciated, determining the sequence of a target sequence of interest that may be mutated compared to wild type will find particular utility in identifying diseases which are associated with mutation, such as cancer, developmental conditions and diseases for which a subject may be genetically predisposed. Thus, in a yet further aspect there is provided a method of predicting a subject’s risk of having and / or developing a disease or condition associated with a mutation, said method comprising determining the sequence of, i.e. genotyping, a target sequence in a sample obtained from the subject as described herein. Said predicting may be used for diagnosing and / or prognosing the disease or condition. In some embodiments, the disease and / or condition is cancer, such as a cancer associated with a mutation causing the activation of an oncogene and / or the inactivation of a tumour suppressor. In further embodiments, the method of predicting a subject’s risk of having and / or developing a disease or condition further comprises administering to the subject a treatment for said disease or condition if a mutation associated with the disease / condition is present. Thus, the present method may be used as a companion diagnostic. In some BAB-C-P3666PCT embodiments a suitable treatment for the disease or condition may be selected based on the identification of a mutation using the method described herein. mRNA Transcript Capture & Reverse Transcription

[0116] In some embodiments, the method described herein comprises the step of (iii b): capturing mRNA transcripts in the sample. Such capture may be by any method known to the skilled person, including using probe or bait sequences which are complementary to the transcripts of interest. Thus in one embodiment, capturing mRNA transcripts in step (iii b) is performed using a probe or bait sequence. For example, wherein it is desired to capture mRNA transcripts, such as mRNA transcripts of actively transcribed genes, a poly-T sequence may be used which is complementary to and binds the poly-A tail of mRNA transcripts. Thus in a further embodiment, capturing mRNA transcripts in step (iii b) comprises using a poly-T sequence. In a yet further embodiment, capturing mRNA transcripts in step (iii b) comprises using sequences complementary to mRNA transcripts of interest. Such complementary sequences may include without limitation, sequences complementary to and thus which bind promoter sequences, enhancer sequences, other regulatory elements present in the transcript, non-coding sequences, exonic sequences, intronic sequences (e.g. to capture nonspliced transcripts), SNP-containing sequences and mutated sequences. The probes, baits and / or complementary sequences may be short, e.g. around 10 nucleotides in length, such as 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 500, 750, 1000 or more nucleotides in length.

[0117] In certain embodiments, capturing the mRNA transcripts comprises a poly-T sequence such that the transcriptome comprises mRNA transcripts of actively transcribed regions of the genome. Thus in one embodiment, the transcriptome comprises mRNA transcripts of actively transcribed regions of the genome of the cell. In a further embodiment, the transcriptome comprises mRNA transcripts of actively expressed genes. According to this embodiment, the actively expressed genes will be predicted to be associated with euchromatin and chromatin marks, genome binding proteins and / or genomic modifications associated therewith (e.g. with the transcription machinery and / or histone modifications which ‘open up’ chromatin). Thus in a particular embodiment, the transcriptome comprises mRNA transcripts of actively transcribed regions of the genome associated with euchromatin.

[0118] After capture, the mRNA transcripts may be reverse transcribed in order to generate nucleic acid for sequencing. Such nucleic acid for sequencing is cDNA. Thus in further embodiments, the method described herein comprises the step of: (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library. Reverse BAB-C-P3666PCT transcription is known in the art and may be performed by any method and using any techniques known to the skilled person. In some embodiments, reverse transcription is performed in situ as described hereinbefore. Such in situ reverse transcription thus generates a transcriptome library for sequencing. The transcriptome library may comprise unique identifiers / barcodes as described hereinbefore to uniquely identify the transcriptome of individual single cells or of individual samples if samples from e.g. different sources are pooled. Thus in one embodiment, the reverse transcription step incorporates unique identifiers / barcodes into the transcriptome library. In a further embodiment, the reverse transcription step may incorporate adapter sequences for sequencing. As described hereinbefore, the specific identity of such adapters will be subject to the method of sequencing and machine used. The addition of such identifiers / barcodes and adapters when preparing libraries for sequencing is well known in the art. Sequencing barcodes are in addition to those unique for the individual anti-chromatin mark, anti-genome binding protein and / or anti-genome modification antibody-transposase-adapter complexes.

[0119] Thus in some embodiments, the method herein further comprises the steps of:

[0120] (iii b) capturing mRNA transcripts in the sample;

[0121] (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and

[0122] (iv b) sequencing the transcriptome library together with the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0123] In certain embodiments, steps (iii b) and (iii c) of the method described herein are performed after the single step fragmentation and adapter of step (iii). In other words, in certain embodiments the method described herein comprises performing in situ single step fragmentation and adapter sequence insertion, followed by capture of mRNA transcripts in the sample and subsequent in situ reverse transcription of the captured mRNA transcripts. As demonstrated herein, while performing mRNA transcript capture and reverse transcription prior to tagmentation did not negatively affect the quality of transcriptome sequencing data, it led to high background signal in the sequencing data from chromatin mark / genome binding protein / genomic modification-associated libraries. Therefore as demonstrated herein, performing mRNA transcript capture and reverse transcription after tagmentation provides low background signal in the data generated from sequencing both transcriptome and chromatin mark / genome binding protein / genomic modification-associated libraries, allowing more accurate profiling / identification of genomic regions associated with said features.

[0124] Thus in a certain embodiment, the method herein comprises the steps of: BAB-C-P3666PCT

[0125] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0126] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0127] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification;

[0128] (iii b) capturing mRNA transcripts in the sample;

[0129] (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and

[0130] (iv b) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and the transcriptome library.

[0131] Sequencing

[0132] The method described herein comprises the step of: (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library, and / or optionally step (iv b) sequencing the transcriptome library together with the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification- associated sequence library or sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and the transcriptome library. As will be appreciated, sequencing may be by any suitable method known in the art, such as any of the high-throughput or next generation sequencing technologies mentioned hereinbefore. Further, sequencing is possible by the presence of the adapter sequences tagged onto the libraries in tagmentation step (iii) and / or added in reverse transcription step (iii c), i.e. the adapters described herein allow or enable subsequent library preparation and sequencing of the adapter-tagged libraries. In some embodiments, the adapters are “paired end adapters” following adapter switching as described hereinbefore. Thus, herein a pair of MEA and MEB adapters may be considered “paired end”. “Paired end adapters” are any primer pair set that allow automated high throughput sequencing to read from both ends. For example, such high throughput sequencing devices that are compatible with these adapters include, but are not limited to Solexa (Illumina), the 454 System, and / or the ABI SOLiD system. Sequencing may also comprise parallel, including massively-parallel, sequencing of multiple samples and / or multiple cells within the same sequencing ‘well’. For example, the simultaneous sequencing of multiple single cells. According to embodiments comprising parallel sequencing, adapters include barcode or unique identifier sequences as BAB-C-P3666PCT described hereinbefore. Such barcodes for sequencing are in addition to those unique for the individual anti-chromatin mark, anti-genome binding protein and / or anti-genome modification antibody-transposase-adapter complexes.

[0133] Uses, Methods and Diagnostic Methods

[0134] The present method provides the identification of genomic regions which may be associated with chromatin marks, genome binding proteins and / or genomic modifications known to be associated with a disease or developmental state, or changes of the chromatin marks, binding of genome binding proteins and / or genomic modifications which may be associated with disease or developmental states.

[0135] Thus in a further aspect of the invention, there is provided a method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0136] (i) performing the method described herein on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0137] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with actively transcribed regions of the genome, such as actively expressed genes, wherein said actively transcribed regions are identified by sequencing the transcriptome library and wherein said actively transcribed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0138] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with actively transcribed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0139] According to this aspect, identifying step (ii) comprises correlating the genomic location associated with chromatin marks, genome binding proteins and / or genomic modifications with the transcriptome of the sample (e.g. cells, such as a single cell), such that actively transcribed regions of the genome are identified in the transcriptome. These actively transcribed regions will be positively represented in the transcriptome by virtue of being actively transcribed into mRNA (which is captured and reverse transcribed in steps (iii b) and (iii c) of the method described hereinbefore), or by virtue of being associated with marks of ‘active’ transcription, BAB-C-P3666PCT such as positive transcription factors, RNA polymerase and its subunits and chromatin / DNA modifications. Such actively transcribed regions may include genes associated with disease, e.g. disease driver genes, mutations known to be associated with disease, e.g. driver mutations in cancer, or genes associated with development / differentiation, e.g. genes associated with lineage commitment / developmental checkpoints. Being actively transcribed in the disease sample, it will be appreciated that the regions are positively associated with disease, i.e. they contribute to the development / progression of disease and / or increased severity. Furthermore, such actively transcribed regions in differentiation / developing samples will be positively associated with differentiation / development, such that their expression correlates with a progression through a developmental checkpoint, commitment to a lineage, or shutting down of a previous developmental / differentiation stage. Actively transcribed regions may also include non-coding or regulatory regions such as those which regulate the expression of genes. Such expression may be positive or negative.

[0140] In alternative embodiments, the identified chromatin marked, genome binding protein bound and / or modified nucleic acid regions are associated with transcriptionally repressed regions of the genome. Such association may be in the case of genes associated with earlier developmental / differentiation stages or with non-diseased states (i.e. when a gene associated with a healthy state is transcriptionally repressed, disease occurs).

[0141] Thus in an alternative aspect, there is provided a method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0142] (i) performing the method described herein on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0143] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with transcriptionally repressed regions of the genome, wherein said transcriptionally repressed regions are identified by sequencing the transcriptome library and wherein said transcriptionally repressed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0144] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with transcriptionally BAB-C-P3666PCT repressed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0145] According to this aspect, identifying step (ii) comprises correlating the genomic location associated with chromatin marks, genome binding proteins and / or genomic modifications with the transcriptome of the sample (e.g. cells, such as a single cell), such that transcriptionally repressed regions of the genome are identified in the transcriptome as described hereinbefore. These regions will be negatively represented in the transcriptome by virtue of being transcriptionally repressed and thus not transcribed into mRNA or by virtue of being associated with marks of ‘repressed’ transcription, such as negative transcription factors, transcription repressors (e.g. CTCF) and chromatin / DNA modifications. Such transcriptionally repressed regions may include genes and / or mutations associated with non-diseased / healthy states, e.g. tumour-suppressor genes in cancer, or genes associated with earlier developmental / differentiation states. Being transcriptionally repressed in the disease sample, it will be appreciated that these regions are negatively associated with disease, i.e. they contribute to the maintenance of a non-diseased / healthy state or they contribute to maintaining mild / lesser disease severity. Furthermore, such transcriptionally repressed regions in differentiation / developing samples will be negatively associated with differentiation / development, such that their expression correlates with an earlier stage of differentiation / development or the maintenance of a developmental checkpoint. Transcriptionally repressed regions may also include non-coding or regulatory regions such as those which regulate the expression of genes. Such expression may be positive or negative.

[0146] The samples obtained may be from any individual, sample (e.g. a biopsy sample) or e.g. in vitro culture. Thus in certain embodiments, the method is performed in vitro and / or ex vivo. According to these embodiments, the method is not performed in vivo, i.e. is not performed on the human or animal body. Samples obtained from an individual include biopsies of tissues and / or organs, as well as blood samples. The individual may be suspected of having a disease or is known to have a disease, such as a disease of the tissue / organ from which the sample is obtained. As described hereinbefore, the sample may contain cells, such as one or more cells, in particular a single cell. Also as described hereinbefore, the present methods find particular utility when working with small sample sizes which often occurs with biopsy samples, as well as single cells which is useful when e.g. working with cancer samples in which the cells may be highly heterogenous and the ability to detect mutations, expression and genomic associations in single cells is beneficial. The disease may be known or unknown. Thus in some embodiments, the methods described herein may be useful in disease identification. In further embodiments, the methods described herein may be useful in diagnosis. In yet further BAB-C-P3666PCT embodiments, the methods described herein may be useful in differential diagnosis, such as patient stratification to determine the likelihood of an individual’s response to treatment. In one embodiment, the method is used as a differential diagnosis. In a further embodiment, the method is used to determine the likelihood of a patient’s response to treatment, such as the likelihood the patient will be responsive to treatment. In another embodiment, the method is used to determine the likelihood that a patient will be unresponsive to treatment. Thus in still further embodiments, the individual has a disease, is suspected of having a disease and / or is in need of treatment. In yet further embodiments, the tissue or organ from which the sample is obtained is diseased, suspected of being diseased and / or is in need of treatment. In an alternative embodiment, the tissue / organ from which the sample is obtained is not diseased or suspected of being diseased, such that it is used as a control comparator to the sample obtained from a diseased / potentially diseased tissue / organ.

[0147] In one embodiment, the disease is cancer. Many cancers will be known to the skilled person, and the present methods will be appreciated to find utility in their diagnosis / differential diagnosis when driver genes are ‘switched on’ and thus become more actively transcribed or mutated such that they are more actively transcribed by virtue of changes in the chromatin marks, genomic binding proteins and / or genomic modifications with which they are associated. For example, several cancers are known to be associated with changes in epigenetic modifications which lead to increased expression (e.g. in cancer driver genes) or reduced expression (e.g. in tumour-suppressor genes). Thus in one embodiment, the disease is a cancer associated with epigenetic modifications. In a specific example, it is known that the ‘normal’ CpG methylation profile of the genome is often inverted in cells that become tumorigenic (Esteller (2007) Nat. Rev. Genetics, 8(4):286-298, doi: htps; / / oi,org / 10.1038%2Fnrg.2005). In non-cancerous cells, CpG islands preceding gene promoters are generally unmethylated (associated with active transcription), while other individual CpG dinucleotides throughout the genome tend to be methylated. However, in cancer cells CpG islands preceding tumour suppressor gene promoters are often hypermethylated (leading to a repression of their transcription), while CpG methylation of oncogene promoter regions and parasitic repeat sequences is often decreased (leading to their transcription and expression).

[0148] In another embodiment, the disease is an autoimmune disease. In a further embodiment, the disease is a developmental disease. In a still further embodiment, the disease is a genetic disorder. In other embodiments, the disease is any disease or disorder associated with changes in epigenetic modifications, preferably epigenetic modifications leading to changes in gene expression associated with disease. Examples of such diseases and conditions include, without limitation: blood pathologies; lymphoma and leukaemia (i.e. liquid cancers); BAB-C-P3666PCT liver disease; kidney disease; cardiovascular disease; Alzheimer’s disease; Parkinson’s disease; inflammatory bowel conditions, such as an including gut dysbiosis, coeliac disease, inflammatory bowel disease, irritable bowel syndrome ulcerative colitis, Chron’s disease, diverticulitis and Hirschsprung’s disease; obesity; epilepsy; diabetes; and motor neurone disease.

[0149] In further embodiments, the sample is one or more cells from a developing tissue. Such samples are analogous to those from diseased tissues / organs described hereinbefore and may therefore be obtained from biopsies of developing tissues or organs. Developing tissues / organs include embryonic tissue, extraembryonic tissue and tissues of the reproductive organs. However, in certain embodiments the methods described herein do not comprise the destruction of a human embryo. Thus according to these embodiments, the sample may be a stem cell-based embryo model (e.g. blastoids, gastruloids). A “blastoid” is an embryoid, a stem cell-based embryo model which, morphologically and transcriptionally resembles the early, pre-implantation mammalian conceptus (the blastocyst). The first blastoids were created by combining mouse embryonic stem cells and mouse trophoblast stem cells. Upon in vitro development, blastoids generate analogues of the hypoblast, thus comprising analogues of the three founding cell types of the conceptus (epiblast, trophoblast and hypoblast cells), and recapitulate aspects of implantation. However, blastoids do not display the capacity to support the development of a foetus and are therefore not considered as an embryo but as a model. Compared to other stem cell-based embryo models (e.g. gastruloids), blastoids model the preimplantation stage and the integrated development of the conceptus including the embryo proper and the two extraembryonic tissues (trophectoderm and hypoblast). Also according to these embodiments, the sample may be another stem cellbased model, such as a gastruloid, axioloid, peri-gastruloid, bilaminoid, E-assembloid or stem cell-based embryo model (SEM). In other embodiments, the sample may be a blastocyst. Blastocysts may be primary human blastocysts donated from a subject undergoing IVF treatment, e.g. surplus blastocyts from the IVF procedure. Such blastocyts may not be fertilised such that they cannot become viable foetuses. Thus in further embodiments, the developmental differentiation is embryogenesis and / or embryo model development.

[0150] In yet further embodiments, the sample is one or more cells from a differentiation culture, in particular an in vitro differentiation culture. Such cultures are known to the skilled person and include stem cell differentiation cultures, including reprogramming, directed differentiation and forward programming. Thus in some embodiments, the method identifies chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular state of developmental differentiation. In further embodiments, the developmental BAB-C-P3666PCT differentiation is stem cell differentiation. As described hereinbefore, such identified regions may thus be associated with developmental progression and / or developmental checkpoints. For example, wherein the culture is a directed differentiation or forward programming culture, the identified genomic region may be associated with the more differentiated state towards which the cells are being differentiated / programmed and / or with the progression through checkpoints of said differentiation / programming. Such genomic regions may be associated with the expression of lineage commitment genes as identified from the transcriptome. Alternatively, wherein the culture is a reprogramming culture or a culture maintaining stem cells in an undifferentiated state, the identified genomic region may be associated with a less differentiated / more pluripotent differentiation state and / or with the maintenance of differentiation / developmental checkpoints. Such genomic regions may be associated with the expression of pluripotency genes as identified from the transcriptome.

[0151] Thus, whether a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region is associated with a particular state of developmental differentiation is determined by its association with actively transcribed or transcriptionally repressed regions of the genome as identified from the transcriptome in the method herein. According to this embodiment, wherein an identified genomic region is associated with actively transcribed genes known to be associated with pluripotency, it can be identified as being associated with a less differentiated / more pluripotent developmental state. Conversely, wherein an identified genomic region is associated with actively transcribed genes known to be associated with lineage commitment, it can be identified as being associated with a more differentiated / lineage committed developmental state. Furthermore, wherein an identified genomic region is associated with transcriptionally repressed genes known to be associated with pluripotency, it can be identified as being associated with a more differentiated / lineage committed developmental state, or wherein an identified genomic region is associated with transcriptionally repressed genes known to be associated with lineage commitment, it can be identified as being associated with a less differentiated / more pluripotent developmental state.

[0152] Additionally, whether a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region is associated with a particular disease state is determined by its association with actively transcribed or transcriptionally repressed regions of the genome associated with said disease as identified by the transcriptome in the method herein. For example, wherein an identified genomic region is associated with actively transcribed genes known to be involved in disease (e.g. cancer driver genes / oncogenes), it can be identified as being associated with the development / progression of disease and / or its severity. Conversely, wherein an identified genomic region is associated with actively transcribed BAB-C-P3666PCT genes known to be involved in reversal of disease / a healthy state (e.g. tumour suppressor genes), it can be identified as being associated with the prevention of disease and / or lessening of its severity. Furthermore, wherein an identified genomic region is associated with transcriptionally repressed genes known to be associated with disease, it can be identified as being associated with the prevention of disease and / or lessening of severity, or wherein an identified genomic region is associated with transcriptionally repressed genes known to be associated with disease, it can be identified as being associated with the development / progression of disease and / or severity.

[0153] Thus, the method described herein comprises the step of: (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with actively transcribed or transcriptionally repressed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0154] The association between the chromatin marked, genome binding protein bound and / or modified genomic region and the actively transcribed or transcriptionally repressed genes may be determined qualitatively or quantified. Thus in a certain embodiment, the method additionally comprises quantifying the frequency of association of the chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with actively transcribed or transcriptionally repressed regions of the genome in step (ii). Quantification allows the comparison with control samples and / or data sets, such as those from non-diseased tissues / organs or from individuals without disease. Comparison may also be with differentiation cultures, organs or cells at different states of development. Such comparisons may therefore be considered as being to a ‘control’. Thus in a further embodiment, the frequency of association in the sample is compared with the frequency of associations in a sample obtained from a non-diseased individual or a non-diseased tissue or organ, or in a sample of one or more cells at a different state of developmental differentiation from a developing tissue or organ or from a differentiation culture in step (iii). The comparison of frequencies in the sample with a those in a ‘control’ allow, for example the determination of an association as being associated with disease if it is seen in the diseased sample but not in the healthy / non-diseased control. Furthermore, in differentiation cultures, comparison of frequencies allows the determination of an association as being associated with a more developed / differentiated state when it is seen in the differentiated sample but not in the less differentiated / more pluripotent control. Thus in a certain embodiment, the method additionally comprises quantifying the frequency of association of the chromatin marked, genome binding BAB-C-P3666PCT protein bound and / or modified genomic nucleic acid regions associated with actively transcribed or transcriptionally repressed regions of the genome in step (ii), and comparing the frequency of association in the sample with the frequency of associations in a sample obtained from a non-diseased individual or a non-diseased tissue or organ, or in a sample of one or more cells at a different state of developmental differentiation from a developing tissue or organ or from a differentiation culture in step (iii).

[0155] Kits

[0156] In a further aspect of the invention, there is provided a kit for identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample. Said kit is suitable for performing the methods described herein. In one embodiment, the kit comprises individual pre-assembled and blocked antibody-transposase-adapter complexes. In a further embodiment, each of said complexes comprises one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence. In a yet further embodiment, the kit comprises two or more individual preassembled and blocked antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence. In a still further embodiment, the kit comprises three or more, preferably four of more, five or more, six or more, seven or more or eight individual pre-assembled and blocked antibody- transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence. The antibody-transposase-adapter complexes may be as described hereinbefore. In one embodiment, the antibody-transposase- adapter complexes provided as part of the kit are blocked (i.e. pre-blocked) as described herein.

[0157] In further embodiments, the kit additionally comprises buffers and reagents capable of performing the methods described herein. The identity of such buffers and reagents will be readily recognised by the skilled person, but may include, without limitation storage buffers, blocker reagents (e.g. when the antibody-transposase-adapter complexes are not provided pre-blocked), tagmentation buffer, capture buffer, bait / probe sequences for mRNA transcript capture, reverse transcription buffers and reagents (e.g. reverse transcriptase enzyme), dNTPs suitable for incorporation into cDNA during reverse transcription and adapter sequences as described herein (e.g. for incorporation during reverse transcription or for BAB-C-P3666PCT ligating to tagged libraries, comprising barcode sequences and / or comprising sequencing adapter sequences for sequencing).

[0158] Thus in a certain embodiment, the kit comprises individual pre-assembled and blocked antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, in addition to buffers and reagents capable of performing the method described herein. In a further embodiment, the antibody- transposase-adapter complexes are optionally as described herein.

[0159] It will be appreciated that references herein to a patient, individual or subject relate equally to animals and humans and that the invention finds particular utility in veterinary treatment of any of the above mentioned diseases, disorders and conditions which are also present in said animals.

[0160] It will also be appreciated that references herein to “treatment” and “amelioration” include such terms as “prevention”, “reversal” and “suppression”. Similarly, references to “diagnosis” and “identification of disease state” may be used interchangeably herein and include terms such as “differential diagnosis”, “patient stratification” and the like. Furthermore, such terms include reference to the “patient”, “subject” and “individual” (e.g. an individual / subject in need of treatment) which may be used interchangeably herein. Treatment and / or diagnosis may be prior to the onset of the disease or disorder, e.g. wherein the subject is at risk of the disease or disorder. Alternatively, treatment and / or diagnosis may also be anticipated after the induction event of the injury, damage, disease or disorder, either before clinical presentation of said disease or disorder, or after symptoms manifest. Such references further include performing the methods of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state either prior to the onset of the disease, or after the induction event of the disease.

[0161] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. As used herein, the term “about” when used herein includes up to and including 10% greater and up to and including 10% lower than the value specified, suitably up to and including 5% greater and up to and including 5% lower than the value specified, especially the value specified. The term “between” as used herein includes the values of the specified boundaries. BAB-C-P3666PCT

[0162] Throughout the specification and the claims which follow, unless the context requires otherwise, the word “comprise”, and variations thereof such as “comprises” and “comprising”, will be understood to imply the inclusion of a stated integer, step, group of integers or group of steps but not to the exclusion of any other integer, step, group of integers or group of steps.

[0163] In addition, as used herein and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise (e.g. in a context where the term “single” is clearly required). Thus, for example reference to “an indexed adapter sequence” includes two or more such adapters as appropriate in the context herein (e.g. provided they comprise the same unique identifier if comprised in the same antibody- transposase-adapter complex, or may comprise different identifiers if comprised within different antibody-transposase-adapter complexes and reference to “an indexed adapter sequence” is used), or reference to “a cell” includes two or more such cells, unless specifically described otherwise (e.g. in the specific context of single cells).

[0164] It will be understood that all embodiments described herein may be applied to all aspects of the invention and vice versa, and such combinations would be readily apparent from the description provided herein and to those skilled in the art.

[0165] Other features and advantages of the present invention will be apparent from the description provided herein. It should be understood, however, that the description and the specific examples while indicating preferred embodiments of the invention are given by way of illustration only, since various changes and modifications will become apparent to those skilled in the art.

[0166] CLAUSES

[0167] A set of clauses defining the invention, its aspects and embodiments is as follows:

[0168] 1 . A method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising the steps of:

[0169] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0170] (ii) blocking the assembled antibody-transposase-adapter complexes; BAB-C-P3666PCT

[0171] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and

[0172] (iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0173] 2. The method of claim 1 , further comprising the steps of:

[0174] (iii b) capturing mRNA transcripts in the sample;

[0175] (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and

[0176] (iv b) sequencing the transcriptome library together with the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

[0177] 3. A method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising the steps of:

[0178] (i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;

[0179] (ii) blocking the assembled antibody-transposase-adapter complexes;

[0180] (iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification;

[0181] (iii b) capturing mRNA transcripts in the sample;

[0182] (iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and

[0183] (iv b) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and the transcriptome library.

[0184] 4. The method of any one of clauses 1 to 3, wherein the sample comprises one or more cells, such as a single cell, in particular wherein the sample is one or more cells.

[0185] 5. The method of any one of clauses 1 to 4, wherein the sample is a single cell. BAB-C-P3666PCT

[0186] 6. The method of any one of clauses 1 to 5, wherein the genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications are identified simultaneously in a single sample, such as in a single population of cells or in a single cell.

[0187] 7. The method of any one of clauses 1 to 6, wherein the transcriptome of the cell is identified simultaneously in the single sample, such as in the single population of cells or in the single cell.

[0188] 8. The method of any one of clauses 1 to 7, wherein two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i), each complex differentially comprising one or more antibodies recognising a single chromatin mark, genome binding protein or genomic modification of interest and a single indexed transposase-adapter sequence, and the two or more anti-chromatin mark, anti-genome binding protein and / or anti genomic modification antibody-transposase-adapter complexes are used together simultaneously in step (iii).

[0189] 9. The method of clause 8, wherein three or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii), in particular four or more, five or more, six or more, seven or more or eight anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled.

[0190] 10. The method of clause 8 or clause 9, wherein each anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complex comprises one or more antibodies recognising different chromatin marks, genome binding proteins or genomic modifications of interest and different indexed transposase-adapter sequences, such that each indexed transposase-adapter sequence is unique to each chromatin mark, genome binding protein and genomic modification of interest.

[0191] 11. The method of any one of clauses 8 to 10, wherein following step (iii), one or more further anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are used sequentially or together simultaneously for further in situ single step fragmentation and adapter sequence insertion in the sample. BAB-C-P3666PCT

[0192] 12. The method of any one of clauses 1 to 11 , wherein the chromatin marks are histone modifications, such as phosphorylation, acetylation, ubiquitination, GIcNAcylation, citrullination, crotonylation, sumoylation, isomerisation and / or methylation, including monomethylation, di-methylation and tri-methylation.

[0193] 13. The method of any one of clauses 1 to 12, wherein the anti-chromatin mark antibody recognises and binds a histone modification and / or a modified histone selected from any of: H3K4me1 , H3K4me2, H3K4me3, H3K36me3, H3K79me2, H3K9Ac, H3K27AC, H4K16AC, H3K27me1 , H3K27me3, H3K27me3, H2AK119ub, H3K122ac and H3K9me3.

[0194] 14. The method of clause 12 or clause 13, wherein the histone modification is associated with euchromatin or heterochromatin, in particular euchromatin.

[0195] 15. The method of any one of clauses 1 to 14, wherein the indexed transposase-adapter sequence comprises protein A, protein G, protein A / G or protein L, with a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence.

[0196] 16. The method of clause 15, wherein the indexed transposase-adapter sequence comprises protein A, a transposases enzyme and a transposase adapter sequence.

[0197] 17. The method of clause 15 or clause 16, wherein blocking step (ii) comprises blocking free protein A, protein G, protein A / G or protein L in the assembled antibody-transposase- adapter complexes, such as using free IgG antibodies.

[0198] 18. The method of any one of clauses 1 to 14, wherein the indexed transposase-adapter sequence comprises streptavidin, a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence, and wherein the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies are biotinylated.

[0199] 19. The method of clause 18, wherein blocking step (ii) comprises blocking free streptavidin in the assembled antibody-transposase-adapter complexes, such as using free biotin.

[0200] 20. The method of any one of clauses 1 to 14, wherein the indexed transposase-adapter sequence is directly conjugated to the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies. BAB-C-P3666PCT

[0201] 21 . The method of clause 20, wherein blocking step (ii) is omitted.

[0202] 22. The method of any one of clauses 1 to 21 , wherein in situ single step fragmentation and adapter sequence insertion in step (iii) comprises tagmentation.

[0203] 23. The method of any one of clauses 1 to 22, wherein the transposase is Tn5 transposase.

[0204] 24. The method of clause 23, wherein the Tn5 transposase is a mutant Tn5 transposase such as hyperactive Tn5 transposase.

[0205] 25. The method of any one of clauses 1 to 24, wherein blocking step (ii) reduces off-target binding of the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes.

[0206] 26. The method of clause 25, wherein two or more antibody-transposase-adapter complexes are used together simultaneously in step (iii) as defined in clauses 8 to 10.

[0207] 27. The method of clause 25 or clause 26, wherein two or more antibody-transposase- adapter complexes are used in further in situ single step fragmentation and adapter sequence insertion in the sample as defined in clause 11.

[0208] 28. The method of any one of clauses 1 to 27, wherein the indexed transposase-adapter sequence comprises only Mosaic End B (MEB) adapter sequences, and the method additionally comprises adding Mosaic End A (MEA) adapter sequences to the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library following step (iii).

[0209] 29. The method of clause 28, wherein the indexed transposase-adapter sequence comprising only Mosaic End B (MEB) adapter sequences and the adding of Mosaic End A (MEA) adapter sequences results in the sequencing of all of the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library in step (iv) and optionally step (iv b) and / or prevents the generation of chimeric sequence reads in the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification- associated sequence library. BAB-C-P3666PCT

[0210] 30. The method of any one of clauses 2 to 29, wherein the transcriptome comprises mRNA transcripts of actively transcribed regions of the genome of the cell, such as of actively expressed genes, in particular actively transcribed regions of the genome associated with euchromatin.

[0211] 31. The method of any one of clauses 2 to 30, wherein capturing mRNA transcripts in step (iii b) is performed using a probe or bait sequence.

[0212] 32. The method of clause 31, wherein the probe or bait sequence is a poly-T sequence and / or sequences complementary to mRNA transcripts of interest.

[0213] 33. The method of any one of clauses 2 to 32, wherein step (iv) is performed after the single step fragmentation and adapter sequence ligation of step (iii).

[0214] 34. The method of cany one of clauses 2 to 33, wherein step (iv b) is performed after in situ reverse transcription of the captured mRNA transcripts of step (iii c).

[0215] 35. The method of any one of clauses 1 to 34, wherein step (iii) is performed on permeabilised cells and / or intact cell nuclei.

[0216] 36. The method of any one of clauses 2 to 35, wherein steps (iii b) and (iii c) are performed on permeabilised cells and / or intact cell nuclei.

[0217] 37. The method of clause 35 or clause 36, wherein the cells and / or cell nuclei are crosslinked.

[0218] 38. The method of clause 37, wherein the method additionally comprises lysing the cells and / or nuclei and reversing cross-linking following step (iii).

[0219] 39. The method of clause 37 or clause 38, wherein the method additionally comprises lysing the cells and / or nuclei and reversing cross-linking following step (iii c).

[0220] 40. The method of any one of clauses 1 to 39, wherein the method additionally comprises splitting the sample following step (iii) into distinct samples and adding unique molecular identifiers to each distinct sample, followed by re-pooling the uniquely identified distinct samples. BAB-C-P3666PCT

[0221] 41. The method of any one of clauses 2 to 40, wherein the method additionally comprises splitting the sample following step (iii c) into distinct samples and adding unique molecular identifiers to each distinct sample, followed by re-pooling the uniquely identified distinct samples.

[0222] 42. The method of clause 40 or clause 41 , wherein the splitting and re-pooling is repeated one or more times, in particular repeated two times.

[0223] 43. The method of any one of clauses 1 to 42, wherein the method additionally comprises amplifying the adapter-tagged chromatin mark-associated sequence library and the transcriptome library prior to step (iv).

[0224] 44. The method of any one of clauses 2 to 43, wherein the method additionally comprises amplifying the adapter-tagged chromatin mark-associated sequence library and the transcriptome library prior to step (iv b).

[0225] 45. The method of clause 43 or clause 44, wherein amplification is by PCR.

[0226] 46. The method of any one of clauses 1 to 45, wherein the method additionally comprises amplifying and / or selecting a target sequence of interest, optionally within the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and / or the transcriptome library, prior to step (iv) and optionally step (iv b).

[0227] 47. The method of clause 46, wherein said amplifying and / or selecting comprises one or more primer sequences complementary to a region surrounding the target sequence of interest.

[0228] 48. The method of clause 46 or clause 47, wherein the primer sequences further comprise a selection marker, such as biotin.

[0229] 49. The method of any one of clauses 46 to 48, wherein said amplification is by PCR.

[0230] 50. A method of determining the sequence of a target sequence in a sample, said method comprising performing the method of any one of clauses 1 to 45 and further amplifying or selecting genomic nucleic acid regions or transcriptome library sequences comprising the target sequence prior to step (iv) and optionally step (iv b). BAB-C-P3666PCT

[0231] 51. A method of genotyping a target sequence in a sample, said method comprising performing the method of any one of clauses 1 to 45 and further amplifying or selecting genomic nucleic acid regions or transcriptome library sequences comprising the target sequence prior to step (iv) and optionally step (iv b).

[0232] 52. The method of clause 50 or clause 51, wherein amplifying and / or selecting a target sequence comprises amplifying and / or selecting the target sequence within the adapter- tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and / or the transcriptome library.

[0233] 53. The method of any one of clauses 50 to 52, wherein amplifying and / or selecting comprises one or more primer sequences complementary to a region surrounding the target sequence of interest.

[0234] 54. The method of clause 53, wherein the primer sequences further comprise a selection marker, such as biotin.

[0235] 55. The method of any one of clauses 50 to 54, wherein said amplification is by PCR.

[0236] 56. The method of any one of clauses 46 to 55, wherein the sequence of the target sequence is determined during sequencing step (iv) and optionally step (iv b).

[0237] 57. A method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0238] (i) performing the method of any one of clauses 1 to 49 on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0239] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with actively transcribed regions of the genome, such as actively expressed genes, wherein said actively transcribed regions are identified by sequencing the transcriptome library and wherein said actively transcribed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0240] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with actively transcribed BAB-C-P3666PCT regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0241] 58. A method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:

[0242] (i) performing the method of any one of clauses 1 to 49 on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;

[0243] (ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with transcriptionally repressed regions of the genome, wherein said transcriptionally repressed regions are identified by sequencing the transcriptome library and wherein said transcriptionally repressed regions are associated with the particular disease state and / or state of developmental differentiation; and

[0244] (iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with transcriptionally repressed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

[0245] 59. The method of clause 57 or clause 58, additionally comprising quantifying the frequency of association of the chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with actively transcribed or transcriptionally repressed regions of the genome in step (ii), and comparing the frequency of association in the sample with the frequency of associations in a sample obtained from a non-diseased individual or a non-diseased tissue or organ, or in a sample of one or more cells at a different state of developmental differentiation from a developing tissue or organ or from a differentiation culture in step (iii).

[0246] 60. A method of predicting a subject’s risk of having and / or developing a disease or condition associated with a mutation, said method comprising determining the sequence of and / or genotyping a target sequence in a sample obtained from the subject as defined in any one of clauses 46 to 56, wherein the presence of a mutation associated with the disease or condition in the target sequence indicates the subject is at risk of or has said disease or condition. BAB-C-P3666PCT

[0247] 61 . The method of clause 60, further comprising administering to the subject a treatment for the disease or condition if a mutation associated with said disease / condition is present.

[0248] 62. The method of clause 61 , wherein the treatment for the disease or condition is selected based on the identification of the mutation.

[0249] 63. The method of any one of clauses 57 to 62, wherein the disease is cancer, an autoimmune disease, a developmental disease and / or a genetic disorder, such as a cancer associated with epigenetic modifications.

[0250] 64. The method of clause 63, wherein the cancer is associated with a mutation causing the activation of an oncogene and / or the inactivation of a tumour suppressor.

[0251] 65. The method of any one of clauses 57 to 63, wherein the developmental differentiation is stem cell differentiation, embryogenesis and / or embryo model development.

[0252] 66. A kit for identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said kit comprising individual pre-assembled and blocked antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, in addition to buffers and reagents capable of performing the method of any one of clauses 1 to 56.

[0253] 67. The kit of clause 66, wherein the antibody-transposase-adapter complexes are as defined in any one of clauses 8 to 29.

[0254] The invention will now be described using the following, non-limiting examples:

[0255] EXAMPLES

[0256] Materials & Methods

[0257] Biological Samples

[0258] Cell Lines: Human pluripotent stem cells were grown in 5% CO2, 5% O2 at 37°C and maintained in Essentials medium (ThermoFisher Scientific) on 0.5% Geltrex (ThermoFisher Scientific). For definitive endoderm induction, undifferentiated hPSCs were seeded at a density of 15,000 cells per cm2on Geltrex in Essentials medium. The next day, cells were washed with 0.1% BSA in DM EM, and then cultured in endoderm induction medium comprised BAB-C-P3666PCT of 50% IMDM (ThermoFisher Scientific), 50% F-12 Nutrient Mix (ThermoFisher Scientific), 1 % Chemically Defined Lipids (ThermoFisher Scientific), 0.1% BSA Fraction V (ThermoFisher Scientific), 15pg / mL Transferrin (Merck), 450pM Monothioglycerol (Merck), 0.7pg / mL Insulin (Merck), 100ng / mL Activin-A (Cambridge Stem Cell Institute), 100nM PI-103 (Tocris Bio- Techne), 3pM CHIR99021 (Cambridge Stem Cell Institute), 10ng / mL FGF2 (Cambridge Stem Cell Institute), 3ng / mL BMP4 (R&D Systems) and 10pg / mL Heparin (Merck). After 24 hours, cells were cultured in modified endoderm induction medium whereby CHIR99021 and BMP4 were removed, the concentration of FGF2 was increased to 20ng / mL, and LDN193189 (Axon Medchem) was added to 250nM.

[0259] Mouse Embryos: All mouse experimentation was approved by the Babraham Institute Animal Welfare and Ethical Review Body. Animal husbandry and experimentation complied with existing European Union and United Kingdom Home Office legislation and local standards. Mice were bred and maintained in the Babraham Institute Biological Support Unit. B6CBAF1 mice bred in-house were used experimentally between 8- and 12-weeks of age. Female mice were superovulated by injection with 5 IU of pregnant mare serum gonadotropin followed with 5 IU human chorionic gonadotropin 46 hours later. Embryos at approximately E3.5 were collected by flushing the uterus with M2 medium (Sigma-Aldrich) 94- to 98-hours after human chorionic gonadotropin injection. Embryos were cultured in KSOM medium (Millipore, MR-121-D) for 24 hours at 37°C in a 5% CO2 incubator under light mineral oil (Millipore, ES-005-C). If present, the zona pellucidae was removed from embryos using Acidic Tyrode’s solution (Sigma) and washed through three drops of M2 media. Embryos were washed in PBSI (PBS without calcium or magnesium supplemented with 1x Protease Inhibitor Cocktail [PIC], 0.05 U / pL Superaseln RNase Inhibitor, 0.1 U / pL Protector RNAse Inhibitor and 0.2% BSA) and placed in 1.5 mL tubes containing PBSI. Embryos were incubated in TrypLE Express (ThermoFisher Scientific) for 10 min at 37°C. Embryos were washed with PBSI and dissociated into single cells by pipetting. Nuclei were prepared as described in the ‘Sample Preparation’ section and stored at -80°C until use.

[0260] Assembly of ProteinA-Tn5 with MEDS: To assemble proteinA-Tn5 with MEDS, 2 pL of 200pM phosphorylated Mosaic end adapter B dU (MEBdU) with different indexes (5 bp) or Mosaic end adapter A (MEA) was mixed with 2 pL of 200 pM blocked Mosaic end-reverse (MER) and annealed using the following programme: 95°C for 5 min, cooled down to 20°C with ramp rate of -1°C / min. Annealed adapters (4 pL) were mixed with 20 pL of proteinA-Tn5 (~5.5 pM) and incubated on a rotator at room temperature for 1 hour and then stored at -20°C for up to 1 year. BAB-C-P3666PCT

[0261] Sample Preparation: Samples were dissociated into single cells and washed twice with PBSI (DPBS without calcium or magnesium, 1x PIC, 0.05 U / pL Superaseln RNase Inhibitor, 0.1 U / pL Protector RNAse Inhibitor, 0.04% BSA). Cells were collected by centrifugation in a swing-out rotor at 300 x g for 3 min. Liquid was carefully aspirated, leaving ~20 pL remaining in the tube. Samples were mixed with the same volume of 2x Nuclei Extraction (NE) buffer (40 mM HEPES-KOH pH 7.9, 20 mM KCI, 0.5 mM spermidine, 0.2% Triton-X100, 40% Glycerol, 2x PIC, 0.25 LI / pL Superaseln RNase Inhibitor, 0.5 LI / pL Protector RNAse Inhibitor) and incubated on ice for 10 min. Nuclei were fixed with 0.2% formaldehyde at room temperature for 5 min and quenched with 1x Fixation Quench Buffer (5x Fixation Quench Buffer: 625 mM Glycine, 250 mM Tris-HCL pH 8.0, 0.1 % Trixton-X100, 0.5% BSA in DPBS without calcium or magnesium) on ice for 5 min. Nuclei were collected by centrifugation in a swing-out rotor at 4°C, 600 x g for 3 min and resuspended in 100 pL Wash buffer (20 mM HEPES-NaOH pH 7.5, 150 mM NaCI, 0.5 mM spermidine, 1 % BSA, 1% PIC, 0.125 U / pL Superaseln RNase Inhibitor, 0.25 U / pL Protector RNAse Inhibitor). Nuclei were either used fresh or, if they were to be frozen, DMSO was added to a final concentration of 10%, and samples were placed in a slow freezing container and stored at -80°C.

[0262] In situ DNA Tagmentation: To combine the selected antibody with the indexed proteinA- Tn5-MEBdll, 0.5 pL antibody (1 pg / pL) and 0.5 pL proteinA-Tn5 were mixed with 4 pL Complete Wash 300 buffer (20 mM HEPES-NaOH pH 7.5, 300 mM NaCI, 0.5 mM spermidine, 1 % BSA, 1% PIC, 0.125 U / pL Superaseln RNase Inhibitor, 0.25 U / pL Protector RNAse Inhibitor) and incubated on an end-over-end rotator at room temperature for 1 hour. Rabbit IgG antibody (1 pL of 1 pg / pL) was added to each antibody-ProteinA-Tn5-MEBdll complex to block any uncombined proteinA-Tn5 and incubated at room temperature for 30 min. Blocked, preassembled antibody-proteinA-Tn5-MEBdll were pooled together and brought to a final volume of 100 pL with Complete Wash 300 buffer in the presence of 2 mM EDTA.

[0263] Preactivated paramagnetic concanavalin A (ConA; 20-100 pL) beads were added to the cell samples (either freshly prepared or thawed from frozen, as described above) and incubated at room temperature for 10 min. Nuclei bound to the ConA beads were washed with Wash buffer containing 2 mM EDTA and then incubated with blocked, preassembled antibody- proteinA-Tn5-MEBdll on a rotator at RT for 1 hour or overnight at 4°C. Cells were washed three times with modified Complete Wash 300 buffer (20 mM HEPES-NaOH pH 7.5, 300 mM NaCI, 0.5 mM spermidine, 1% BSA, 1% PIC, 0.05 U / pL Superaseln RNase Inhibitor, 0.1 U / pL Protector RNAse Inhibitor). Cells were suspended in 100 pL Tag buffer (20 mM HEPES- NaOH pH 7.5, 300 mM NaCI, 0.5 mM spermidine, 0.1% BSA, 1% PIC, 0.25 U / pL Superaseln RNase Inhibitor, 0.5 U / pL Protector RNAse Inhibitor and 10 mM MgCI2) and incubated at 37°C BAB-C-P3666PCT for 1 hour. To stop tagmentation, samples were mixed with 3.34 pL of 0.5 M EDTA and 16 pL of 7.5% BSA and incubated on a rotator at RT for 5 min.

[0264] For sequential tagmentation, samples were washed twice with modified Complete Wash 300 buffer supplemented with 2 mM EDTA before incubated with preassembled antibody-proteinA- Tn5-MEBdll for the second round tagmentation. The same steps of DNA tagmentation were repeated as described above for the first round tagmentation. Further rounds of DNA tagmentation were repeated as required.

[0265] In situ Reverse Transcription: Tagmented nuclei were washed twice with 100 pL NIB-RI buffer (10 mM Trisma buffer pH 7.5, 10 mM NaCI, 3 mM MgCI2, 0.1% NP-40, 1 % BSA, 0.05 LI / pL Superaseln RNase Inhibitor, 0.1 LI / pL Protector RNAse Inhibitor, 1x PIC) then resuspended in 50 pL RT mix (1x RT buffer, 15% PEG6000, 10 pM RT primer, 0.5 mM dNTPs, 0.5 LI / pL Superaseln RNase Inhibitor, 0.5 LI / pL Protector RNAse Inhibitor, 20 LI / pL Maxima H Minus RT). Reverse transcription was carried out by incubation at 50°C for 10 min and then three cycles of annealing (8°C 12s, 15°C 45s, 20°C 45s, 30°C 30s, 42°C 120s and 50°C 180s), followed by 50°C for 5 min. Samples were diluted with 50 pL NIB-RI buffer and washed once with 100 pL NIB-RI buffer.

[0266] Preparing Oligonucleotides for Single Cell Barcoding: Linker oligonucleotides and barcode oligonucleotides for hybridisation were prepared in RNase-free 96-well plates (Twin- Tec PRC plate, 96 Lobind, Semi-skirted) with the combined oligonucleotides for each round of hybridisation prepared in separate plates. Round 1 linker (final concentration 9 pM) and BC1 barcodes (final concentration 10 pM) were combined in a total volume of 10 pl per well in STE buffer (10 mM Tris pH 8.0, 50 mM NaCI, 1 mM EDTA). In a new plate, round two hybridisation reagents combined 11 pM round 2 linker and 12 pM BC2 barcodes in 10 pL STE buffer per well; and in a new plate, round three hybridisation reagents combined 13 pM round 3 linker and 14 pM BC3 barcodes in 10 pl STE buffer per well. Oligonucleotides were annealed in a PCR machine using the programme: 95 °C for 5 min, and cooling to 20°C at -1°C / min. Annealed oligonucleotides were stored at -20°C.

[0267] Split-Pool Barcoding and Ligation: Samples after in situ reverse transcription were resuspended in 2 mL hybridisation mix (1x T4 ligation buffer, 0.05 U / pL Superaseln RNase Inhibitor, 0.32 U / pL Protector RNAse Inhibitor, 1x PIC, 0.1% Triton-X100, 0.1 % BSA, 0.25x NIB). Samples were then split and pooled for a total of three rounds with 48 barcodes in each round. Cells (40 pL) were aliquoted into each well of the Round 1 plate (containing 10 pL of annealed round 1 linker and a unique BC1 , as described above) and incubated at room BAB-C-P3666PCT temperature for 30 min in a thermomixer with agitation at 300 rpm. Then, 10 pL Blocker 1 solution (22 pM blocker 1 , 2x T4 ligation buffer) was added to each well of the Round 1 plate and incubated at room temperature for 30 min in a thermomixer with agitation at 300 rpm. All cells in the plate were then combined and mixed. Cells (55 pl) were aliquoted into each well of the Round 2 plate (containing round 2 linker and BC2) and the same procedure was followed as above, expect that Blocker 2 solution contained 26.4 pM blocker 2, 2x T4 ligation buffer. Following this second round of barcoding and pooling, 70 pL of cells were aliquoted into each well of the Round 3 plate (containing round 3 linker and BC3). Blocker 3 solution contained 23 pM blocker 3, 0.1 % Triton-X100. After Blocker 3 solution was added to the plate, cells were combined and centrifuged at 4°C, 600xg for 3 min. Nuclei were transferred into an Eppendorf protein lobind tube and washed twice with 0.5 mL NIB buffer. Nuclei were resuspended in 0.5 mL Ligation buffer (1x T4 ligation buffer, 0.05 U / pL Superaseln RNase Inhibitor, 0.32 U / pL Protector RNAse Inhibitor, 1x PIC, 0.1 % Triton-X100, 0.1 % BSA, 0.2x NIB, T4 DNA ligase 20 U / pL) and incubated at 25°C for 30 min in a thermomixer with agitation at 300rpm. The tube containing the nuclei was attached to a magnetic rack and the nuclei were washed with 0.5 mL NIB and then resuspended in 0.5 mL NIB and passed through a 30 pm strainer. The nuclei were counted with a haemocytometer and aliquoted into multiple tubes to generate sublibraries, with a defined number of cells per sublibrary based on the estimated collision rate (RT / MEBdU x BC1 x BC2 x BC3).

[0268] Reverse Crosslinking and Lysis: The volume of each sublibrary was brought to 20 pL with NIB buffer. Twenty microliters of 2x reverse crosslinking buffer (100 mM Tris pH 8.0, 100 mM NaCI, 0.04% SDS, 2 pL of 20 mg / mL proteinase K, 1.6 pL of Rl mix [SUPERase RkProtector Rl 1 :1]) were mixed with each sample and incubated at 55°C for 1 hour in a PCR machine. PMSF (2.5 pl of 100 mM) was added to the reverse crosslinked sample to inactivate proteinase K and incubated at room temperature for 10 min. Samples were attached to a magnetic rack and the supernatant was transferred into a new tube, leaving the ConA beads behind.

[0269] Separation of cDNA from gDNA: MyOne C1 Streptavidin Dynabeads (10 pL) were washed with 800 pL 1 x B&W-T buffer three times and resuspended in 45 pL 2x B&W buffer with 1 pL SUPERase Rl. Prepared C1 beads were mixed with lysed, reverse crosslinked samples and incubated on an end-to-end rotator at 10 rpm for 60 min at room temperature to bind the biotin labeled cDNA. The tubes were then attached to a magnetic rack to separate tagmented gDNA (in supernatant) from cDNA (attached to the beads).

[0270] DNA Library Amplification: For DNA library preparation, tagmented gDNA was purified with 1.2x volume of Ampure XP beads and eluted with 21 pL 0.1x EB buffer. Eight microliters of BAB-C-P3666PCT

[0271] 2x NEBNext master mix was combined with 20 pL eluted sample from the supernatant and incubated at 72°C for 10 min in a PCR machine. Three microliters of 1 pM MEALNA was added to the mixture to perform adapter switching with the following programme: 98°C for 30s; 10 cycles of 98°C for 10s, 59°C for 20s, and 72°C for 10s. After this step, 2 pL of 25 pM indexed P7, 2 pL of 25 pM indexed Ad1 , 50 pL Q5U master mix and 15 pL H2O were added and the libraries were amplified with the following programme: 98°C for 30s; 16 cycles of 98°C for 10s, 55°C for 20s, and 72°C for 30s; and 72°C for 3 min. PCR products were purified with 0.7x Ampure XP beads and eluted with 32 pL EB. The concentration of libraries was measured with a Qubit. Samples were run on a Bioanalyzer to check the size distribution of library fragments. Libraries were sequenced on Illumina NovaSeq 6000, Nextseq 500 or Element Biosciences AVITI instruments at the Babraham Institute Genomics Facility and the CRLIK Cambridge Institute Genomics Facility.

[0272] RNA Library Preparation:

[0273] Template switching: cDNA bound to the C1 beads was washed three times with 100 pL 1 x B&W-T buffer containing SUPERase inhibitor and once with 100 pL STE buffer containing SUPERase inhibitor. Beads were then resuspended in 50 pL Template Switching mix (10 pL 5x RT buffer, 10 pL 20% Ficoll PM-400, 15 pL 50% PEG6000, 1.25 pL 100 pM TSO, 2 pL 25mM dNTPs, 0.625 pL Protector RNase inhibitor, 0.625 pL Superase RNase inhibitor, 2.5 pL 200U / pL Maxima H Minus RT, 8 pL H2O). The reaction was incubated on an end-to-end rotator at 10 rpm for 30 min at room temperature and then incubated for 90 min at 42°C in a PCR machine with mixing every 30 min.

[0274] Pre-amplification: After template switching, the cDNA-bound C1 beads were diluted with 100 pL of STE and washed with 200 pL of STE without disturbing the bead pellet. Then, the beads were resuspend in 50 pL pre-Amplification mix (25 pL Kapa HiFi PCR mix, 2 pL of 10 pM PA- F, 2 pL of 10 pM PA-R, and 21 pL H2O) and cycled in a PCR machine using following programme: 95°C for 3 min; 11 cycles of 98°C for 30s, 65°C for 45s and 72°C for 3 min; then 72°C for 5 min. PCR products were purified by 0.8x AM Pure XP beads and eluted to 32 pL 0.1x EB. The concentration of libraries was measured with a Qubit. Samples were run on a Bioanalyzer to check the size distribution of library fragments.

[0275] Tagging MEA and PCR amplification: MEA adapters were next added to the pre-amplified RNA library. Libraries (50 ng) were mixed with 2 pL of 1 :20 diluted blocked Tn5-MEAin 5 pL of 4x Tagmentation buffer, and H2O was added to a final volume of 20 pL. Samples were incubated at 37°C for 30 min. Two microliters of 0.2% SDS were added to the reaction and incubated at room temperature for 10 min to release tagged fragments. One microliter of 4% BAB-C-P3666PCT

[0276] TritonXIOO was added to neutralise the SDS before being mixed with 25 pL of 2x NEBnext master mix, 1 pL of 25 pM indexed Ad1 , 1 pL of 25 pM indexed P7. PCR was performed with the following programme: 72°C for 10 min, 98°C for 3 min, 11 cycles of 98°C for 10s, 65°C for 30s and 72°C for 1 min; and 72°C for 5 min. PCR products were purified using 0.7* AMpure beads and eluted with 32 pL of EB. The concentration of libraries was measured with a Qubit. Samples were run on a Bioanalyzer to check the size distribution of library fragments. Libraries were sequenced on Illumina NovaSeq 6000, Nextseq 500 or Element Biosciences AVITI instruments at the Babraham Institute Genomics Facility and the CRLIK Cambridge Institute Genomics Facility.

[0277] Data Processing: After sequencing, cell barcode, sample ID, antibody ID and UMI information was extracted from Read2 with modified Reachtools. Cell barcodes (including sample ID and antibody ID) were mapped to all possible cellular barcode combinations using Bowtie. Reads from Readl and / or the residual Read2 were trimmed with TrimGalore and mapped to GRCh38 for human samples and mm 10 for mouse samples using Bowtie2 for DNA data and STAR for RNA data. Duplicates were removed by the mapped portion and UMIs. RNA data were feature counted by splitpoolquantitation. DNA data were counted within 5kb nonoverlapping bins or other length as specified. Deeptools was used to generate bigwig files with bamCoverage and bigwigCompare for visualisation in Genome browser with IGV tools. ComputMatrix, plotProfile and plotHeatmap functions from Deeptools were used to plot genome-wide signal of DNA data with the set of regions as specified. SEACR was used for peak calling. Single cell datasets were loaded to Seurat and Signac for single cell clustering, marker identification (fold change>4) and dimensional reduction. ChromHMM was used for chromatin states analyses with bulk data or aggregated single cells datasets. Chromatin states were annotated based on the emission probability of combined histone modifications and the distance to TSSs. SCENIC+ was used for inferring gene regulatory networks with single cell H3K27ac data and its corresponding transcriptome data. Gene regulatory networks were visualised with Cytoscape.

[0278] Example 1: Combined Profiling of Multiple Histone Modifications and Transcriptome Modalities

[0279] To profile multiple histone modifications and transcriptome in the same single cells, primary antibodies specific for each target histone modification were pre-assembled with different indexed protein A-Tn5-adapters, followed by in situ Tn5-mediated tagmentation with these indexed complexes, and capture of mRNA with a poly-T primer with in situ reverse transcription (Figure 1A). A series of optimisation steps were performed to increase the robustness and sensitivity of the method as described below. BAB-C-P3666PCT

[0280] First, standard Tn5-mediated DNA tagmentation requires genomic fragments to be simultaneously tagged by two different adapters (MEA and MEB), which results in half of the tagmented fragments being undetectable due to incompatible adapters (as described hereinbefore). Therefore, an adapter switching strategy was implemented that uses only the MEB adapter for antibody-specific tagmentation and then adds the MEA adapter on to the end of all MEB-tagged fragments (Figure 1A). This approach additionally avoids the possibility of generating confounding chimeric sequence reads that can arise during multi-target chromatin profiling when MEA and MEB adapters are tagmented into a genomic locus by two different antibody-Tn5 complexes. Using the adapter switching strategy, all Tn5-tagged fragments can be detected in the final sequencing library, thereby substantially increasing library complexity.

[0281] It was next tested whether the adapter-switching step affects the accurate detection of histone modification signals. H3K27me3 and H3K27ac in human pluripotent stem cells (hPSCs) were chosen to profile because these two histone modifications are mutually exclusive on chromatin and so errors in their detection would be easily identifiable. It was found that including an adapter-switching step does not affect the profiled localisation of the histone modifications (Figures 1 B-1 D). Furthermore, including adapter-switching improved the signal-to- background ratio, as shown by the higher fraction of reads in peaks (FriP; H3K27me3, 69% without adapter switching vs 76% with adapter switching; H3K27ac, 11% in without adapter switching vs 14% with adapter switching; Figure 1 E). Down-sampling analysis was also performed to test how many unique reads could be detected at the same sequencing depth. Notably, Cut&Tag with adapter switching detected more unique reads in both libraries, indicating Cut&Tag with adapter switching libraries achieved higher library complexity (H3K27me3, estimated library size is 4 million reads without adapter switching vs 12 million reads with adapter switching; H3K27ac, estimated library size is 41 million reads in without adapter switching vs 94 million reads with adapter switching; data not shown).

[0282] Second, to assess the performance of the method in mapping multiple histone modifications, different conditions for simultaneously profiling H3K27me3 and H3K27ac in hPSCs were compared. H3K27me3 and H3K27ac on-target peaks were successfully captured (comparing with individual profiles, adjusted R-squared for H3K27me3, -0.98; H3K27ac, 0.87-0.91 ; Figure 1 B-1 D). However, consistent with previous reports, directly combining multiple preassembled antibody-proteinA-Tn5-adapter complexes to simultaneously tagment cells resulted in off-target signals, as H3K27ac-assigned reads aberrantly overlapped with H3K27me3 signal; a problem that was not alleviated in conditions with extended preincubation times (Figures 1B-1 D). Notably, only H3K27ac signal was detected in H3K27me3 reads and BAB-C-P3666PCT not the other way around, indicating directional off-target tagmentation that was not simply caused by sequencing errors (Figures 1B-1C).

[0283] It was hypothesised that the off-target signals might be caused by excess free protein A after assembly with histone modification antibodies. In line with this prediction, it was found that adding IgG blocking antibodies to the post-assembled protein A-antibody mixture successfully prevented off-target signals (Figures 1 B-1C) and increased the on-target FRiP (from 9% to 23% for H3K27me3; from 12% to 14% for H3K27ac) (Figure 1E). Although adding more IgG helped to further reduce off-target signals (Figures 1 B-1C) it did not further improve the detection of H3K27me3 and H3K27ac enriched signals on-target peaks (Figure 1 E).

[0284] Third, to capture transcriptome information together with multi-target chromatin binding, it was assessed whether transcriptome profiling and chromatin profiling would affect each other. It was found that performing mRNA reverse transcription prior to DNA tagmentation or after DNA tagmentation did not affect transcriptome profiles (pairwise correlations >0.99) or sensitivity (similar numbers of genes detected). However, performing reverse transcription prior to DNA tagmentation was detrimental to chromatin profiling with increased background signal (Figures 1B, 1 D and 1 F). Therefore, it was chosen to perform reverse transcription after DNA tagmentation in all subsequent multi-omic experiments.

[0285] Lastly, to reduce the amount of required starting material and to increase the recovery rate, nuclei preparation steps were optimised, washing steps reduced and nuclei immobilised with conA beads. Taken together, an efficient method for jointly profiling transcriptome and multiple histone modifications, with low background signal, good reproducibility and high sample recovery has been established. The method is referred to as scMTR-seq.

[0286] Example 2: Benchmarking of scMTR-seq to Previously Described Methods

[0287] To compare scMTR-seq as described herein with previously described profiling methods, the key metrics were compared and the summary is shown in Table 1.

[0288] Comparative data is shown in Figures 2A and 2B. These results demonstrate that scMTR- seq has higher levels of specific and on-target signals compared to CoTarget, NTT-seq or MulTI-Tag. BAB-C-P3666PCT

[0289] Table 1: Key Metrics of scMTR-seq and Previous Profiling Methods Described Herein

[0290] * represents estimated cell numbers

[0291] BAB-C-P3666PCT

[0292] Example 3: scMTR-seq Profiles Six Targets and the Transcriptome in Single Cells

[0293] Next it was sought to combine the multi-omic profiling with a highly scalable and high- throughput approach applicable to single cells. To achieve this, following in situ DNA tagmentation and reverse transcription on intact nuclei, three rounds of split-pool combinatorial barcoding were applied (Figure 3A). With 48 unique barcodes for each barcoding round, over 110,000 unique barcode combinations can be generated after three rounds of barcoding, which is sufficient to uniquely label 1 ,100 to 11 ,000 single cells with an estimated collision rate of 1 % to 10%. Although this barcoding capacity was sufficient for the present experiments, capacity could be scaled up using multiple sub-libraries, providing the opportunity to profile millions of single cells in one experiment.

[0294] To test the robustness of our method, scMTR-seq was applied to simultaneously profile six targets (five histones including H3K4me1 , H3K4me3, H3K27ac, H3K27me3 and H3K36me3, and an IgG control) and the transcriptome in conventional hPSCs. Starting with 18,000 cells, final data sets on 7,479 individual cells were successfully obtained after stringent quality control filters for both DNA (unique reads per cell >2,500) and RNA (unique reads per cell >2,000), thereby achieving a recovery rate of >40% for combinatorial histone modification and transcriptome profiling. With a sequencing depth of 75,000 total raw reads per cell for DNA and 22,000 raw reads per cell for RNA, on average 9,699 unique reads were obtained per cell for DNA (735 reads for H3K27me3, 3,024 for H3K27ac, 2,543 for H3K4me1 , 1 ,255 for H3K4me3, 1 ,959 for H3K36me3 and 185 for IgG), and 11 ,827 unique reads with 2,559 genes per cell for RNA (Figure 3B). FRiP varied for each histone modification target, which suggests there were different signal to background ratios depending on the modification that was profiled (Figure 3E). Different numbers of unique reads for histone modifications were detected compared to IgG, although the duplication rates were similar for all samples, which indicates that the method might directly reflect library complexity for each modality (data not shown). The IgG sample had the lowest library complexity, which also supports the specificity of our multi-target chromatin profiling (Figure 3B).

[0295] To benchmark the quality of the scMTR-seq data, the single cell profiles were computationally aggregated to generate pseudobulk data for each histone modification individually. The single cell pseudobulk data were highly concordant with bulk ENCODE ChlP-seq and CUT&Tag data sets for the same histone modifications (Figures 3C, 3D and 3F). Importantly, increasing the number of single cells included in the pseudobulk data set helped to more accurately recapitulate bulk assay profiles, although this improvement started to level off above -500 cells (Figure 3H). Thus, combining the profiles of relatively few single cells is sufficient to reproduce large-scale bulk datasets. BAB-C-P3666PCT

[0296] Sporadic, weak peaks were observed in the IgG samples that colocalised with strong peaks in the histone modification samples; a feature that was evident even in the single target, bulk CUT&Tag samples. This indicated the presence of a technical bias that could potentially contribute to some of the signal in the multi-target chromatin profiling samples (Figure 3C). It was found that computationally removing the IgG signal from each dataset could clean target profiles. The IgG profiles were also used as a background dataset for de novo peak calling using the aggregated scMTR-seq data. Although the number of peaks for each histone modification varied in different techniques, the majority peaks are overlapped for the same histone modifications across different techniques (Figure 3F).

[0297] Next, to assess whether the multi-targets chromatin profiles could be used to faithfully identify chromatin states with different histone modification combinations, ChromHMM analysis was performed. With five histone modifications and IgG control, seven chromatin states were identified, including active promoter, bivalent promoter, active enhancer, primed enhancer, transcription, repressed polycomb and unmarked or unknown states (Figure 3I). More importantly, ChromHMM analysis with scMTR-seq multi-targets chromatin profiles successfully identified more than 86% of the chromatin states that were reproducibly identified in single target CUT&Tag data (Figure 3J).

[0298] Dimensional reduction analysis was also performed for each modality. Consistent with previous single cell sequencing results (Han et al. (2018) Genome Biol., 19(1):47, doi: https: / / doi.org / Q 1186 / s 13059-018- 426-0), hPSCs are relatively homogeneous, without any clear multiple clustering pattern across all histone profiles and transcriptome modalities (Figure 3K). To use the multi-modal single cell data for dissecting the interplay between different modalities, Cramer’s V analysis was performed (Figure 3G). At the gene level, gene expression is concurrence with the presence of active histone marks, while H3K27me3 has low associations with active histone marks or gene expression, concordant with previous knowledge.

[0299] The ability of scMTR-seq to profile eight targets simultaneously was then tested. The results are shown in Figure 4. As is demonstrated, scMTR-seq achieves good accuracy and specificity of signal even when profiling 8 different targets (either histone modifications or transcription factors, such as CTCF). BAB-C-P3666PCT

[0300] Together, these results demonstrate that scMTR-seq can faithfully profile multiple modality chromatin profiles and the transcriptome with high sensitivity and high recovery, even when working with small sample sizes such as single cells.

[0301] Example 4: scMTR-seq can be Performed in Either Co-Tagmentation or Sequential- Tagmentation Steps

[0302] The ability of scMTR-seq to profile multiple targets either sequentially or simultaneously was also tested. The data is shown in Figure 5 and demonstrates that profiling six histone modifications using scMTR-seq has high specificity and accuracy in both co-tagmentation (all six targets together; T6-co) and sequential-tagmentation formats (either six rounds of one target (T6-seq6) or three rounds of multiple targets (T6-seq3)). The target signals are further improved when normalised to IgG signals.

[0303] Example 5: scMTR-seq Identifies Chromatin States and Gene Regulatory Network During Definitive Endoderm Differentiation

[0304] To further test whether the present method could dissect continuous development processes, scMTR-seq was applied to profile human definitive endoderm differentiation from human primed stem cells with four time points (H9 primed pluripotent stem cell, DayO; anterior primitive streak, Day1 ; Definitive endoderm, Day2; Definitive endoderm, Day3; Figure 6A). From 18,000 starting nuclei for each time point, 7309 cells were successfully recovered for DayO, 8844 cells for Day1 , 8576 cells for Day2 and 7098 cells for Day3, achieving high recovery rate arrange from 39.4% to 49.1% with 44.2% on average (data not shown). RNA sample information was further used to estimate the collision rate with RNA data. The collision rates are between 0.68% and 0.86%, much lower than that in scMulTi-Tag (collision rates from 9.9% to 11.0%; Figure 6E). A similar level of unique reads were detected for all four time points, with on average of 9,570 for DayO, 8,911 for Day1 , 9,592 for Day2, and 7,931 for Day3 of all DNA unique reads per cell, with on average of 11 ,966 / 2,578 for DayO, 13,399 / 2,777 for Day1 , 8,747 / 2,051 for Day2, and 9,258 / 2,066 for Day3 of unique RNA reads or genes per cell (Figure 6F). Dimensional reduction analysis with RNA data revealed that cells from different time points were clearly separated by different clusters and very few cells were clustered with different time points, suggesting a high efficiency and synchronised definitive endoderm induction (Figures 6A-6B). Lineage marker genes, such as pluripotent genes (SOX2, POLI5F1 , PRDM14, NANOG), anterior primitive streak genes (TBXT, PDGFR1 , EOMES, WNT3), Definitive endoderm genes (SOX17, GATA6, PRDM1 , and CER1), are specifically expressed in corresponding time points and RNA clusters (Figure 6C and 6G). However, dimensional reduction analysis with different histone modification modality data cannot give the same confidence to separate cells by clear time points (Figures 6A-6B). Notably, single BAB-C-P3666PCT cell data of enhancer related marks, such as H3K4me1 and H3K27ac were able to distinguish major cell types, pluripotent stem cells (Day1), anterior primitive streak (Day2) and definitive endoderm (Day2 and Day3; Figures 6A-6B). Promoter related marks (H3K4me3, H3K27me3) and a transcribed gene body marker (H3K36me3) were not able to separate cells by time points, but still showed clear pattern with slow sliding along differentiation time points (Figures 6A-6B).

[0305] Further, single cell histone modification data were aggregated from the same time point to generate pseudobulk data for each time point. The enrichment of active histone markers (H3K4me1 , H3K4me3, H3K27ac, H3K36me3) were positively associated gene expression, while the repressive marker (H3K27me3) was negatively associated gene expression (Figure 6C).

[0306] To identify enhancer regions and capture chromatin state dynamics along definitive endoderm differentiation, ChromHMM analysis was performed with time points pseudobulk data of all five histone modifications (Figure 3C-3D). Interestingly, we noticed the length of different chromatin states have clear pattern (Figure 6H). Active enhancer regions tend to be narrow, sharp regions (on average approximately 350bp), however repressive regions tend to be broader (on average 933bp). Chromatin states of most promoters maintained the same states across different differentiation time points, while enhancer states were drastic changed from time point to time point (Figure 6I). This may also explain the LIMAP analysis with enhancer related markers and promoter related markers showed different confidence to separate cells by cell types.

[0307] This data supports that shown in Example 2, and demonstrates the ability of scMTR-seq to profile novel chromatin profiles with the transcriptome in developing cells.

[0308] Example 6: scMTR-seq can Profile Six Targets Plus Transcriptome in Small Sample Sizes

[0309] Further to the results shown in Example 4, single cell dissociated E4.5 mouse embryos were profiled using scMTR-seq to test the ability to profile small sample sizes (in this example, starting with approximately 10,000 cells per replicate sample). After combining replicate samples and filtering cells with sufficient DNA and RNA unique reads, a total of 5,705 cells were retained in the final dataset, yielding an average of 27 cells per blastocyst.

[0310] The E4.5 mouse embryo is formed of three cell lineages: epiblast, trophectoderm and primitive endoderm. Data is shown in Figure 7 and demonstrates the ability of scMTR-seq to profile BAB-C-P3666PCT the histone modifications for all three cell types in the E4.5 mouse embryo. As shown in Figure 7C, some histone modifications, particularly those associated with gene enhancers, identify and separate the three embryo cell lineages.

[0311] The data further revealed lineage-specific asymmetries in gene regulation. Although the promoters of EPI-associated genes were strongly marked by active histone modifications in EPI cells, the promoters of PE- and TE-associated genes unexpectedly carried intermediate levels of active modifications in EPI cells (Figure 7E). A similar pattern was observed for PE cells (Figure 7E). In contrast, TE cells had a distinct pattern of histone modifications at promoters, whereby active histone modifications were only enriched at TE-specific promoters and were low or absent at EPI- and PE-associated gene promoters (Figure 7E). Similar epigenetic asymmetries were detected for enhancers, whereby EPI- and PE-associated enhancers were marked by H3K27ac in EPI cells, but TE-associated enhancers were not (Figure 7E). In PE cells, enhancers associated with all three lineages had high or intermediate levels of H3K27ac, and in TE cells, only TE-associated enhancers were marked by H3K27ac (Figure 7E). These results show that TE is more epigenetically diverse and separated as compared to EPI and PE.

[0312] Furthermore, using these data regulatory networks for each embryo cell lineage can be created (Figure 7D). Individual gene regulatory sub-networks included expected and also novel transcription factor-centred networks, such as Klf-4, Sox2 and Nanog for EPI; Gata4 and Nr3c2 for PE; Tead4, Tfap2c and Gata3 for TE (Figure 7D). These connections are exemplified by Gata3, which is highly expressed in TE cells, and enhancers predicted to be targeted by Gata3 (n=138 enhancers) have high activity specifically in TE cells (Figure 7F). Furthermore, genes predicted to be targets of those enhancers (n=125 genes) are highly expressed in TE cells and are positively correlated with Gata3 expression in individual cells (Figure 7F).

[0313] Example 7: Genotyping Genomic and / or Transcriptomic Sequences in scMTR-seq Libraries

[0314] Specific primers were designed to anneal to the upstream (5’) region of the target for transcriptomic RNA libraries, and to both the upstream and downstream regions for genomic DNA libraries. This design accounts for the fact that cell barcode information is located on the downstream (3’) end of RNA library fragments, but can be on either side for DNA library fragments. Preamplified cDNA library and amplified DNA library were used as template for genotyping. PCR was performed to specifically enrich fragments containing the targeted region, using the following components: 1.25 pL of 10 pM biotinylated specific primers, 1.25 pL BAB-C-P3666PCT of 10 pM primer PA-F (which anneals to the cell barcode region), 12.5 pL of KAPA HiFi HotStart mix, and 10 pL of 1 :2 diluted template. The PCR program was as follows: 98°C for 3 min; 12 cycles of 98°C for 20 s, 64°C for 15 s, and 72°C for 30 s; followed by a final extension at 72°C for 1 min. To isolate the enriched fragments, 10 pL of MyOne C1 Streptavidin Dynabeads were prepared by washing with 1x BWT buffer twice and resuspending in 25 pL of 2x BW buffer. The beads were added to the PCR products and incubated at room temperature for 20 minutes to allow binding of the specific fragments. After two washes with 1x BWT, the beads were air-dried for 2-5 minutes and resuspended in 23 pL of nuclease-free water. For library amplification, 25 pL of 2x KAPA HiFi HotStart mix, 1 pL of 25 pM indexed Ad1 , and 1 pL of 25 pM indexed P7 primers were added. The amplification was performed using the following program: 98°C for 3 min; 12 cycles of 98°C for 20 s, 64°C for 15 s, and 72°C for 30 s; with a final extension at 72°C for 1 min. PCR products were purified using 0.7x AMPure XP beads and eluted in 32 pL of EB buffer. Library concentrations were quantified using a Qubit fluorometer, and fragment size distributions were assessed on a Bioanalyzer. Libraries were sequenced on an Illumina NovaSeq X platform.

[0315] As shown in Figure 8, a primer sequence complementary to a region upstream of a target sequence within an exemplary gene (JAK2) can be used to amplify transcripts of said gene in a transcriptome library generated by the present method to successfully identify the point mutation causing the amino acid mutation V617F in the JAK2 protein (Figure 8A; JAK2 WT is represented by SEQ ID NO: 1 and JAK2 V617F is represented by SEQ ID NO: 2). In particular, the rate of V617F, a common mutation in myeloproliferative neoplasm (MPN) cancers of the bone marrow (in particular MPNs such as polycythemia vera (PV), essential thrombocythemia (ET) and myelofibrosis (MF)) can be robustly determined in both bone marrow cells and peripheral blood obtained from MPN subjects (Figure 8B).

Claims

BAB-C-P3666PCTCLAIMS1 . A method of identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said method comprising the steps of:(i) pre-assembling individual anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence;(ii) blocking the assembled antibody-transposase-adapter complexes;(iii) performing in situ single step fragmentation and adapter sequence insertion in the sample using the assembled antibody-transposase-adapter complexes to generate an adapter-tagged library of sequences associated with the chromatin mark, genome binding protein and / or genomic modification; and(iv) sequencing the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

2. The method of claim 1 , further comprising the steps of:(iii b) capturing mRNA transcripts in the sample;(iii c) performing in situ reverse transcription of the captured mRNA transcripts to generate a transcriptome library; and(iv b) sequencing the transcriptome library together with the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

3. The method of claim 1 of claim 2, wherein the sample comprises one or more cells, such as a single cell, in particular wherein the sample is one or more cells, such as a single cell, and / or wherein the genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications are identified simultaneously in a single sample, such as in a single population of cells or in a single cell, and optionally wherein the transcriptome of the cell is identified simultaneously in the single sample, such as in the single population of cells or in the single cell.

4. The method of any one of claims 1 to 3, wherein two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i), each complex differentially comprising one or more antibodies recognising a single chromatin mark, genome binding protein orBAB-C-P3666PCT genomic modification of interest and a single indexed transposase-adapter sequence, and the two or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are used together simultaneously in step (iii), optionally wherein three or more anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled in step (i) and used together simultaneously in step (iii), in particular four or more, five or more, six or more, seven or more or eight anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes are individually pre-assembled.

5. The method of claim 4, wherein each anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complex comprises one or more antibodies recognising different chromatin marks, genome binding proteins or genomic modifications of interest and different indexed transposase-adapter sequences, such that each indexed transposase-adapter sequence is unique to each chromatin mark, genome binding protein and genomic modification of interest.

6. The method of claim 4 or claim 5, wherein following step (iii), one or more further antichromatin mark, anti-genome binding protein and / or anti-genomic modification antibody- transposase-adapter complexes are used sequentially or together simultaneously for further in situ single step fragmentation and adapter sequence insertion in the sample.

7. The method of any one of claims 1 to 6, wherein the chromatin marks are histone modifications, such as phosphorylation, acetylation, ubiquitination, GIcNAcylation, citrullination, crotonylation, propionylation, sumoylation, isomerisation and / or methylation, including mono-methylation, di-methylation and tri-methylation, and / or optionally wherein the anti-chromatin mark antibody recognises and binds a histone modification and / or a modified histone selected from any of: H3K4me1 , H3K4me2, H3K4me3, H3K36me3, H3K79me2, H3K9Ac, H3K27AC, H4K16AC, H3K27me1 , H3K27me2, H3K27me3, H2AK119ub, H3K122ac, H3K9me2,H3K9me3, H3R2me3, H3R8me3, H3R17me3, H3R26me3, H3R42me3, H3S10P, H3S28P, H4K5me3, H4K8me3, H4K12me3, H4K16me3, H4K20me3, H4R3me3 and gamma H2A.X.

8. The method of claim 7, wherein the histone modification is associated with euchromatin or heterochromatin, in particular euchromatin.BAB-C-P3666PCT9. The method of any one of claims 1 to 8, wherein the indexed transposase-adapter sequence comprises protein A, protein G, protein A / G or protein L, with a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence, in particular wherein the indexed transposase-adapter sequence comprises protein A, a transposases enzyme and a transposase adapter sequence, and optionally wherein blocking step (ii) comprises blocking free protein A, protein G, protein A / G or protein L in the assembled antibody-transposase-adapter complexes, such as using free IgG antibodies.

10. The method of any one of claims 1 to 8, wherein the indexed transposase-adapter sequence comprises streptavidin, a transposase enzyme and a transposase adapter sequence comprising a unique identifying indexing sequence, and wherein the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies are biotinylated.11 . The method of claim 10, wherein blocking step (ii) comprises blocking free streptavidin in the assembled antibody-transposase-adapter complexes, such as using free biotin.

12. The method of any one of claims 1 to 8, wherein the indexed transposase-adapter sequence is directly conjugated to the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibodies, and optionally wherein blocking step (ii) is omitted.

13. The method of any one of claims 1 to 12, wherein in situ single step fragmentation and adapter sequence insertion in step (iii) comprises tagmentation, and / or the transposase is Tn5 transposase, optionally a mutant Tn5 transposase such as hyperactive Tn5 transposase.

14. The method of any one of claims 1 to 13, wherein blocking step (ii) reduces off-target binding of the anti-chromatin mark, anti-genome binding protein and / or anti-genomic modification antibody-transposase-adapter complexes, in particular wherein two or more antibody-transposase-adapter complexes are used together simultaneously in step (iii) as defined in claim 4 or claim 5 or are used in further in situ single step fragmentation and adapter sequence insertion in the sample as defined in claim 6.

15. The method of any one of claims 1 to 14, wherein the indexed transposase-adapter sequence comprises only Mosaic End B (MEB) adapter sequences, and the methodBAB-C-P3666PCT additionally comprises adding Mosaic End A (MEA) adapter sequences to the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library following step (iii), optionally wherein the indexed transposase-adapter sequence comprising only Mosaic End B (MEB) adapter sequences and the adding of Mosaic End A (MEA) adapter sequences results in the sequencing of all of the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library in step (iv) and optionally step (iv b) and / or prevents the generation of chimeric sequence reads in the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library.

16. The method of any one of claims 2 to 15, wherein the transcriptome comprises mRNA transcripts of actively transcribed regions of the genome of the cell, such as of actively expressed genes, in particular actively transcribed regions of the genome associated with euchromatin.

17. The method of any one of claims 2 to 16, wherein capturing mRNA transcripts in step (iii b) is performed using a probe or bait sequence, such as a poly-T sequence and / or sequences complementary to mRNA transcripts of interest.

18. The method of any one of claims 1 to 17, wherein step (iv) is performed after the single step fragmentation and adapter sequence ligation of step (iii), and / or optionally wherein step (iv b) is performed after step (iii c).

19. The method of any one of claims 1 to 18, wherein step (iii) and optionally steps (iii b) and (iii c) are performed on permeabilised cells and / or intact cell nuclei, optionally wherein the cells and / or cell nuclei are cross-linked, and optionally the method additionally comprises lysing the cells and / or nuclei and reversing cross-linking following step (iii) and optionally step (iii c).

20. The method of any one of claims 1 to 19, wherein the method additionally comprises splitting the sample following step (iii) and optionally step (iii c) into distinct samples and adding unique molecular identifiers to each distinct sample, followed by re-pooling the uniquely identified distinct samples, optionally wherein the splitting and re-pooling is repeated one or more times, in particular repeated two times.BAB-C-P3666PCT21 . The method of any one of claims 1 to 20, wherein the method additionally comprises amplifying the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and the transcriptome library prior to step (iv) and optionally step (iv b), optionally wherein said amplification is by PCR.

22. The method of any one of claims 1 to 21 , wherein the method additionally comprises amplifying and / or selecting a target sequence of interest, optionally a target sequence within the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification- associated sequence library and / or the transcriptome library, prior to step (iv) and optionally step (iv b), wherein said amplifying and / or selecting comprises one or more primer sequences complementary to a region surrounding the target sequence of interest, optionally wherein said amplification is by PCR, and / or optionally wherein the primer sequences further comprise a selection marker, such as biotin.

23. A method of determining the sequence of and / or genotyping a target sequence in a sample, said method comprising performing the method of any one of claims 1 to 21 and further amplifying or selecting genomic nucleic acid regions or transcriptome library sequences comprising the target sequence prior to step (iv) and optionally step (iv b).

24. A method of identifying chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with a particular disease state and / or a particular state of developmental differentiation, said method comprising:(i) performing the method of any one of claims 1 to 22 on a sample of one or more cells obtained from an individual suspected of having or with a particular disease or from a tissue or organ suspected of being or being diseased, or on a sample of one or more cells from a developing tissue or organ or from a differentiation culture;(ii) identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with actively transcribed regions of the genome, such as actively expressed genes, wherein said actively transcribed regions are identified by sequencing the transcriptome library and wherein said actively transcribed regions are associated with the particular disease state and / or state of developmental differentiation, or identifying the chromatin marked, genome binding protein bound and / or modified nucleic acid regions associated with transcriptionally repressed regions of the genome, wherein said transcriptionally repressed regions are identified by sequencing theBAB-C-P3666PCT transcriptome library and wherein said transcriptionally repressed regions are associated with the particular disease state and / or state of developmental differentiation; and(iii) determining a chromatin marked, genome binding protein bound and / or modified genomic nucleic acid region as being associated with a particular disease state and / or a particular state of developmental differentiation when it associates with actively transcribed or transcriptionally repressed regions of the genome associated with the particular disease state and / or state of developmental differentiation.

25. The method of claim 24, additionally comprising quantifying the frequency of association of the chromatin marked, genome binding protein bound and / or modified genomic nucleic acid regions associated with actively transcribed or transcriptionally repressed regions of the genome in step (ii), and comparing the frequency of association in the sample with the frequency of associations in a sample obtained from a non-diseased individual or a nondiseased tissue or organ, or in a sample of one or more cells at a different state of developmental differentiation from a developing tissue or organ or from a differentiation culture in step (iii).

26. The method of claim 24 or claim 25, wherein the method additionally comprises amplifying and / or selecting a target sequence of interest, optionally within the adapter-tagged chromatin mark-, genome binding protein- and / or genomic modification-associated sequence library and / or the transcriptome library, prior to step (iv) and optionally step (iv b) as defined in claim 22 or claim 23.

27. The method of any one of claims 24 to 26, wherein the disease is cancer, an autoimmune disease, a developmental disease and / or a genetic disorder, such as a cancer associated with epigenetic modifications, or wherein the developmental differentiation is stem cell differentiation, embryogenesis and / or embryo model development.

28. A kit for identifying genomic nucleic acid regions associated with chromatin marks, genome binding proteins and / or genomic modifications, and the transcriptome in a sample, said kit comprising individual pre-assembled and blocked antibody-transposase-adapter complexes, each of said complexes comprising one or more antibodies specific for a chromatin mark, genome binding protein and / or genomic modification of interest with an indexed transposase-adapter sequence, in addition to buffers and reagents capable of performing the method of any one of claims 1 to 22,BAB-C-P3666PCT optionally wherein the antibody-transposase-adapter complexes are as defined in any one of claims 4 to 15.