Methods to profile protein binding events on DNA with single-cell resolution
A fusion protein with a base-editing enzyme and immunoglobulin-binding protein allows single-cell profiling of DNA-binding proteins, addressing the limitations of existing methods and revealing cellular regulation and disease mechanisms.
Patent Information
- Application Number
- PCT/US2025/037707
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Current methods for profiling protein-DNA interactions, such as ChIP-seq and CUT&Tag, are not easily adaptable to single-cell workflows, limiting the ability to understand regulatory control by chromatin-binding factors and epigenetic dysregulation across diverse cell types and states, which is crucial for understanding gene expression and disease mechanisms.
A fusion protein comprising a base-editing enzyme and a protein that binds to the Fc region of an immunoglobulin, such as a nanobody, is used to identify DNA-binding protein interactions by inducing base changes in the vicinity of the target genomic area, enabling single-cell profiling through sequencing.
Enables high-throughput, single-cell profiling of DNA-binding proteins, revealing their binding events and chromatin structures, thereby providing insights into cellular regulation and disease-associated dysregulation.
Smart Images

Figure IMGF000100_0001 
Figure IMGF000103_0001 
Figure IMGF000103_0002
Abstract
Description
[0001] Methods to Profile Protein Binding Events on DNA with Single-Cell Resolution
[0002] Related Application
[0003] This application claims the benefit of U.S. Provisional Patent Application serial number 63 / 671,380, filed on July 15, 2024, which is hereby incorporated by reference herein in its entirety.
[0004] Government Support
[0005] This invention was made with government support under RO 1 HL 157387-01 Al awarded by the National Heart Lung and Blood Institute, R33 CA267219 awarded by National Cancer Institute, UG3NS132139-01 awarded by the National Institutes of Health Common Fund Somatic Mosaicism Across Human Tissues, and RM1HG011014 awarded by the National Human Genome Research Institute, Center of Excellence in Genomic Science. The government has certain rights in the invention.
[0006] Background of the Invention
[0007] The genomic DNA is responsible for storing genetic information, which is maintained and decoded by proteins that read, regulate, replicate, recombine, and repair it. Understanding where and how proteins interact with DNA can help provide valuable insights into how they function, as well as malfunction, in healthy and diseased cells. To map the interaction of individual target proteins with DNA genome-wide, several powerful approaches have been developed, including DamID, chromatin immunoprecipitation with sequencing (ChlP-seq), cleavage under targets & tagmentation (CUT&TAG), and their derivatives. These approaches involve selectively amplifying short DNA fragments from regions bound by a particular protein of interest, determining the sequence of those DNA molecules using next-generation sequencing (NGS), and mapping those sequences back to a reference genome. Although these methods are extremely helpful for studying DNA- binding proteins and chromatin modifications, they suffer from several limitations. In particular, none of these methods can be easily implemented in existing single-cell workflows.
[0008] Regulatory control orchestrated by chromatin-binding factors, including transcription factors (TFs) and chromatin remodelers, underlies the gene expression programs responsible for maintaining cell identity, executing cellular functions and responding to environmental stimuli. These DNA:protein interactions are directed through epigenetic features, such as histone modifications and DNA methylation (DNAme), which establish chromatin landscapes, modulating the binding of specific factors, and thereby sculpting the functional genome according to the needs of the cell. Importantly, epigenetic dysregulation has been linked to cellular dysfunction in disease, cancer and aging, where aberrant chromatin landscapes transform the binding landscapes of TFs, altering the normal biological processes of the cell. Furthermore, >90% of GW AS hits fall within non-coding regions, underscoring the need to define cis-regulatory elements for insights into human genetic disease. As such, understanding the complexity and dysregulation of the TF binding code across different cell types and cell states remains a central challenge in human biology.
[0009] While transcriptional readout through RNAseq identifies the active genes and pathways that contribute to cellular responses or phenotypes, understanding the regulatory factors that direct these expression programs is crucial for both biological insights as well as for the identification of factors that are dysregulated during disease. There are over 1,600 likely TFs encoded in the human genome, with DNA binding motif sequence information, obtained from various in vitro, bulk or bioinformatic analyses, established for -1,100 of these. While the presence of a canonical motif is a key determinant of TF binding, motif information alone is often insufficient to accurately predict binding in vivo, as TF occupancy is influenced by the local chromatin context, TF concentration, and cofactor availability45. Importantly, chromatin accessibility itself is both a consequence and a determinant of TF binding: many TFs preferentially bind accessible regions5, while others, particularly pioneer factors, actively shape chromatin states through their binding activity. TF-TF interactions6, TF-nucleosome competition7, and the spatial syntax of motifs within enhancers also influence binding specificity.
[0010] Thus, profiling TF binding in native chromatin contexts is crucial for unraveling the physiologic regulation, and disease-associated dysregulation, of gene expression. However, profiling of these networks across diverse cell types and states is lacking, due to technical limitations of current methods for mapping DNA:Protein interactions in single cells. Chromatin immunoprecipitation with sequencing (ChlP-seq) and cleavage under targets & tagmentation (CUT&Tag) are powerful approaches for mapping genome-wide interactions of target proteins with DNA. These methods identify genomic regions bound by an individual protein of interest through sequencing short DNA fragments that are cross linked or tagmented. In addition, DamlD8,9molecular footprinting techniques harnessing DNA methyltransferase activity have been used to map TF binding. However, due to their various protocol requirements, these methods cannot be easily incorporated into available high- throughput single-cell workflows, limiting application to bulk analysis, or in single cells to the profiling of only the chromatin factors that display the strongest interactions with DNA, such as histones.
[0011] Summary of the Invention
[0012] In some aspects, provided herein is a fusion protein comprising a base-editing enzyme and a protein that binds to a fragment crystallizable region (Fc region) of an immunoglobulin .
[0013] In some embodiments, the Fc region is an Fc region from IgGl, IgG2, IgG3, or IgG4. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an antibody or antigen-binding fragment thereof, a protein G, and / or a protein A.
[0014] In some embodiments, the protein that binds to a Fc region of an antibody is or comprises an antibody or antigen-binding fragment thereof. In some embodiments, antigenbinding fragment thereof is selected from a Fab, a single chain variable region (scFv), a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody. In some embodiments, the antigen-binding fragment thereof is a nanobody. In some embodiments, the nanobody recognizes a rabbit, mouse, goat or rat immunoglobulin.
[0015] In some aspects, provided herein is a fusion protein comprising a base-editing enzyme and a protein that binds to a DNA binding protein. In some embodiments, the protein that binds to a DNA binding protein is an antibody or antigen-binding fragment thereof.
[0016] In some embodiments, the antigen-binding fragment thereof is selected from a Fab, a single chain variable region (scFv), a nanobody, a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody. In some embodiments, the antigen-binding fragment thereof is a nanobody. In some embodiments, the nanobody comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 1-4. In some embodiments, the nanobody comprises the amino acid sequence of any one of SEQ ID NOs: 1-4.
[0017] In some embodiments, the protein that binds to a Fc region of an immunoglobulin is or comprises a protein G. For example, in some embodiments, the protein is a protein G. In some embodiments, the protein is a fusion protein comprising a protein G, e.g., a protein AG. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to an amino acid sequence of SEQ ID NO: 16 or 20. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence of SEQ ID NO: 16 or 20.
[0018] In some embodiments, the protein that binds to a Fc region of an immunoglobulin is or comprises a protein A, e.g., a protein AG. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to an amino acid sequence of SEQ ID NO: 5 or 20. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence of SEQ ID NO: 5 or 20.
[0019] In some embodiments, the base-editing enzyme is a base-editing deaminase. In some embodiments, the base-editing deaminase deaminates cytosine to uracil (C>U). In some embodiments, the base-editing enzyme is selected from DddA, CbDaOl, AcDaOl, non-DddA-like OU dsDNA deaminases, or DddA-like. In some embodiments, the DddA base-editing enzyme is DddAl 1. In some embodiments, the DddA-like base-editing enzyme is MGYPDa829 or LbsDaOl. In some embodiments, the base-editing enzyme comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 6-8. In some embodiments, the base-editing enzyme comprises an amino acid sequence selected from SEQ ID NOs: 6-8. In some embodiments, the base-editing enzyme is a non-split full-length base editor.
[0020] In some embodiments, the base-editing enzyme is a split base editor. In some embodiments, the base-editing enzyme is a DddA split base editor. In some embodiments, the DddA split base editor comprises an N-terminal portion of DddA (DddA_NT). In some embodiments, the N-terminal portion of DddA (DddA_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 9. In some embodiments, the N-terminal portion of DddA (DddA_NT) comprises the amino acid sequence of SEQ ID NO: 9. In some embodiments, the DddA split base editor is activated by the addition of a C-terminal portion of DddA (DddA_CT). In some embodiments, the C- terminal peptide comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 10 or 17. In some embodiments, the C-terminal peptide comprises the amino acid sequence of SEQ ID NO: 10 or 17.
[0021] In some embodiments, the base-editing enzyme is a MGYPDa829 split base editor. In some embodiments, the MGYPDa829 split base editor comprises an N-terminal portion of MGYPDa829 (M829_NT). In some embodiments, the N-terminal portion of MGYPDa829 (M829_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 11. In some embodiments, the N-terminal portion of MGYPDa829 (M829_NT) comprises the amino acid sequence of SEQ ID NO: 11. In some embodiments, the MGYPDa829 split base editor is activated by the addition of a C-terminal portion of MGYPDa829 (M829_CT). In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence at least 95 % homology to the amino acid sequence of SEQ ID NO: 12 or 19. In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises the amino acid sequence of SEQ ID NO: 12 or 19.
[0022] In some embodiments, the base-editing deaminase deaminates A to I (A>I). In some embodiments, the A>I deaminase is hADA. In some embodiments, the A>I deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. In some embodiments, the adenine base editor is fused to the N-terminus of the DddA variant. In some embodiments, the adenine base editor is TadA, optionally wherein the adenine base editor is TadA8e. In some embodiments, the DddA variant is DddA E1347A. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13 or 22. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13 or 22. In some embodiments, the AM deaminase is a non-split full length AM deaminase.
[0023] In some embodiments, the AM deaminase is a split AM deaminase. In some embodiments, the AM deaminase split base editor comprises TadA8e fused to an N- terminal portion of DddA E1347A. In some embodiments, the N-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14. In some embodiments, the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14. In some embodiments, the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A. In some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18. In some embodiments, the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18.
[0024] In some embodiments, the fusion protein is capable of binding to a DNA binding protein. In some embodiments, the fusion protein is capable of binding to a DNA binding protein- specific antibody. In some embodiments, the fusion protein is capable of catalyzing cytosine (C) or adenosine (A) deamination in the vicinity of the DNA binding protein- targeted genomic area. In some embodiments, the DNA-binding protein binds to a DNA directly. In some embodiments, the DNA-binding protein binds to a DNA indirectly. In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
[0025] In some aspects, provided herein is a nucleic acid encoding the fusion protein described herein.
[0026] In some aspects, provided herein is a vector comprising the nucleic acid described herein.
[0027] In some aspects, provided herein is a kit comprising the fusion protein described herein. In some embodiments, the kit further comprises a primary antibody that binds the DNA-binding protein. In some embodiments, the kit further comprises the C-terminal portion of DddA (DddA_CT) or the C-terminal portion of MGYPDa829 (MGYPDa829_CT). In some embodiments, the kit further comprises a divalent cofactor. In some embodiments, the divalent cofactor is Zn2+.
[0028] In some aspects, provided herein is a method of identifying a binding event of a DNA- binding protein in a cell or a population of cells, comprising: (a) contacting the cell or the population of cells with a fusion protein described herein; (b) incubating the cell or the population of cells at a condition for a period of time; and (c) detecting base changes induced by the fusion protein by sequencing, thereby identifying the binding event.
[0029] In some embodiments, the DNA-binding protein binds to a DNA directly. In some embodiments, the DNA-binding protein binds to a DNA indirectly. In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the method profiles all binding events of the DNA-binding protein in the cell or the population of cells. In some embodiments, the population of cells comprises at least 100 cells, at least 150 cells, at least 200 cells, or at least 250 cells. In some embodiments, the cell is a primary cell, optionally wherein the cell is a primary human cell, e.g., primary human PBMCs. In some embodiments, the cell is a cell line.
[0030] In some embodiments, the cell or the population of cells are contacted with an antibody specific to the DNA-binding protein prior to step (a). In some embodiments, the cell or the population of cells are fixed and permeabilized prior to step (a).
[0031] In some embodiments, step (b) occurs in the presence of the C-terminal portion of DddA (DddA_CT). In some embodiments, step (b) occurs in the presence of a divalent cofactor. In some embodiments, step (b) occurs in the presence of Zn2+. In some embodiments, step (b) occurs at a temperature from 37°C to 50 °C.
[0032] In some embodiments, the sequencing at step (c) is next generation sequencing. In some embodiments, the sequencing at step (c) is single cell sequencing. In some embodiments, the sequencing at step (c) is tagmentation-based single cell sequencing. In some embodiments, the sequencing at step (c) is Assay for Transposase- Accessible Chromatin using sequencing (ATAC-seq), single-cell (sc) ATAC-seq, Nanobody Tethered Tagmentation followed by sequencing (NTT-seq), CUT&Tag-related seq, or Dogma-seq. In some embodiments, the sequencing at step (c) comprises Tn5 or TnY tagmentation. In some embodiments, the sequencing at step (c) uses a lOx Genomics single-cell (sc) ATAC- seq kit. In some embodiments, the sequencing at step (c) is a droplet-based single cell sequencing method. In some embodiments, the sequencing at step (c) is a split and pool single-cell sequencing method. In some embodiments, the sequencing at step (c) is a spatially resolved chromatin profiling method. In some embodiments, the sequencing at step (c) is Slide-seq or dBIT-seq. In some embodiments, the sequencing at step (c) is whole-genome sequencing. In some embodiments, the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing. In some embodiments, the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PTA) protocol. In some embodiments, the method can be used to predict 3D chromatin structure in the cell or the population of cells.
[0033] In some aspects, provided herein is a method of identifying a binding event of a DNA-binding protein in a cell or a population of cells, comprising: (a)fixing and permeabilizing the cell or the population of cells; (b) contacting the cell or the population of cells with a primary antibody that specifically binds the DNA-binding protein; (c) contacting the cell or the population of cells with a fusion protein comprising (1) a protein that binds the Fc region of the primary antibody, optionally wherein the protein is a secondary nanobody, a protein A, or a protein G; and (2) an N-terminal portion of a baseediting enzyme, optionally wherein the base-editing enzyme is selection from DddAl 1, MGYPDa829, LbsDaOl, or TadA8e-DddA El 347 A; (d) adding a C-terminal portion of the base-editing enzyme and a divalent cofactor; (e) performing Tagmentation, optionally wherein the Tagmentation comprises Tn5 or TnY tagmentation; (f) extracting genomic DNA from the cell and constructing a library for sequence; and (g) sequencing the genomic DNA library to detect base changes induced by the fusion protein, thereby identifying the binding event.
[0034] In some embodiments, the method further comprises conducting a genotyping method. In some embodiments, the genotyping method is a high-throughput single-cell genotyping method. In some embodiments, the genotyping method is Genotyping of Transcriptomes (GoT) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA). In some embodiments, the method further comprises conducting cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), NTT-seq, ATAC with select antigen profiling by sequencing (ASAP-seq), simultaneous high-throughput ATAC and RNA expression with sequencing (SHARE-seq), Paired-Tag, and / or combined assay of transcriptome and enriched chromatin binding (CoTECH).
[0035] In some aspects, provided herein is a system or kit comprising: (a) a first fusion protein comprising a first nanobody and a first deaminase; and (b) a second fusion protein comprising a second nanobody and a second deaminase, wherein (1) the first deaminase and the second deaminase have different motif preferences; and (2) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species. In some embodiments, the first deaminase prefers CpG context, and the second deaminase prefers TpC context. In some embodiments, the first deaminase is CbDaOl or AcDaOl; and the second deaminase is DddA or DddA-like.
[0036] In some aspects, provided herein is a system or a kit comprising: a first fusion protein comprising a first nanobody and a A>I deaminase; and a second fusion protein comprising a second nanobody and a C>U deaminase, wherein the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
[0037] In some embodiments, the AM deaminase is hADA. In some embodiments, the AM deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. In some embodiments, the adenine base editor is fused to the N-terminal of the DddA variant. In some embodiments, the adenine base editor is TadA8e. In some embodiments, the DddA variant is DddA E1347A. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13 or 22. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13 or 22. In some embodiments, the A>I deaminase is a non-split full-length AM deaminase.
[0038] In some embodiments, the AM deaminase is a split AM deaminase. In some embodiments, the AM deaminase split base editor comprises TadA8e fused to an N- terminal portion of DddA E1347A. In some embodiments, the N-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14. In some embodiments, the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14. In some embodiments, the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A. In some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18. In some embodiments, the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18.
[0039] In some embodiments, the system or kit further comprises primary antibodies to the two different DNA-binding proteins. In some embodiments, the system or kit further comprises the C-terminal portion of DddA E1347A.
[0040] In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the two DNA-binding proteins are cofactors, or different components from a DNA-binding protein complex.
[0041] In some aspects, provided herein is a method of identifying binding events of two DNA- binding proteins in a cell or a population of cells, comprising: (a) contacting the cell or the population of cells with the system or kit described herein; (b) incubate the cell or the population of cells at a condition for a period of time; and (c) detecting base changes induced by the system or kit by sequencing, thereby identifying the binding events of two DNA-binding proteins.
[0042] In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the method profiles all binding events of the two DNA-binding proteins in the cell or the population of cells. In some embodiments, the cell is a primary cell or a cell line.
[0043] In some embodiments, the cell or the population of cells are contacted with two primary antibodies specific to the two DNA-binding proteins, respectively, prior to step (a). In some embodiments, the cell or the population of cells are fixed and permeabilized prior to step (a). In some embodiments, the sequencing at step (c) is single cell sequencing. In some embodiments, the sequencing at step (c) is tagmentation-based single cell sequencing. In some embodiments, the sequencing at step (c) is ATAC-seq, NTT-seq, CUT&Tag-related seq, or Dogma-seq. In some embodiments, the sequencing at step (c) comprises Tn5 or TnY tagmentation. In some embodiments, the sequencing at step (c) uses a lOx Genomics scATAC-seq kit. In some embodiments, the sequencing at step (c) is a droplet-based single cell sequencing method. In some embodiments, the sequencing at step (c) is a split and pool single-cell method. In some embodiments, the sequencing at step (c) is a spatially resolved chromatin profiling method. In some embodiments, the sequencing at step (c) is Slide-seq or dBIT-seq. In some embodiments, the sequencing at step (c) is whole genome sequencing. In some embodiments, the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing. In some embodiments, the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PTA) protocol.
[0044] In some aspects, provided herein is a method of identifying binding events of two DNA-binding proteins in a cell or a population of cells, comprising: (a) fixing and permeabilizing the cell or the population of cells; (b) contacting the cell or the population of cells with two primary antibodies specific to the two DNA-binding proteins, respectively; (c) contacting the cell or the population of cells with (1) a first fusion protein comprising a first nanobody and a first deaminase; and (2) a second fusion protein comprising a second nanobody and a second deaminase, wherein (i) the first deaminase and the second deaminase have different motif preferences or the first deaminase is a A>I deaminase and the second deaminase is a C>U deaminase; and (ii) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species; and (d) sequencing genomic DNA of the cell or the population of cells to detect base changes induced by the first and the second fusion proteins, respectively, thereby identifying the binding events of the two DNA-binding proteins.
[0045] In some embodiments, the first deaminase and / or the second deaminase is a split enzyme, and the step (c) further comprises contacting the cell with a C-terminal portion of the first deaminase to activate the first deaminase and / or with a C-terminal portion of the second deaminase to activate the second deaminase. Brief Description of the Drawings
[0046] Fig. 1. a. Schematic representation of scD&D-seq technology, b. Illustration of the split DddA base editor fused to nanobodies that is activated by addition of a C-terminal peptide, c. In vitro deamination assay for DddA 11 on different genetic context, d. In vitro deamination assay for split DddAl 1. e. Motif enrichment analysis of bulk D&D-seq for CTCF on K562 cells, f. Molecular footprint of bulk D&D-seq for CTCF on K562 cells, g. Motif enrichment analysis of bulk D&D-seq for GATA1 on K562 cells, h. Molecular footprint of bulk D&D-seq for GATA1 on K562 cells, i. Motif enrichment analysis of bulk D&D-seq for GATA2 on K562 cells, j. Molecular footprint of bulk D&D-seq for GATA2 on K562 cells, k, 1. Genome browser tracks for a representative region of the human genome. D&D-seq was performed on K562 cells for CTCF, GATA1 or GATA2. Encode ChlP-seq data for the 3 proteins as reference. Sequencing data were normalized as bins per million (BPM) mapped reads.
[0047] Fig. 2. a. UMAP of cell line mixing study of CA46 and K562 cells in the AT AC space, b. Violin plot showing the number of ATAC fragments per cell for each subpopulation from a. c. Accessibility scores for marker genes defining the subpopulation lineage (FDR < 0.05). d. UMAP projection of gene accessibility score for CA46 marker genes RUBCNL and KLF12. e. UMAP projection of gene accessibility score for K562 marker genes NUP214 and PRAME. f. Projection of CTCF D&D edit events into the UMAP space, g. Violin plot of CTCF D&D edit events, h. Projection of GATA1 D&D edit events into the UMAP space, i. Same as G for GATA1. j. Motif enrichment analysis for D&D-seq reads in CA46 cells, k. Distribution of C to U D&D events in CA46 cells around CTCF binding site (top) and ATAC (bottom) in all the genome. 1. Motif enrichment analysis for D&D-seq reads in K562 cells, m. Distribution of C to U D&D events in K562 cells around GATA1 binding site (top) and ATAC (bottom) in all the genome, n, o. Pseudobulk genome browser tracks for a representative region of the human genome. scD&D-seq was performed on a mixture of K562 and CA46 cells for GATA1 or CTCF, respectively. Encode ChlP-seq data for the 2 proteins (black) as reference. Sequencing data were normalized as bins per million (BPM) mapped reads, p. Pearson correlation matrix between pseudobulk scD&D-seq signal and Encode reference ChlP-seq signal for CTCF and GATA1.
[0048] Fig. 3. Evaluation of DnD-mediated genome edits.
[0049] Fig. 4. Defining gene regulation across molecular layers in single cells. Fig- 5. Illustration of the split DddA base editor fused to nanobodies (nb) that is activated by addition of a C-terminal peptide.
[0050] Fig. 6: a. Schematic representation of scD&D-seq technology, b. UMAP of cell line mixing study in the ATAC space, c. Violin plot showing the number of fragments per cell for each subpopulation, d. Accessibility scores for marker genes defining the subpopulation lineage (FDR < 0.05). e. Genome browser tracks for a representative region of the human genome. Pseudobulk accessible regions for each cell line (gray). Pseudobulk clusterspecific D&D-seq track for CTCF and GATA1. ChlP-seq tracks from ENCODE for CTCF and GATA1 (black) were used as reference. Sequencing data were normalized as bins per million (BPM) mapped reads, f. Motif enrichment analysis for D&D-seq reads in CA46 cells, g. Distribution of C to U D&D events in CA46 cells around CTCF binding site (top) and ATAC (bottom) in all the genome, h. Same as F for K562 cells, i. Distribution of C to U D&D events in K562 cells around GATA1 binding site (top) and ATAC (bottom) in all the genome, j. Projection of CTCF D&D edit events into the UMAP space, k. Violin plot of CTCF D&D edit events. 1. Same as J for GATA1 m. Same as K for GATA1.
[0051] Fig. 7: a. UMAP of human primary PBMCs in the ATAC space, b. Projection of CTCF D&D edit events onto the UMAP space, c. Motif enrichment analysis for D&D-seq reads.
[0052] Fig. 8: Schematic for multiplexed D&D-seq.
[0053] Fig. 9 (adapted from36). a. UMAP constructed using a weighted combination of H3K27me3 and H3K27ac measurement from single-cell NTT-seq applied to bone marrow mononuclear cells, showing pseudotime trajectory for B cell development. Cells are colored by their pseudotime value and labeled by their annotated cell type. b. Heatmap showing H3K27me3 and H3K27ac signal for 10 kb genomic bins correlated with B cell pseudotime progression.
[0054] Fig. 10: a. Schematic of GoT-ChA-D&D. b. Genotyping of PBMCs from an IDH2- mutant CHIP patient. Left, UMAP based on chromatin accessibility annotated with cell types. Right, 1DH2 genotypes projected on the UMAP. MUT, IDH2R140Q, HET, heterozygous; WT, wild type; NA, no genotyping information, c. CTCF binding is disrupted in cells with IDH2 mutation.
[0055] Fig. 11. No existing tool that profiles TFs in single cells
[0056] Fig. 12. DddA: the first known dsDNA deaminase
[0057] Fig. 13. nb-DddA split version Fig. 14. Reconstituted DddAl 1 deaminates dsDNA in vitro
[0058] Fig. 15. DddAl 1 prefers to deaminate TpC context
[0059] Fig. 16. Bulk Dnd-seq Workflow
[0060] Fig. 17. Most mutations are C>T and G>A
[0061] Fig. 18. Bulk DnD-seq have high on-target activity
[0062] Fig. 19. Bulk DnD-seq have high on-target activity
[0063] Fig. 20. TF motifs enriched in peaks with C>T mut
[0064] Fig. 21. scDnD-seq cell mixing experiment
[0065] Fig. 22. High specificity and expected footprint
[0066] Fig. 23. Expected footprint recovered from on-target edit events and low background
[0067] Fig. 24. High specificity, enriched motif, DNA footprint
[0068] Fig. 25. Bulk DnD-seq recapitulates ChlP-seq signal
[0069] Fig. 26. scDnD-seq cell mixing experiment
[0070] Fig. 27. scDnD-seq shows high specificity and DNA footprint
[0071] Fig. 28. Additional optimizations of DnD-seq (to include in the patent but not in preprint)
[0072] Fig. 29. New DnD-seq enzme candidates
[0073] Fig. 30. High sequence homology with DddAl 1
[0074] Fig. 31. Dnd-seq2 with stronger DddAs
[0075] Fig. 32. M829 & LbsOl are stronger DddAs
[0076] Fig. 33. M829 split version worked
[0077] Fig. 34. Multiplexed DnD-seq using two deaminases coupled to two different nb
[0078] Fig. 35. Multiplexed DnD-seq using AM fusion deaminase
[0079] Fig. 36a.-Fig. 36c. TadA8e-DddA(E1347A): dimeric monomeric (adapted from Cho et al. Cell 2022; Cho et al. Cell 2024).
[0080] Fig. 37a-Fig. 37 e. FigureS 1 adapted from Cho et al cell 2022; Cho et al cells 2024.
[0081] Fig. 38. AM Deaminase worked
[0082] Fig. 39. Docking & Deamination (D&D-seq) allows for mapping ProteintDNA binding via molecular footprinting, a. Schematic illustration of D&D-seq technology. TF: transcription factor. BE: base editor (deaminase), b. Illustration of the split DddA base editor comprising N-terminal (DddA_NT) and C-terminal (DddA-CT) regions. DddA_NT is fused to nanobodies and is activated by addition of the C-terminal peptide, c. In vitro deamination assay (Fig. 43d) to quantify deamination efficiency in the presence or absence of Zn2+. The bars represent the average DNA intensity (n = 2) of the cleaved product quantified by ImageJ for split DddAl 1 Rb-DddA_NT, DddA_CT, or both. d. In vitro deamination assay (Fig. 43e) to quantify deamination efficiency for the split DddAl 1 deaminase and the uracil-targeting USER enzyme in different dinucleotide contexts. The bars represent the average DNA intensity (n = 2) of the cleaved product quantified by ImageJ. e.-g. Averaged edit counts in (e) CTCF, (f) GATA1, and (g) GATA2 binding sites in K562. h. Molecular footprint of bulk D&D-seq for CTCF in K562 cells showing count of C-to-U edits at aggregated CTCF sites compared to a background AT AC region with no CTCF binding site. i. Same as h. for GATA1. j. Same as h. for GATA2. k. Motif enrichment analysis of bulk D&D-seq for CTCF in K562 cells. Genome- wide D&D-seq reads were analyzed for enrichment of the CTCF binding motif using MEME. 1. Same as k. for GATA1. m. Same as k. for GATA2. n. Rank plot of CTCF motif in AT AC peaks harboring C-to-U edits, o. Same as n. for GATA1. p. Same as n. for GATA2. q. Genome browser tracks for representative regions of the human genome. D&D-seq was performed on K562 cells for CTCF, GATA1 or GATA2. ENCODE ChlP-seq data for the 3 proteins are used as reference. Sequencing data were normalized as bins per million (BPM) mapped reads, r. Copy number alterations in K562 cells, s. Genome browser tracks for representative regions of the human genome. D&D signals in open and closed chromatin regions were extracted based on K562 ATAC-seq and CTCF ChlP-seq from Encode, t. Molecular footprint of bulk D&D-seq for CTCF in K562 cells showing count of C-to-U edits at aggregated open chromatin CTCF regions, for closed chromatin CTCF regions, and for CTCF negative regions.
[0083] Fig. 40. D&D-seq allow for single cell mapping of DNA:Protein interactions, a. UMAP of cell line mixing study of CA46 (n = 4,108 cells) and K562 (n = 1,732 cells) cells in the ATAC space, b. Violin plot showing the number of ATAC fragments per cell for each subpopulation from (a), c. Accessibility scores for marker genes defining the subpopulation lineage (FDR < 0.05). d. UMAP projection of gene accessibility score for CA46 marker genes RUBCNL and KLF12. e. UMAP projection of gene accessibility score for K562 marker genes NUP214 and PRAME. f. Projection of CTCF D&D edit events into the UMAP space, g. Violin plot of CTCF D&D edit events in CA46 and K562 cells, h. Projection of GATA1 D&D edit events into the UMAP space, i. Violin plot of GATA1 D&D edit events in CA46 and K562 cells, j. Motif enrichment analysis for D&D-seq reads in CA46 cells, k. Molecular footprint of pseudo bulk D&D-seq for CTCF in CA46 cells showing count of C-to-T edits at aggregated CTCF sites compared to a background ATAC region with no CTCF binding site. 1. Motif enrichment analysis for D&D-seq reads in K562 cells, m. Molecular footprint of pseudobulk D&D-seq for GATA1 in K562 cells showing count of C-to-T edits at aggregated GATA1 sites compared to a background ATAC region with no GATA1 binding site. n. Pseudobulk genome browser tracks for representative regions of the human genome (chrl). scD&D-seq was performed on a mixture of K562 and CA46 cells for GATA1 or CTCF, respectively. ENCODE ChlP-seq data for the 2 proteins (black) as reference. ATAC peaks are in gray. Sequencing data were normalized as bins per million (BPM) mapped reads, o. Genome-wide Pearson correlation between pseudobulk scD&D-seq and ENCODE ChlP-seq for CTCF and GATA1 in K562 cells. Peaks are defined by ENCODE ChlP-seq for CTCF and GATA1. p. Fraction of fragments falling in ENCODE ChlP-seq peak regions for CTCF for D&D-seq (top) and uliCUT&RUN (bottom). No antibody (Ab) control indicates cells that were not stained with the CTCF antibody.
[0084] Fig. 41. D&D-seq allows for DNA:Protein interaction mapping and 3D chromatin inference in primary human cells, a. UMAP of primary PBMCs study (n = 5,358 cells) in the ATAC space. HSPC, hematopoietic stem / progenitor cells; eDC, conventional dendritic cells; NK, natural killer cells, b. Projection of CTCF D&D edit events into the UMAP space, c. Distribution of C-to-U D&D events around CTCF binding site (top) and ATAC (bottom) across all the genome, d. Motif enrichment analysis for D&D-seq reads, e. D&D-C.Origami-predicted (left) and experimental (right) Hi-C matrices of CD8 T cells, f. Pearson correlation of 100 randomly chosen CD8 T cells regions experimental Hi-C data (ENCODE) and: C.Origami 3D chromatin predictions generated with CTCF ChlP-seq data (ENCODE), CTCF D&D-seq or in the absence of CTCF information. Boxes represent the interquartile range, lines represent the mean, and error bars represent the 1.5 interquartile ranges of the lower and upper quartile.
[0085] Fig. 42. Multiomic application of D&D-seq to human primary clonal hematopoiesis samples, a. Integrated chromatin accessibility uniform manifold approximation and projection (UMAP) after reciprocal latent semantic indexing (LSI) integration of patient samples (n = 2 replicates, 15,807 cells) illustrating the expected hematopoietic populations, b. Integrated UMAP colored by GoT-ChA-assigned 1DH2 genotype as wild type (WT; n = 4,177 cells) or mutant (MUT; n = 1,114 cells) and not assignable (NA; n = 10,516 cells), c. Percentage of cells genotyped as IDH2 mutant (MUT) for each cell subtype, d. Validation of GoT-ChA results showing percentage of mutant cells genotyped using bulk targeted sequencing for 1DH2 in each cell subtype, e. Projection of CTCF D&D edit events into the UMAP space, f. Number of D&D edit events recorded for each cell subtype, g. Motif enrichment analysis for D&D-seq reads, h. Uniform manifold approximation and projection (UMAP) of the CD8 T cell fraction of the sample (n = 7,546 cells) with genotypes annotated as mutant-enriched or wild type-enriched, i. Differential CTCF binding regions in mutant versus wild-type CD8 T cells. Total of 1,102 CTCF peaks with more than five D&D edits and five cells were tested using LMM (see Methods). Gene regions with false discovery rate < 0.25 are highlighted, j. Heatmap displaying D&D C. Origami Hi-C predicted interactions enriched in mutant versus wild- type CD8 T cells at the genomic region on chr 17 encompassing the G1T1 locus, k. Co-accessible regions determined from AT AC data (purple lines) in wild-type CD8 T cells at the GIT1 locus. CTCF binding events and AT AC signal detected with D&D-seq are displayed in the bottom tracks. 1. Co-accessible regions determined from ATAC data (purple lines) in mutant CD8 T cells at the GIT1 locus. CTCF binding events and ATAC signal detected with D&D-seq are displayed in the bottom tracks, m. Heatmap displaying pairwise ATAC-seq peak coaccessibility in wild-type CD8 T cells at the region corresponding to k and 1 co-accessibility plots (top-left). Zoom-in heatmap displaying pairwise ATAC-seq peak co-accessibility within a, b, c, and d peaks (bottom-right). Co-accessibility score was smoothed by averaging three adjacent peaks, n. Heatmap displaying pairwise ATAC-seq peak coaccessibility in mutant CD8 T cells at the region corresponding to k and 1 co-accessibility plots (top-left). Zoom-in heatmap displaying pairwise ATAC-seq peak co-accessibility within a, b, c, and d peaks (bottom-right). Co-accessibility score was smoothed by averaging three adjacent peaks.
[0086] Fig. 43. Docking & Deamination allow for mapping of transcription factor binding via molecular footprinting, a. Schematic representation of the pTXB 1 -nanobody - DddAl 1_NT plasmid, b. SDS-PAGE of purified Rb-DddA_NT protein. The upper band is uncleaved Rb-DddA_NT with intein and chitin-binding domain (CBD) (54 KDa). The lower band is cleaved Rb-DddA_NT (26 KDa). c. Scheme of in vitro deamination assay using 36-bp 6-carboxyfluorescein (FAM)-labeled DNA oligos treated with D&D enzyme for conversion of C-to-U, followed by treatment with uracil- specific USER enzyme to create single-stranded nicks at U sites. Successful deamination can be detected by assessing single- stranded DNA, including generation of an 18 nt ssDNA fragment, on a denaturing gel. d. TBE-urea gel electrophoresis of DNA oligo treated with Rb-DddA_NT with or without Rb-DddA_CT and Zn2+. Rb-DddA_NT was activated upon addition of Rb- DddA_CT and Zn2+ further increased the activity. dsDNA, double- stranded DNA; ssDNA, single- stranded DNA. e. TBE-urea gel electrophoresis of DNA oligos containing different 5’ DNA contexts for the target cytosine treated with Rb-DddA_NT and Rb-DddA_CT. Rb- DddA preferentially deaminates cytosines that are preceded by thymines, f. Scheme of in vitro deamination assay using lambda phage DNA. g. E-gel of lambda phage DNA treated with Rb-DddA_NT+CT (DddA) overnight at different temperatures (°C). RT, room temperature, h. E-gel of lambda phage DNA treated with Rb-DddA_NT+CT (DddA) at 37 °C for different duration (min). In lane 4, only Rb-DddA_NT is added as a negative control, i. DNA mutation signature of bulk K562 CTCF D&D-seq in the background regions. See Fig. 39e (Methods), j. Same as j. for GATA1. See Fig. 39f., k. Same as j. for GATA2. See Fig. 39g., 1. Dinucleotide context frequency of edited cytosine in bulk K562 CTCF D&D-seq data. m. Trinucleotide context frequency of edited cytosine in bulk K562 CTCF D&D-seq data. n. DNA mutation signature of bulk K562 P300 D&D-seq. Target regions are defined by ChlP-seq (Methods), o. Footprint plot of p300 D&D-seq showing distribution of C-to-U edit counts in ChlP-seq defined target peaks and background peaks, p. Motif enrichment analysis to identify frequently found binding sites in top 10% most edited p300 interacting regions, q. Top 3 transcription factor binding motifs enriched in p300 interacting regions, r. PWM motifs and rank plots of JUN (top), SP1 (middle), and TALI (bottom) of D&D edit peaks, s. Footprint plots of JUN (top), SP1 (middle), and TALI (bottom) showing distribution of C-to-U edit counts in motif-defined target peaks and background peaks.
[0087] Fig. 44. D&D-seq is compatible with existing droplet-based single-cell chromatin profiling technologies, a. Quality control metrics of cell mixing scATAC-seq study, including number of counts in peaks (nCount_peaks), TSS enrichment scores (TSS. enrichment), fraction of reads in blacklist (blacklist_fraction), nucleosome signal scores (nucleosome_signal), and percentage of reads in peaks (pct_reads_in_peaks). b. DNA mutation signature of pseudobulk D&D-seq for GATA1 in K562 and CTCF in CA46. Target regions are defined by ChlP-seq. c. Molecular footprint of pseudobulk D&D-seq for GATA1 in CA46 cells and CTCF in K562 cells showing count of C-to-T edits at aggregated sites and suggesting no GATA 1 binding signal in CA46 cells and no CTCF binding signal in K562 cells, d. Genome-wide Pearson correlation between CUT&Tag from26and ENCODE ChlP-seq for H3K27me3 and H3K4me3 in K562 cells. Peaks are defined by ENCODE ChlP-seq for H3K27me3 and H3K4me3. e. Genome track plot of a CTCF binding site in CA46 with D&D edit reads from down- sampling sizes, f. Genome track plot of a GATA1 binding site in K562 with D&D edit reads from down- sampling sizes, g. Footprint signals of CTCF binding in CA46 down-sampling analyses. D&D edits were counted from 200 randomly selected peaks and average edit counts and standard deviations of 100 replicates were plotted, h. Footprint signals of GATA1 binding in K562 down-sampling analyses. D&D edits were counted from 200 randomly selected peaks and average edit counts and standard deviations of 100 replicates were plotted, i. Recall of D&D edits or target peaks in CA46 (left) or K562 (right) in various down-sampling sizes.
[0088] Fig. 45. D&D-seq allows for DNA:Protein interaction mapping in primary human cells, a. Quality control metrics of human primary PBMCs scATAC-seq, including number of counts in peaks (nCount_peaks), TSS enrichment scores (TSS. enrichment), fraction of reads in blacklist (blacklist_fraction), nucleosome signal scores (nucleosome_signal), and percentage of reads in peaks (pct_reads_in_peaks). b. Fragment length distribution of nucleosome signal scores (NS) <4 and > 4. c. Number of CTCF D&D edits per cell, separated by cell types, d. C. Origami-predicted (left) and experimental (right) Hi-C matrices of B cells, e. Pearson correlation of experimental Hi-C data (ENCODE) and: c. Origami 3D chromatin predictions generated with CTCF ChlP-seq data (ENCODE), CTCF D&D-seq, or in the absence of CTCF information in 100 randomly chosen regions. Boxes represent the interquartile range, lines represent the mean, and error bars represent the 1.5 interquartile ranges of the lower and upper quartile, f. C. Origami-predicted (left) and experimental (right) Hi-C matrices of monocytes, g. Same as e. but for monocytes.
[0089] Fig. 46. Multiomic application of D&D-seq to human primary clonal hematopoiesis samples, a. Schematic of D&D-GoT-ChA. b. Quality control metrics of scATAC-seq, including number of counts in peaks (nCount_peaks), TSS enrichment scores (TSS. enrichment), fraction of reads in blacklist (blacklist_fraction), nucleosome signal scores (nucleosome_signal), and percentage of reads in peaks (pct_reads_in_peaks). c. Differential gene activity score based on chromatin accessibility for cluster-specific marker genes, d. Distribution of WT and MUT read counts from GoT-ChA, separated by cell type. Thresholds are used to define MUT, WT and NA (Methods), with mutant cells in the top two quadrants, e. UMAP plot of AT AC fragment counts, f. Average variant counts per peak of CTCF from two replicates of the single-cell primary blood cells, g. DNA footprint of on- target and background peaks, h. Rank plots of CTCF binding sites for each cell type. Significant CTCF sites (FDR < 0.05) were marked. Chipseeker was used to annotate proximal genes of CTCF binding sites, i. UMAP of the CD8 T cell fraction of the sample (n = 7,546 cells) with genotypes annotated (left) and violin plot of D&D edits in the two genotypes (right), j. Heatmap displaying D&D C. Origami Hi-C predicted interactions enriched in mutant versus wild-type CD8 T cells at the genomic region on chr 11 encompassing the CNIH2 locus, k. Co-accessible regions determined from ATAC data (purple lines) in wild-type CD8 T cells at the CNIH2 locus. CTCF binding events detected with D&D-seq are displayed. 1. Same as k. for mutant CD8 T cells, m. Heatmap displaying pairwise ATAC-seq peak co-accessibility in wild-type CD8 T cells at the CNIH2 locus. Coaccessibility score was smoothed by averaging three adjacent peaks, n. Same as m. for mutant CD8 T cells.
[0090] Fig. 47. a. Number of peaks called in each condition (left) and footprint analysis of CTCF binding in K562 detected using primary antibody only (middle) and primary antibody with secondary antibody (right), b. Number of peaks called in each condition (left) and footprint analysis of GATA1 binding in K562 detected using primary antibody only (middle) and primary antibody with secondary antibody (right), c. Analytical workflow of scD&D-seq data. Reads from the same cell types were extracted from the bam file, followed by preprocessing and quality controls, variant pile-ups and filtering, peak calling, annotation, and finally evaluation of D&D reads, d. Number of variants in each analysis step for CA46 (top) and K562 (bottom) from the single-cell D&D-seq analysis, e. Annotation of residual somatic mutations in CA46 (top) and K562 (bottom). The final D&D edits were compared with cell line-specific variants from DepMap and COSMIC, or public sequencing data (ENCODE) from matching cell lines including CTCF NTT-seq, ATAC-seq, or whole-genome sequencing, f. Alignment of D&D edited reads from GATA1 and GATA2 experiments in the TALI locus (chrl:47, 232, 150-47 ,232, 550). Raw ATAC reads were binned in a lObp window. GATA1 motif position was indicated with dotted lines and C>T or G>A variants were marked with different colors.
[0091] Fig. 48. Secondary antibody boosts signal 3x.
[0092] Fig. 49. MGYPDa829 has no sequence motif bias.
[0093] Fig. 50. MGYPDa829 has no sequence motif bias.
[0094] Fig. 51. Gating strategy of flow sorting for different cell types. Detailed Description of the Invention
[0095] The present disclosure is based, at least in part, on the discovery of a new method to accurately map DNA:protein interactions with single-cell precision and sufficient power for application to human samples. In some embodiments, to directly profile TF or chromatin factor binding in single cells, a base-editing deaminase was tethered to nanobodies targeting a DNA binding protein of interest and combine this with single-cell ATAC-seq. Upon genome- wide antibody binding to the TF, the base editor enzyme deaminates cytosine to uracil, leaving a permanent genomic signature at regions of TF binding that can be identified through downstream single-cell sequencing of the tagmented loci. Facile incorporation into existing droplet-based single-cell sequencing workflows further allow for the addition of other molecular modalities onto this backbone technology, which was called D&D-seq (Docking and Deamination followed by sequencing). The sensitivity and specificity of the D&D enzyme in vitro, and perform bulk experiments to obtain genomewide TF binding profiles that have high concordance with ChlP-seq data was demonstrated. The robustness of single-cell D&D-seq through cell-line mixing experiments, identifying canonical GATA1 binding patterns in erythroid cells was tested. D&D-seq was further applied to primary human peripheral blood mononuclear cells (PBMCs), recovering high enrichment of CTCF binding motifs coinciding with D&D edits. In some embodiments, the single-cell genotype-aware GoT-ChA approach was integrated into the D&D-seq framework to simultaneously analyze genotype together with TF binding and chromatin accessibility, profiling CTCF binding in IDH2 -mutant clonal hematopoiesis of indeterminant potential (CHIP) samples. The tools and analyses developed here can be broadly applied for the interrogation of the interplay between TF binding and chromatin landscapes, as well as genotypes, in single cells directly from primary samples, opening up new avenues for the study of gene regulation and dysregulation across physiological or disease contexts.
[0096] In some embodiments, a new technique has been developed to profile protein binding events in cells. Briefly, novel protein A- or nanobody-base editor fusion protein was designed that, upon activation, can edit the DNA bases nearby. To profile protein-DNA binding, the target is incubated with primary antibodies, and then with protein A- or nanobody-base editor fusion proteins, followed by activation of the base editor that edits the DNA nearby. The recorded events are read by sequencing methods including ATAC-seq, whole genome sequencing, or other sequencing methods.
[0097] In some embodiments, an enzyme called D&D-deaminase was designed and engineered. Briefly, this enzyme is the fusion of DddA split base editor with secondary nanobodies that can recognize rabbit, mouse, goat, and rat Immunoglobulins. The enzyme is capable of binding to protein- specific antibodies and catalyzing cytosine deamination in the vicinity of the targeted genomic area. This deamination event results in the conversion of cytosine to thymidine on the genomic DNA. By using NGS sequencing, this conversion can be identified, providing a molecular footprint of the target of interest on the genomic DNA.
[0098] In some embodiments, with this enzyme, a method called Docking and Deamination followed by sequencing (D&D-seq) was developed. In some embodiments, isolated cells are first fixed and permeabilized to allow all the reagents to enter the cell. Then, the sample is incubated with an antibody specific to the DNA binding target of interest. After washing the antibody in excess, the sample is incubated with the D&D-deaminase specific to the antibody used to label the target protein. The nanobody domain of the D&D-deaminase specifically and stably bind with the antibody, localizing the enzyme in close proximity with the target. After washing the D&D-deaminase in excess, the sample can be processed with standard single-cell ATAC-seq methods. In ATAC-seq, genomic DNA is exposed to Tn5, a highly active transposase. Tn5 simultaneously fragments DNA, preferentially inserts into open chromatin sites, and adds sequencing primers (a process known as tagmentation). Open chromatin is identified from the sequenced DNA and data analysis can provide insight into gene regulation. The modified ATAC-buffer allows for simultaneous tagmentation and deamination in the same step. In this way, it is feasible to profile accessible chromatin and at the same time record the protein binding events on the genomic DNA (scheme).
[0099] Notably, during D&D-deaminase incubation, it is possible to maintain the enzyme in an inactive state, avoiding non-specific deamination resulting from non-tethered random interaction of the enzyme with the genomic DNA. In some embodiments, to achieve a switch-like control over enzyme activity, a split enzyme was use, where the DddA protein is separated into two polypeptides. The activity of the enzyme can be controlled by the addition of the C-terminal small peptide to the reaction, which promotes the dimerization of the two domains to reconstitute the split enzyme and its activity. In some embodiments, a novel method to map the DNA binding of Transcription Factors (TFs) or Chromatin Remodelers (CR) with single-cell resolution using commercially available single-cell techniques was developed. In some embodiments, an enzyme that can record protein binding events on accessible promoters have been designed and engineered. Once the targeted protein binding event has been recorded, the sample is processed with a standard ATAC-seq or other tagmentation-based protocols. In some embodiments, computational frameworks that can analyze data to identify the unique TFs / CR fingerprints generated by the protocol was additionally developed.
[0100] The present disclosure is unique in at least two aspects: (1) use of novel Nanobody- Base Editor fusion to map protein binding; and (2) use of split enzyme to control enzymatic activity after targeted tethering. The present disclosure can at least be used for profiling protein binding to genomic DNA with tagmentation-based single-cell multi-omics methods such as ATAC-seq, NTT-seq, and Dogma- seq; droplet-based single-cell methods, Split&Pool single-cell methods; or spatially resolved Chromatin profiling methods, such as Slide-seq, and dBIT-seq.
[0101] Definitions
[0102] The articles “a” and “an” are used herein to refer to one or to more than one (e.g., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0103] The term “amino acid” is intended to embrace all molecules, whether natural or synthetic, which include both an amino functionality and an acid functionality and capable of being included in a polymer of naturally-occurring amino acids. Example amino acids include naturally-occurring amino acids; analogs, derivatives and congeners thereof; amino acid analogs having variant side chains; and all stereoisomers of any of the foregoing.
[0104] The term “binding” or “interacting” refers to an association, which may be a stable association, between two molecules, e.g., between a peptide and a binding partner or agent, e.g., small molecule, due to, for example, electrostatic, hydrophobic, ionic and / or hydrogenbond interactions under physiological conditions.
[0105] The term “linker” refers to a molecule or group of molecules connecting two compounds, such as two polypeptides. The linker may be comprised of a single linking molecule or may comprise a linking molecule and a spacer molecule, intended to separate the linking molecule and a compound by a specific distance. The term “operably linked to” refers to the functional relationship of a nucleic acid with another nucleic acid sequence. Promoters, enhancers, transcriptional and translational stop sites, and other signal sequences are examples of nucleic acid sequences operably linked to other sequences. For example, operable linkage of DNA to a transcriptional control element refers to the physical and functional relationship between the DNA and promoter such that the transcription of such DNA is initiated from the promoter by an RNA polymerase that specifically recognizes, binds to and transcribes the DNA.
[0106] The terms “polynucleotide”, and “nucleic acid” are used interchangeably. They refer to a natural or synthetic molecule, or some combination thereof, comprising a single nucleotide or two or more nucleotides linked by a phosphate group at the 3’ position of one nucleotide to the 5’ end of another nucleotide. The polymeric form of nucleotides is not limited by length and can comprise either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure, and may perform any function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. A polynucleotide may be further modified, such as by conjugation with a labeling component. In all nucleic acid sequences provided herein, U nucleotides are interchangeable with T nucleotides. The polynucleotide is not necessarily associated with the cell in which the nucleic acid is found in nature, and / or operably linked to a polynucleotide to which it is linked in nature.
[0107] The term “specifically binds” or “specific binding”, as used herein, when referring to a polypeptide (including CAR polypeptides) refers to a binding reaction which is determinative of the presence of the protein or polypeptide or receptor in a heterogeneous population of proteins and other biologies. Thus, under designated conditions (e.g. immunoassay conditions in the case of an antibody), a specified ligand or antibody “specifically binds” to its particular “target” (e.g. an antibody specifically binds to an endothelial antigen) when it does not bind in a significant amount to other proteins present in the sample or to other proteins to which the ligand or antibody may come in contact in an organism. Generally, a first molecule that “specifically binds” a second molecule has an affinity constant (Ka) greater than about 105M1(e.g., 106M1. 107M1, 108M 109M IO10M1, 1011M1, and 1012M1or more) with that second molecule. For example, in the case of the ability of a PIG-specific CAR to bind to a peptide presented on an MHC (e.g., class I MHC or class II MHC); typically, a CAR specifically binds to its peptide / MHC with an affinity of at least a KD of about IO"4M or less, and binds to the predetermined antigen / binding partner with an affinity (as expressed by KD) that is at least 10 fold less, at least 100 fold less or at least 1000 fold less than its affinity for binding to a non-specific and unrelated peptide / MHC complex (e.g., one comprising a BSA peptide or a casein peptide).
[0108] The term “vector” refers to the means by which a nucleic acid can be propagated and / or transferred between organisms, cells, or cellular components. Vectors include plasmids, viruses, bacteriophage, pro-viruses, phagemids, transposons, and artificial chromosomes, and the like, to which the nucleic acid has been linked, and may or may not be able to replicate autonomously or integrate into a chromosome of a host cell. Such vectors may include any vector, (e.g., a plasmid, cosmid or phage chromosome) containing a gene construct in a form suitable for expression by a cell (e.g., linked to a transcriptional control element).
[0109] The percent identity between two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity= # of identical positions / total # of positions x 100), taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm, as described in the non-limiting examples below.
[0110] Fusion Proteins
[0111] In some aspects, provided herein is a fusion protein comprising a base-editing enzyme and a protein that binds to a fragment crystallizable region (Fc region) of an immunoglobulin .
[0112] In some embodiments, the Fc region is an Fc region from IgGl, IgG2, IgG3, or IgG4. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is or comprises an antibody or antigen-binding fragment thereof, a protein G, and / or a protein A. In some embodiments, the protein that binds to a Fc region of an antibody is or comprises an antibody or antigen-binding fragment thereof. In some embodiments, antigenbinding fragment thereof is selected from a Fab, a single chain variable region (scFv), a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody. In some embodiments, the antigen-binding fragment thereof is a nanobody. In some embodiments, the nanobody recognizes a rabbit, mouse, goat or rat immunoglobulin.
[0113] In some aspects, provided herein is a fusion protein comprising a base-editing enzyme and a protein that binds to a DNA binding protein. In some embodiments, the protein that binds to a DNA binding protein is an antibody or antigen-binding fragment thereof.
[0114] In some embodiments, the antigen-binding fragment thereof is selected from a Fab, a single chain variable region (scFv), a nanobody, a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody. In some embodiments, the antigen-binding fragment thereof is a nanobody. In some embodiments, the nanobody comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 1-4. In some embodiments, the nanobody comprises the amino acid sequence of any one of SEQ ID NOs: 1-4. For example, in some embodiments the nanobody comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 1. For example, in some embodiments the nanobody comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 2. For example, in some embodiments the nanobody comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 3. For example, in some embodiments the nanobody comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 4.
[0115] In some embodiments, the protein that binds to a Fc region of an immunoglobulin is or comprises a protein G. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is a protein G. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is a fusion protein that comprises a protein G, e.g., a protein AG. In some embodiments, the protein G comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 16. In some embodiments, the protein G comprises the amino acid sequence of SEQ ID NO: 16. For example, in some embodiments, the protein G comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 16.
[0116] In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 16. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises the amino acid sequence of SEQ ID NO: 16. For example, in some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 16. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is or comprises a protein A. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is a protein A. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is a fusion protein that comprises a protein A, e.g., a protein AG.
[0117] In some embodiments, the protein A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the protein A comprises the amino acid sequence of SEQ ID NO: 5. For example, in some embodiments, the protein A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 5.
[0118] In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises the amino acid sequence of SEQ ID NO: 5. For example, in some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the protein that binds to a Fc region of an immunoglobulin is a fusion protein that comprises a protein A and a protein G, e.g., pAG.
[0119] In some embodiments, the pAG comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 20. In some embodiments, the pAG comprises the amino acid sequence of SEQ ID NO: 20. For example, in some embodiments, the pAG comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 20.
[0120] In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 20. In some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises the amino acid sequence of SEQ ID NO: 20. For example, in some embodiments, the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 20.
[0121] In some embodiments, the base-editing enzyme is a base-editing deaminase. In some embodiments, the base-editing deaminase deaminates cytosine to uracil (C>U). In some embodiments, the base-editing enzyme is selected from DddA, CbDaOl, AcDaOl, non-DddA-like C>U dsDNA deaminases or DddA-like. In some embodiments, the DddA base-editing enzyme is DddAl 1. In some embodiments, the DddA-like base-editing enzyme is MGYPDa829 or LbsDaOl. In some embodiments, the DddA-like base-editing enzyme is MGYPDa829. In some embodiments, the base-editing enzyme comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 6-8. In some embodiments, the base-editing enzyme comprises an amino acid sequence selected from SEQ ID NOs: 6- 8. For example, in some embodiments, the base-editing enzyme comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 6. For example, in some embodiments, the base-editing enzyme comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 7. For example, in some embodiments, the base-editing enzyme comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 8. In some embodiments, the base-editing enzyme is a nonsplit full length base editor.
[0122] In some embodiments, the base-editing enzyme is a split base editor. In some embodiments, the base-editing enzyme is a DddA split base editor. In some embodiments, the DddA split base editor comprises an N-terminal portion of DddA (DddA_NT). In some embodiments, the N-terminal portion of DddA (DddA_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 9. In some embodiments, the N-terminal portion of DddA (DddA_NT) comprises the amino acid sequence of SEQ ID NO: 9. For example, in some embodiments, the N-terminal portion of DddA (DddA_NT) comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 9.
[0123] In some embodiments, the DddA split base editor is activated by the addition of a C- terminal portion of DddA (DddA_CT). In some embodiments, the C-terminal peptide comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 10 or 17. In some embodiments, the C-terminal peptide comprises the amino acid sequence of SEQ ID NO: 10 or 17. For example, in some embodiments, the C-terminal peptide comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 10. For example, in some embodiments, the C-terminal peptide comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 17. In some embodiments, the C-terminal peptide is or comprises the amino acid sequence of SEQ ID NO: 17 without 1-5 (e.g., 1, 2, 3, 4, or 5) amino acids at the C-terminus.
[0124] In some embodiments, the base-editing enzyme is a MGYPDa829 split base editor. In some embodiments, the MGYPDa829 split base editor comprises an N-terminal portion of MGYPDa829 (M829_NT). In some embodiments, the N-terminal portion of MGYPDa829 (M829_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 11. In some embodiments, the N-terminal portion of MGYPDa829 (M829_NT) comprises the amino acid sequence of SEQ ID NO: 11. For example, in some embodiments, the N-terminal portion of MGYPDa829 (M829_NT) comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 11.
[0125] In some embodiments, the MGYPDa829 split base editor is activated by the addition of a C-terminal portion of MGYPDa829 (M829_CT). In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 12. In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises the amino acid sequence of SEQ ID NO: 12. For example, in some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 12.
[0126] In some embodiments, the MGYPDa829 split base editor is activated by the addition of a C-terminal portion of MGYPDa829 (M829_CT). In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 19. In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises the amino acid sequence of SEQ ID NO: 19. For example, in some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 19. In some embodiments, the C-terminal portion of MGYPDa829 (M829_CT) is or comprises the amino acid sequence of SEQ ID NO: 19 without 1-6 (e.g., 1, 2, 3, 4, 5, or 6) amino acids at the C-terminus.
[0127] In some embodiments, the base-editing deaminase deaminates A to I (A>I). In some embodiments, the A>I deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. In some embodiments, the adenine base editor is fused to the N-terminus of the DddA variant. In some embodiments, the adenine base editor is TadA, optionally wherein the adenine base editor is TadA8e. In some embodiments, the DddA variant is DddA E1347A. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13. For example, in some embodiments, the AM deaminase comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 13.
[0128] In some embodiments, the AM deaminase is hADA. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 22. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 22. For example, in some embodiments, the AM deaminase comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 22.
[0129] In some embodiments, the AM deaminase is a non-split full length AM deaminase.
[0130] In some embodiments, the AM deaminase is a split AM deaminase. In some embodiments, the AM deaminase split base editor comprises TadA8e fused to an N- terminal portion of DddA E1347A. In some embodiments, the N-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14. In some embodiments, the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14. For example, in some embodiments, the N-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 14.
[0131] In some embodiments, the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A. In some embodiments, the C-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18. In some embodiments, the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18. For example, in some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 15. For example, in some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 18.
[0132] In some embodiments, the C-terminal portion of DddA E1347A is or comprises the amino acid sequence of SEQ ID NO: 18 without 1-5 (e.g., 1, 2, 3, 4, or 5) amino acids at the C-terminus.
[0133] In some embodiments, the fusion protein is capable of binding to a DNA binding protein. In some embodiments, the fusion protein is capable of binding to a DNA binding protein- specific antibody. In some embodiments, the fusion protein is capable of catalyzing cytosine (C) or adenosine (A) deamination in the vicinity of the DNA binding protein- targeted genomic area. In some embodiments, the DNA-binding protein binds to a DNA directly. In some embodiments, the DNA-binding protein binds to a DNA indirectly. In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
[0134] Nucleic acid, Vectors, Host Cells
[0135] In some aspects, provided herein is a nucleic acid encoding the fusion protein described herein. The nucleic acid can be a DNA or an RNA.
[0136] In some aspects, provided herein is a vector comprising the nucleic acid described herein. In some embodiments, the vector is a plasmid. In some embodiments, the vector is a viral vector, for example, a lentiviral vector, or a retroviral vector. In some embodiments, the vector is an Adeno-associated virus (AAV).
[0137] Methods
[0138] In some aspects, provided herein is a method of identifying a binding event of a DNA-binding protein in a cell or a population of cells, comprising: (a) contacting the cell or the population of cells with a fusion protein described herein; (b) incubating the cell or the population of cells at a condition for a period of time; and (c) detecting base changes induced by the fusion protein by sequencing, thereby identifying the binding event.
[0139] In some embodiments, the DNA-binding protein binds to a DNA directly. In some embodiments, the DNA-binding protein binds to a DNA indirectly. In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the method profiles all binding events of the DNA-binding protein in the cell or the population of cells. In some embodiments, the population of cells comprises at least 50 cells, at least 100 cells, at least 150 cells, at least 200 cells, at least 250 cells, at least 300 cells, at least 350, at least 400, at least 450, or at least 500. In some embodiments, the cell is a primary cell, optionally wherein the cell is a primary human cell, e.g., primary human PBMCs. In some embodiments, the cell is a cell line (e.g., a human cell line).
[0140] In some embodiments, the cell or the population of cells are contacted with an antibody specific to the DNA-binding protein prior to step (a). In some embodiments, the cell or the population of cells are fixed and permeabilized prior to step (a).
[0141] In some embodiments, step (b) occurs in the presence of the C-terminal portion of DddA (DddA_CT). In some embodiments, step (b) occurs in the presence of a divalent cofactor. In some embodiments, step (b) occurs in the presence of Zn2+. In some embodiments, step (b) occurs at a temperature from 37°C to 50 °C.
[0142] In some embodiments, the sequencing at step (c) is next generation sequencing. In some embodiments, the sequencing at step (c) is single cell sequencing. In some embodiments, the sequencing at step (c) is tagmentation-based single cell sequencing. In some embodiments, the sequencing at step (c) is Assay for Transposase- Accessible Chromatin using sequencing (ATAC-seq), single-cell (sc) ATAC-seq, Nanobody Tethered Tagmentation followed by sequencing (NTT-seq), CUT&Tag-related seq, or Dogma-seq. In some embodiments, the sequencing at step (c) comprises Tn5 or TnY tagmentation. In some embodiments, the sequencing at step (c) uses a lOx Genomics single-cell (sc) ATAC-seq kit. In some embodiments, the sequencing at step (c) is a droplet-based single cell sequencing method. In some embodiments, the sequencing at step (c) is a split and pool single -cell sequencing method. In some embodiments, the sequencing at step (c) is a spatially resolved chromatin profiling method. In some embodiments, the sequencing at step (c) is Slide- supported DNA sequencing or Slide-based sequencing (Slide- seq) or Deterministic Barcoding in Tissue for Spatial Omics Sequencing (dBIT-seq). In some embodiments, the sequencing at step (c) is whole-genome sequencing. In some embodiments, the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing. In some embodiments, the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PTA) protocol. In some embodiments, the method can be used to predict 3D chromatin structure in the cell or the population of cells.
[0143] In some aspects, provided herein is a method of identifying a binding event of a DNA-binding protein in a cell or a population of cells, comprising: (a)fixing and permeabilizing the cell or the population of cells; (b) contacting the cell or the population of cells with a primary antibody that specifically binds the DNA-binding protein; (c) contacting the cell or the population of cells with a fusion protein comprising (1) a protein that binds the Fc region of the primary antibody, optionally wherein the protein is a secondary nanobody, a protein A, or a protein G; and (2) an N-terminal portion of a baseediting enzyme, optionally wherein the base-editing enzyme is selection from DddAll, MGYPDa829, LbsDaOl, or TadA8e-DddAE1347A; (d) adding a C-terminal portion of the base-editing enzyme and a divalent cofactor; (e) performing Tagmentation, optionally wherein the Tagmentation comprises Tn5 or TnY tagmentation; (f) extracting genomic DNA from the cell and constructing a library for sequence; and (g) sequencing the genomic DNA library to detect base changes induced by the fusion protein, thereby identifying the binding event.
[0144] In some embodiments, the method further comprises conducting a genotyping method. In some embodiments, the genotyping method is a high-throughput single-cell genotyping method. In some embodiments, the genotyping method is Genotyping of Transcriptomes (GoT) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT- ChA). In some embodiments, the method further comprises conducting cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), NTT-seq, ATAC with select antigen profiling by sequencing (ASAP- seq), simultaneous high-throughput ATAC and RNA expression with sequencing (SHARE-seq), Paired-Tag, and / or combined assay of transcriptome and enriched chromatin binding (CoTECH).
[0145] In some aspects, provided herein is a method of identifying binding events of two DNA-binding proteins in a cell or a population of cells, comprising: (a) contacting the cell or the population of cells with the system or kit described herein; (b) incubate the cell or the population of cells at a condition for a period of time; and (c) detecting base changes induced by the system or kit by sequencing, thereby identifying the binding events of two DNA-binding proteins.
[0146] In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the method profiles all binding events of the two DNA-binding proteins in the cell or the population of cells. In some embodiments, the cell is a primary cell or a cell line.
[0147] In some embodiments, the cell or the population of cells are contacted with two primary antibodies specific to the two DNA-binding proteins, respectively, prior to step (a). In some embodiments, the cell or the population of cells are fixed and permeabilized prior to step (a).
[0148] In some embodiments, the sequencing at step (c) is single cell sequencing. In some embodiments, the sequencing at step (c) is tagmentation-based single cell sequencing. In some embodiments, the sequencing at step (c) is ATAC-seq, CUT&Tag-related seq, NTT- seq, or Dogma-seq. In some embodiments, the sequencing at step (c) comprises Tn5 or TnY tagmentation. In some embodiments, the sequencing at step (c) uses a lOx Genomics single-cell Assay for Transposase- Accessible Chromatin using sequencing (scATAC-seq) kit. In some embodiments, the sequencing at step (c) is a droplet-based single cell sequencing method. In some embodiments, the sequencing at step (c) is a split and pool single-cell method. In some embodiments, the sequencing at step (c) is a spatially resolved chromatin profiling method. In some embodiments, the sequencing at step (c) is Slide-seq or dBIT-seq. In some embodiments, the sequencing at step (c) is whole genome sequencing. In some embodiments, the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing. In some embodiments, the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PTA) protocol.
[0149] In some aspects, provided herein is a method of identifying binding events of two DNA-binding proteins in a cell or a population of cells, comprising: (a) fixing and permeabilizing the cell or the population of cells; (b) contacting the cell or the population of cells with two primary antibodies specific to the two DNA-binding proteins, respectively;
[0150] (c) contacting the cell or the population of cells with (1) a first fusion protein comprising a first nanobody and a first deaminase; and (2) a second fusion protein comprising a second nanobody and a second deaminase, wherein (i) the first deaminase and the second deaminase have different motif preferences or the first deaminase is a AM deaminase and the second deaminase is a C>U deaminase; and (ii) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species; and (d) sequencing genomic DNA of the cell or the population of cells to detect base changes induced by the first and the second fusion proteins, respectively, thereby identifying the binding events of the two DNA-binding proteins.
[0151] In some embodiments, the first deaminase and / or the second deaminase is a split enzyme, and the step (c) further comprises contacting the cell with a C-terminal portion of the first deaminase to activate the first deaminase and / or with a C-terminal portion of the second deaminase to activate the second deaminase.
[0152] Kits and Systems
[0153] In some aspects, provided herein is a kit comprising the fusion protein described herein. In some embodiments, the kit further comprises a primary antibody that binds the DNA-binding protein. In some embodiments, the kit further comprises the C-terminal portion of DddA (DddA_CT) or the C-terminal portion of MGYPDa829 (MGYPDa829_CT). In some embodiments, the kit further comprises a divalent cofactor. In some embodiments, the divalent cofactor is Zn2+.
[0154] In some aspects, provided herein is a system or kit comprising: (a) a first fusion protein comprising a first nanobody and a first deaminase; and (b) a second fusion protein comprising a second nanobody and a second deaminase, wherein (1) the first deaminase and the second deaminase have different motif preferences; and (2) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species. In some embodiments, the first deaminase prefers CpG context, and the second deaminase prefers TpC context. In some embodiments, the first deaminase is CbDaOl or AcDaOl; and the second deaminase is DddA or DddA-like. In some aspects, provided herein is a system or a kit comprising: a first fusion protein comprising a first nanobody and a A>I deaminase; and a second fusion protein comprising a second nanobody and a OU deaminase, wherein the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
[0155] In some embodiments, the A>I deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. In some embodiments, the adenine base editor is fused to the N-terminal of the DddA variant. In some embodiments, the adenine base editor is TadA8e. In some embodiments, the DddA variant is DddA E1347A. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13. For example, in some embodiments, the AM deaminase comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 13.
[0156] In some embodiments, the AM deaminase is hADA. In some embodiments, the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 22. In some embodiments, the AM deaminase comprises the amino acid sequence of SEQ ID NO: 22. For example, in some embodiments, the AM deaminase comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 22.
[0157] In some embodiments, the AM deaminase is a non-split full length AM deaminase.
[0158] In some embodiments, the AM deaminase is a split AM deaminase. In some embodiments, the AM deaminase split base editor comprises TadA8e fused to an N- terminal portion of DddA E1347A. In some embodiments, the N-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14. In some embodiments, the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14. For example, in some embodiments, the N-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 14.
[0159] In some embodiments, the A>I deaminase split base editor is activated by the addition of a C-terminal portion of DddA El 347 A. In some embodiments, the C-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18. In some embodiments, the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18. For example, in some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 15. In some embodiments, the C-terminal portion of DddA E1347A comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence of SEQ ID NO: 18.
[0160] In some embodiments, the system or kit further comprises primary antibodies to the two different DNA-binding proteins. In some embodiments, the system or kit further comprises the C-terminal portion of DddA E1347A.
[0161] In some embodiments, the DNA-binding protein is a transcription factor or a chromatin remodeling factor. In some embodiments, the two DNA-binding proteins are cofactors, or different components from a DNA-binding protein complex.
[0162] ADDITIONAL EMBODIMENTS:
[0163] Embodiment 1. A fusion protein comprising a base-editing enzyme and a nanobody. Embodiment 2. The fusion protein of embodiment 1, wherein the nanobody targets a DNA binding protein.
[0164] Embodiment 3. The fusion protein of embodiment 1, wherein the nanobody is a secondary nanobody that recognizes rabbit, mouse, goat and rat immunoglobulins. Embodiment 4. The fusion protein of any one of embodiments 1-3, wherein the nanobody comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 1-4. Embodiment 5. The fusion protein of embodiment 4, wherein the nanobody comprises the amino acid sequence of any one of SEQ ID NOs: 1-4.
[0165] Embodiment 6. A fusion protein comprising a base-editing deaminase and a protein A.
[0166] Embodiment 7. The fusion protein of embodiment 6, wherein the protein A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 5.
[0167] Embodiment 8. The fusion protein of embodiment 7, wherein the protein A comprises the amino acid sequence of SEQ ID NO: 5.
[0168] Embodiment 9. The fusion protein of any one of embodiments 1-8, wherein the baseediting enzyme is a base-editing deaminase.
[0169] Embodiment 10. The fusion protein of embodiment 9, wherein the base-editing deaminase deaminates cytosine to uracil (OU).
[0170] Embodiment 11. The fusion protein of embodiment 10, wherein the base-editing enzyme is selected from DddA, CbDaOl, AcDaOl, or DddA-like.
[0171] Embodiment 12. The fusion protein of embodiment 11, wherein the DddA baseediting enzyme is DddA 11.
[0172] Embodiment 13. The fusion protein of embodiment 11, wherein the DddA-like baseediting enzyme is MGYPDa829 or LbsDaOl.
[0173] Embodiment 14. The fusion protein of any one of embodiments 1-13, wherein the base-editing enzyme comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 6-8.
[0174] Embodiment 15. The fusion protein of embodiment 14, wherein the base-editing enzyme comprises an amino acid sequence selected from SEQ ID NOs: 6-8.
[0175] Embodiment 16. The fusion protein of any one of embodiments 1-13, wherein the base-editing enzyme is a split base editor.
[0176] Embodiment 17. The fusion protein of embodiment 16, wherein the base-editing enzyme is a DddA split base editor.
[0177] Embodiment 18. The fusion protein of embodiment 17, wherein the DddA split base editor comprises an N-terminal portion of DddA (DddA_NT).
[0178] Embodiment 19. The fusion protein of embodiment 18, wherein the N-terminal portion of DddA (DddA_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 9. Embodiment 20. The fusion protein of embodiment 19, wherein the N-terminal portion of DddA (DddA_NT) comprises the amino acid sequence of SEQ ID NO: 9.
[0179] Embodiment 21. The fusion protein of any one of embodiments 17-20, wherein the DddA split base editor is activated by the addition of a C-terminal portion of DddA (DddA_CT).
[0180] Embodiment 22. The fusion protein of embodiment 21, wherein the C-terminal peptide comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 10.
[0181] Embodiment 23. The fusion protein of embodiment 22, wherein the C-terminal peptide comprises the amino acid sequence of SEQ ID NO: 10.
[0182] Embodiment 24. The fusion protein of embodiment 16, wherein the base-editing enzyme is a MGYPDa829 split base editor.
[0183] Embodiment 25. The fusion protein of embodiment 24, wherein the MGYPDa829 split base editor comprises an N-terminal portion of MGYPDa829 (M829_NT).
[0184] Embodiment 26. The fusion protein of embodiment 25, wherein the N-terminal portion of MGYPDa829 (M829_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 11.
[0185] Embodiment 27. The fusion protein of embodiment 26, wherein the N-terminal portion of MGYPDa829 (M829_NT) comprises the amino acid sequence of SEQ ID NO: 11.
[0186] Embodiment 28. The fusion protein of any one of embodiments 24-27, wherein the MGYPDa829 split base editor is activated by the addition of a C-terminal portion of MGYPDa829 (M829_CT).
[0187] Embodiment 29. The fusion protein of embodiment 28, wherein the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 12.
[0188] Embodiment 30. The fusion protein of embodiment 29, wherein the C-terminal portion of MGYPDa829 (M829_CT) comprises the amino acid sequence of SEQ ID NO: 12.
[0189] Embodiment 31. The fusion protein of embodiment 9, wherein the base-editing deaminase deaminates A to I (A>I).
[0190] Embodiment 32. The fusion protein of embodiment 31, wherein the A>I deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. Embodiment 33. The fusion protein of embodiment 32, wherein the adenine base editor is fused to the N-terminus of the DddA variant.
[0191] Embodiment 34. The fusion protein of embodiment 32 or 33, wherein the adenine base editor is TadA or another ssDNA adenine base editor, optionally wherein the adenine base editor is TadA8e.
[0192] Embodiment 35. The fusion protein of any one of embodiments 32-34, wherein the DddA variant is DddA El 347 A.
[0193] Embodiment 36. The fusion protein of any one of embodiments 31-35, wherein the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13.
[0194] Embodiment 37. The fusion protein of embodiment 36, wherein the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13.
[0195] Embodiment 38. The fusion protein of any one of embodiments 31-35, wherein the AM deaminase is a split base editor.
[0196] Embodiment 39. The fusion protein of embodiment 38, wherein the AM deaminase split base editor comprises TadA8e fused to an N-terminal portion of DddA E1347A.
[0197] Embodiment 40. The fusion protein of embodiment 39, wherein the N-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14.
[0198] Embodiment 41. The fusion protein of embodiment 40, wherein the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14.
[0199] Embodiment 42. The fusion protein of any one of embodiments 38-41, wherein the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A.
[0200] Embodiment 43. The fusion protein of embodiment 42, wherein the C-terminal portion of DddA El 347 A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15.
[0201] Embodiment 44. The fusion protein of embodiment 43, wherein the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15.
[0202] Embodiment 45. The fusion protein of any one of embodiments 1-44, wherein the fusion protein is capable of binding to a DNA binding protein.
[0203] Embodiment 46. The fusion protein of any one of embodiments 1-45, wherein the fusion protein is capable of binding to a DNA binding protein- specific antibody. Embodiment 47. The fusion protein of any one of embodiments 1-46, wherein the fusion protein is capable of catalyzing cytosine deamination in the vicinity of the DNA binding protein-targeted genomic area.
[0204] Embodiment 48. The fusion protein of any one of embodiments 1-47, wherein the
[0205] DNA-binding protein is a transcription factor or a chromatin factor.
[0206] Embodiment 49. A nucleic acid encoding the fusion protein of any one of embodiments 1-48.
[0207] Embodiment 50. A vector comprising the nucleic acid of embodiment 49.
[0208] Embodiment 51. A kit comprising the fusion protein of any one of embodiments 1-48.
[0209] Embodiment 52. The kit of embodiment 51 , wherein the kit further comprises a primary antibody that binds the DNA-binding protein.
[0210] Embodiment 53. The kit of embodiment 51 or 52, wherein the kit further comprises the C-terminal portion of DddA (DddA_CT) or the C-terminal portion of MGYPDa829 (MGYPDa829_CT).
[0211] Embodiment 54. The kit of any one of embodiments 51-53, wherein the kit further comprises a divalent cofactor.
[0212] Embodiment 55. A method of identifying a binding event of a DNA-binding protein in a cell, comprising:
[0213] (a) contacting the cell with a fusion protein of any one of embodiments 1-48;
[0214] (b) incubating the cell at a condition for a period of time; and
[0215] (c) detecting base changes induced by the fusion protein by sequencing, thereby identifying the binding event.
[0216] Embodiment 56. The method of embodiment 55, wherein the DNA-binding protein is a transcription factor or a chromatin factor.
[0217] Embodiment 57. The method of embodiment 55 or 56, wherein the method profiles all binding events of the DNA-binding protein in the cell.
[0218] Embodiment 58. The method of any one of embodiments 55-57, wherein the sequencing is single cell sequencing.
[0219] Embodiment 59. The method of any one of embodiments 55-57, wherein the sequencing is tagmentation-based single cell sequencing.
[0220] Embodiment 60. The method of any one of embodiments 55-57, wherein the sequencing is ATAC-seq, Nanobody Tethered Tagmentation followed by sequencing (NTT-seq), Dogma-seq. Embodiment 61. The method of embodiment 60, wherein the sequencing uses a lOx Genomics scATAC-seq kit.
[0221] Embodiment 62. The method of any one of embodiments 55-57, wherein the sequencing is a droplet-based single cell sequencing method.
[0222] Embodiment 63. The method of any one of embodiments 55-57, wherein the sequencing is a split and pool single-cell sequencing method.
[0223] Embodiment 64. The method of any one of embodiments 55-57, wherein the sequencing is a spatially resolved chromatin profiling method.
[0224] Embodiment 65. The method of embodiment 64, wherein the sequencing is Slide- seq or dBIT-seq.
[0225] Embodiment 66. The method of any one of embodiments 55-65, wherein the cell is a primary cell or a cell line.
[0226] Embodiment 67. The method of any one of embodiments 55-66, wherein the step (b) occurs in the presence of the C-terminal portion of DddA (DddA_CT).
[0227] Embodiment 68. The method of any one of embodiments 55-67, wherein the step (b) occurs in the presence of a divalent cofactor.
[0228] Embodiment 69. The method of embodiment 68, wherein the step (b) occurs in the presence of Zn2+.
[0229] Embodiment 70. The method of any one of embodiments 55-69, wherein the step (b) occurs at a temperature from 37 °C to 50 °C.
[0230] Embodiment 71. The method of any one of embodiments 55-70, wherein the cells are contacted with an antibody specific to the DNA-binding protein prior to step (a).
[0231] Embodiment 72. The method of any one of embodiments 55-71, wherein the cells are fixed and permeabilized prior to step (a).
[0232] Embodiment 73. The method of any one of embodiments 55-72, further comprising Tn5 tagmentation after step (b) and before step (c).
[0233] Embodiment 74. The method of any one of embodiments 55-73, further comprising extracting genomic DNA from the cell and constructing a library for sequencing after step (b) and before step (c).
[0234] Embodiment 75. A method of identifying a binding event of a DNA-binding protein in a cell, comprising:
[0235] (a) fixing and permeabilizing the cell; (b) contacting the cell with primary antibody that specifically binds the DNA-binding protein;
[0236] (c) contacting the cell with a fusion protein comprising (1) a nanobody that binds the primary antibody and (2) an N-terminal portion of DddA;
[0237] (d) adding a C-terminal portion of DddA and a divalent cofactor;
[0238] (e) performing Tn5 tagmentation;
[0239] (f) extracting genomic DNA from the cell and constructing a library for sequence; and
[0240] (g) sequencing the genomic DNA library to detect base changes induced by the fusion protein, thereby identifying the binding event.
[0241] Embodiment 76. The method of any one of embodiments 55-75, wherein the method further comprises conducting a genotyping method.
[0242] Embodiment 77. The method of embodiment76, wherein the genotyping method is a high-throughput single-cell genotyping method.
[0243] Embodiment 78. The method of embodiment 75 or 76, wherein the genotyping method is Genotyping of Transcriptomes (GoT) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA).
[0244] Embodiment 79. The method of any one of embodiments 55-75, wherein the method further comprises conducting cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq).
[0245] Embodiment 80. The method of any one of embodiments 55-75, the method further comprises conducting NTT-seq.
[0246] Embodiment 81. A system or kit comprising:
[0247] (a) a first fusion protein comprising a first nanobody and a first deaminase; and
[0248] (b) a second fusion protein comprising a second nanobody and a second deaminase, wherein (1) the first deaminase and the second deaminase have different motif preferences; and (2) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
[0249] Embodiment 82. The system or kit of embodiment 81 , wherein the first deaminase prefers CpG context, and the second deaminase prefers TpC context.
[0250] Embodiment 83. The system or kit of embodiment 81 or 82, wherein the first deaminase is CbDaOl or AcDaOl; and the second deaminase is DddA or DddA-like. Embodiment 84. A system or a kit comprising: (a) a first fusion protein comprising a first nanobody and a A>I deaminase; and
[0251] (b) a second fusion protein comprising a second nanobody and a OU deaminase, wherein the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
[0252] Embodiment 85. The system or kit of embodiment 84, wherein the AM deaminase is a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant. Embodiment 86. The system or kit of embodiment 85, wherein the adenine base editor is fused to the N-terminal of the DddA variant.
[0253] Embodiment 87. The system or kit of embodiment 85 or 86, wherein the adenine base editor is TadA8e.
[0254] Embodiment 88. The system or kit of any one of embodiments 85-87, wherein the DddA variant is DddA El 347 A.
[0255] Embodiment 89. The system or kit of any one of embodiments 84-87, wherein the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13.
[0256] Embodiment 90. The system or kit of embodiment 89, wherein the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13.
[0257] Embodiment 91. The system or kit of any one of embodiments 84-87, wherein the AM deaminase is a split base editor.
[0258] Embodiment 92. The system or kit of embodiment 91, wherein the AM deaminase split base editor comprises TadA8e fused to an N-terminal portion of DddA E1347A.
[0259] Embodiment 93. The system or kit of embodiment 92, wherein the N-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14.
[0260] Embodiment 94. The system or kit of embodiment 93, wherein the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14.
[0261] Embodiment 95. The system or kit of any one of embodiments 91-94, wherein the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A.
[0262] Embodiment 96. The system or kit of embodiment 95, wherein the C-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15. Embodiment 97. The system or kit of embodiment 96, wherein the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15.
[0263] Embodiment 98. The system or kit of any one of embodiments 81-97 further comprising primary antibodies to the two different DNA-binding proteins.
[0264] Embodiment 99. The system or kit of any one of embodiments 81-98 further comprising the C-terminal portion of DddA E1347A.
[0265] Embodiment 100. The system or kit of any one of embodiments 81-99, wherein the DNA-binding protein is a transcription factor or a chromatin factor.
[0266] Embodiment 101. The system or kit of embodiment 100, wherein the two DNA-binding proteins are co-factors, or different components from a DNA-binding protein complex.
[0267] Embodiment 102. A method of identifying binding events of two DNA-binding proteins in a cell, comprising:
[0268] (a) contacting the cell with the system or kit of any one of embodiments 81-101;
[0269] (b) incubate the cell at a condition for a period of time; and
[0270] (c) detecting base changes induced by the system or kit by sequencing, thereby identifying the binding events of two DNA-binding proteins.
[0271] Embodiment 103. The method of embodiment 102, wherein the DNA-binding protein is a transcription factor or a chromatin factor.
[0272] Embodiment 104. The method of embodiment 102 or 103, wherein the method profiles all binding events of the two DNA-binding proteins in the cell.
[0273] Embodiment 105. The method of any one of embodiments 102-104, wherein the sequencing is single cell sequencing.
[0274] Embodiment 106. The method of any one of embodiments 102-104, wherein the sequencing is tagmentation-based single cell sequencing.
[0275] Embodiment 107. The method of any one of embodiments 102-104, wherein the sequencing is ATAC-seq, NTT-seq, Dogma-seq.
[0276] Embodiment 108. The method of embodiment 107, wherein the sequencing uses a lOx Genomics scATAC-seq kit.
[0277] Embodiment 109. The method of any one of embodiments 102-104, wherein the sequencing is a droplet-based single cell sequencing method.
[0278] Embodiment 110. The method of any one of embodiments 102-104, wherein the sequencing is a split and pool single-cell method. Embodiment 111. The method of any one of embodiments 102-104, wherein the sequencing is a spatially resolved chromatin profiling method.
[0279] Embodiment 112. The method of embodiment 111, wherein the sequencing is Slide- seq or dBIT-seq
[0280] Embodiment 113. The method of any one of embodiments 102- 112, wherein the cell is a primary cell or a cell line.
[0281] Embodiment 114. The method of any one of embodiments 102- 113, wherein the cells are contacted with two primary antibodies specific to the two DNA-binding proteins, respectively, prior to step (a).
[0282] Embodiment 115. The method of any one of embodiments 102- 114, wherein the cells are fixed and permeabilized prior to step (a).
[0283] Embodiment 116. The method of any one of embodiments 102- 115, further comprising Tn5 tagmentation after step (b) and before step (c).
[0284] Embodiment 117. The method of any one of embodiments 102- 116, further comprising extracting genomic DNA from the cell and constructing a library for sequencing after step
[0285] (b) and before step (c).
[0286] Embodiment 118. A method of identifying binding events of two DNA-binding proteins in a cell, comprising:
[0287] (a) fixing and permeabilizing the cell;
[0288] (b) contacting the cell with two primary antibodies specific to the two DNA-binding proteins, respectively;
[0289] (c) contacting the cell with (1) a first fusion protein comprising a first nanobody and a first deaminase; and (2) a second fusion protein comprising a second nanobody and a second deaminase, wherein (i) the first deaminase and the second deaminase have different motif preferences or the first deaminase is a A>I deaminase and the second deaminase is a C>U deaminase; and (ii) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species; and
[0290] (d) sequencing the genomic DNA library to detect base changes induced by the first and the second fusion proteins, respectively, thereby identifying the binding events of the two DNA-binding proteins.
[0291] Embodiment 119. The method of embodiment 118, wherein the first deaminase and / or the second deaminase is a split enzyme, and the step (c) further comprises contacting the cell with a C-terminal portion of the first deaminase and / or the second deaminase to activate the first deaminase and / or the second deaminase.
[0292] EXAMPLES
[0293] Example 1: Materials and Methods for Examples 2-6
[0294] Method
[0295] Cell culture
[0296] K562 cells were acquired from the American Type Culture Collection (CCL-243). HEK293FT cells were acquired from Thermo Fisher Scientific (R70007). HEK293FT cells were maintained at 37 °C and 5% CO2 in DIO medium (DMEM with high glucose and stabilized L-glutamine (Caisson, DML23) supplemented with 10% FBS (Thermo Fisher Scientific, 16000044). K562 cells were maintained at 37 °C and 5% CO2 in R10 medium (RPMI with stabilized L-glutamine (Thermo Fisher Scientific, 11875119) supplemented with 10% FBS). Human CA46 (ATCC, CRL-1648), HEL (ATCC, TIB-180), SET-2 (DSMZ, ACC 608), CCRF-CEM (ATCC, CRM-CCL-119), OCI-AML3 (provided by C. Park), HEPG2 with the FOXO1S22W mutation (provided by P. Campbell) and SUM159 (a gift from Y. Yarden) cell lines were maintained according to standard procedures in RPMI- 1640 (Thermo Fisher Scientific, 11-875-119) or in DMEM (HEPG2, Thermo Fisher Scientific, 11965092) with 10% (or 20% for CA46 and SET-2 cells) FBS (Thermo Fisher Scientific, 10-437-028) at 37 °C with 5% CO2. Cell lines in culture were screened biweekly for mycoplasma contamination using the MycoAlert PLUS Mycoplasma Detection Kit (Lonza, LT07-703).
[0297] Primary cell acquisition and processing
[0298] Fresh mobilized PBMCs used for single-cell nanobody-tethered transposition followed by sequencing (scNTT-seq) with cell surface protein measurement were isolated within 48 hours of blood collection using a Ficoll (Thermo Fisher Scientific, 45-001-750) gradient according to the manufacturer’ s recommendations and cryopreserved. Isolated mononuclear cells were thawed and stained according to standard procedures, beginning with resuspension in staining buffer (BioLegend, 420201) and incubation with Human TruStain FcX (10 minutes at 4 °C; BioLegend, 422302) to block Fc receptor-mediated binding. Cells were then stained with a CD34-PE-Vio770 antibody (20 minutes at 4 °C; Miltenyi Biotec, clone AC136, 130-113-180) and DAPI (Invitrogen, D1306). The samples were then sorted for DAPI-, CD34+ cells using a BD Influx cell sorter. Live CD34+ and CD34- cells were mixed 1:10 and processed with NTT-seq. BMMCs and PBMCs profiled by scNTT-seq without cell surface protein measurement were purchased from AllCells. After thawing into DMEM with 10% FBS, the cells were spun down at 4 °C for 5 minutes at 400g and washed twice with PBS with 2% BSA. After centrifugation, the cell pellet was resuspended in staining buffer (2% BSA and 0.01% Tween in PBS).
[0299] Cloning of nb-DddA constructs
[0300] Previously published sequences coding for the DddAl 1 were split at position 1397, and the C terminus of DddAl 1 (1290-1397) was synthesized as a gene fragment (Integrated DNA Technologies (IDT) flanked by restriction enzyme sites EcoRI and Spel. The gene fragments were digested with EcoRI and Spel for 1 hour at 37 °C, ligated for 1 hr at room temperature with pTXBl-alnbOc-Tn5 (Addgene, 184285), pTXBl-nbMm!gGl-Tn5 (Addgene, 184287), pTXBl-nbMm!gG2a-Tn5 (Addgene, 184288), pTXBl-nbMmKappa- Tn5 (Addgene, 184286), or 3xFlag-pA-Tn5-Fl (Addgene, 124601) digested by the same enzymes. The ligated product was transformed into competent cells (New England Biolabs, C2987) per the manufacturer’s instruction. The final plasmid products (pTXBl -nb-DddA) were confirmed by Sanger sequencing (Eton Bio) or Whole Plasmid Sequencing using Oxford Nanopore Technology with custom analysis and annotation (Plasmidsaurus). nb-DddA production
[0301] The pTXBl-nb-DddA_NT vectors were transformed into BL21(DE3)-competent Escherichia coli cells (NEB, C2527), and nb-DddA_NT was produced via intein purification with an affinity chitin-binding tag. First, transformed bacterial colonies were grown overnight in 5 ml Luria broth (LB). Next day, 5 ml of LB culture was added to 400 ml and grown at 37 °C to optical density (OD600) = 0.6. nb-DddA expression was induced with isopropyl-B-D-thiogalactopyranoside (IPTG) 0.5 mM and 50 pM ZnC12 at 30 °C for 4 hours. After induction, cells were pelleted and then frozen at -80 °C overnight. Cells were then lysed by sonication in 30 ml HEGX (20 mM HEPES-KOH pH 7.5, 0.8 M NaCl, 10% glycerol, 0.2% Triton X-100) with a protease inhibitor cocktail (Roche, 04693132001). The lysate was pelleted at 10,000 g for 20 minutes at 4 °C. The supernatant was transferred to a new tube, and 600 pl of neutralized 10% polyethyleneimine (Sigma- Aldrich, P3143) was added drop wise to the bacterial extract, gently mixed and centrifuged at 12,000g for 40 minutes at 4 °C to precipitate DNA. The supernatant was loaded on a 2-ml chitin column (NEB, S665 IS) and followed by washing with 12 ml of HEGX. Then, 3 ml of HEGX containing 100 mM DTT was added to the column with incubation for 48 hours at 4 °C to allow cleavage of nb-DddA from the intein tag. After incubation, 2 ml HEGX was added to elute nb-DddA directly into a 10-kDa molecular weight cutoff (MWCO) spin column (Thermo Scientific, 88527). Protein was dialyzed three times using 15 ml of 2x dialysis buffer (100 HEPES-KOH pH 7.2, 0.2 M NaCl, 0.2 mM EDTA, 2 mM DTT, 20% glycerol) and concentrated to 0.5-1 ml by centrifugation at 5,000g. The protein concentrate was transferred to a new tube, mixed with an equal volume of 100% glycerol, and stored at -20 °C.
[0302] Deamination assay
[0303] DNA substrates, including lambda phage DNA (NEB, N3011) or 5' 6-FAM-labeled 30-mer dsDNA oligonucleotides (IDT) were used for testing nb-DddA deamination activity. Deamination reactions of 250-1000 ng DNA were performed in 50 pl of deamination buffer (40 mM Tris-HCl pH 7.4, 50 mM KC1, 1 mM MgC12, 1 mM dithiothreitol (DTT), 20 pM ZnC12) with 50 pM nb-DddA_NT and 100 pM DddA_CT (Eton bio, HPLC-purified, purity >95%). The reactions were incubated at 37°C for 1 hr unless otherwise indicated. For lambda phage DNA, the products were then purified with 1.6X Ampure beads (Beckman Coulter, A63881), treated with 1 ul of T7 endonuclease (NEB, M0302) in 30 ul T7 buffer, further incubated at 37 °C for 15 min per manufacturer’s instructions, and analyzed by an agarose gel (Invitrogen, A42135). For 6-FAM-labeled dsDNA oligonucleotides, 1 ul USER enzyme (NEB, M5508) was added to the deaminase- treated products, incubated at 37 °C for 1 hr, denatured at 95 °C for 5 min, and analyzed by 15% polyacrylamide TBE-Urea gel (BIORAD, 4566055). The cleaved DNA fragments were imaged by an imager.
[0304] Antibodies
[0305] Antibodies used were CTCF (1:100, Active Motif, 61932), GATA1 (1:100, abeam, abl l852), GATA2 (1:100, Invitrogen, PAI-100), and HA tag (1:100). For DnD-seq with surface markers readout on primary cells, the TotalSeq-A conjugated Human Universal Cocktail version 1.0 panel was obtained from BioLegend (399907).
[0306] DnD-seq
[0307] Cell fixation, Permeabilization, and Antibody binding
[0308] For all the following centrifugation, cells were centrifuged in a swing-bucket centrifuge at 300g for 5 minutes before fixation and 600g for 5 minutes after fixation. About 200K to 1 million cells were collected and resuspended in 400 pl cold PBS. Then, 16% methanol-free formaldehyde (Thermo Fisher Scientific, PI28906) was added for fixation (final concentration 0.1%) at room temperature for 5 minutes. Glycine (final concentration 125 mM) was added to stop the cross-linking, followed by a wash with 1 ml of PBS. Fixed cells were permeabilized in 200 pl permeabilization buffer (20 mM Tris-HCl pH 7.4, 150 mM NaCl, 3 mM MgC12, 0.1% NP40, 0.1% Tween-20, 1% BSA, lx protease inhibitors) on ice (7 minutes for cell lines, 5 minutes for primary cells), followed by a wash with 200 pl cold wash buffer (20 mM HEPES pH 7.6, 150 mM NaCl, 0.5 mM spermidine, 1% BSA, lx protease inhibitors). After cell counting, 100K-400K cells were transferred to PCR tubes and resuspended in 100 pl antibody buffer (20 mM HEPES pH 7.6, 150 mM NaCl, 2 mM EDTA, 0.5 mM spermidine, 1% BSA, lx protease inhibitors), with 1 pl (1:100 dilution) of the antibodies. The cells were incubated at 4°C with slow rotation overnight.
[0309] DnD binding and activation
[0310] The cells were washed once with 200 ul wash buffer, and another time with 200 ul DnD binding buffer. The cells were resuspended in 50 ul DnD binding buffer with 50 uM nb-DddA_NT, and incubated at room temperature on a rotator for 1 hr. The cells were then washed with 200 ul DnD binding buffer, resuspended in 50 ul DnD activation buffer, and incubated at 37°C for 1 hr for deamination.
[0311] Bulk ATAC
[0312] Transposome assembly
[0313] Tn5 adaptors were purchased from IDT. Adaptors (100 uM) were annealed in water to form mosaic-end, double- stranded (MEDS) oligos by incubating at 95°C for 5 minutes and then cooling at 0.2°C per second to 12°C. MEDS-A and MEDA-B were mixed 1:1, and 2 pl was transferred to a new tube and mixed with 18 pl of TnY (in-house) enzyme after 1 hour at room temperature to allow for transposome assembly.
[0314] After DnD activation, the cells were washed once with 200 ul lx Tris-TD buffer (component) and resuspended in 50 ul lx Tris-TD buffer with 9 ul loaded TnY (volume depends on concentration, needs validation). To initiate tagmentation, the reaction was incubated at 37°C for 1 hr. To extract DNA, the reactions were added 1 ul 10% SDS, 2 ul proteinase K (NEB), and 3 ul 0.5M EDTA and incubated at 55°C for 1 hr, followed by column-based DNA purification per manufacturer’s instruction (Zymo, D5205).
[0315] Single-cell ATAC
[0316] After DnD activation, the cells were resuspended in lx nuclei buffer and counted using trypan blue and a Countess II FL Automated Cell Counter. For the rest, we follow single-cell ATAC-seq per lOx protocol (version CG000209 Rev F, lOx Genomics) with the following modifications.
[0317] 1. During the GEM generation and barcoding reaction (step 2.1), 2 ul of uracil- tolerant DNA polymerase (NEBNext Q5U Master Mix, NEB M0597S) was added to the barcoding reaction mixture to facilitate the first few rounds of PCR, where we expect to have some uradine in the reaction.
[0318] 2. For the GEM barcode PCR steps, we modified it to be: 72°C for 5 min; 6 cycles of 98°C for 40 s, 59°C for 2 min 30 s and 72 °C for 1 min; 6 cycles of 98°C for 10 s, 59°C for 30 s and 72°C for 1 min; followed by 72°C for 5 min and ending with hold at 4°C. The additional incubation time in the first 6 cycles allowed for more complete primer annealing and extension reaction.
[0319] Sequencing
[0320] The sequencing libraries were sequenced on a NovaSeq6000, NovaSeqX, or NextSeq2000 with dual indexed, paired-end 100 or 150bp settings. Specifically, i5: 8 bp (16 bp for single-cell sequencing), i7: 8 bp, readl: 100 or 150 bp, read2: 100 or 150 bp.
[0321] Genotyping of Targeted loci with Chromatin Accessibility (GoTChA)
[0322] We perform GoTChA according to the previously published paper with minor modifications . Briefly, we perform whole cell fixation, permeabilization, antibody incubation, DnD binding and activation as aforementioned bulk DnD-seq. After DnD activation, the cells were resuspended in lx diluted nucleus buffer (lOx Genomics) and counted using trypan blue and a Countess II FL Automated Cell Counter. Afterwards, the cells were processed according to the Chromium Next GEM Single Cell ATAC Solution user guide (version CG000209 Rev F, lOx Genomics) with the following modifications:
[0323] 1. During the GEM generation and barcoding reaction (step 2.1), 1 pl of 22.5 p M GoT-ChA primer mix was added to the barcoding reaction mixture. The primers used are IDH2 locus-specific primers IDH2 Fl (or F2) and IDH2 R1 (or R2). These primers allow for exponential amplification of the GoTChA fragments relative to the linear amplification of ATAC fragments. In addition, to facilitate the first few rounds of PCR, where we expect to have some uradine in the reaction, we spiked in 2 ul of uracil-tolerant DNA polymerase (NEBNext Q5U Master Mix, NEB M0597S) into the barcoding reaction mixture.
[0324] 2. For the GEM barcode PCR steps, we modified it to be: 72°C for 5 min; 6 cycles of 98°C for 40 s, 59°C for 2 min 30 s and 72 °C for 1 min; 6 cycles of 98°C for 10 s, 59°C for 30 s and 72°C for 1 min; followed by 72°C for 5 min and ending with hold at 4°C. The additional incubation time in the first 6 cycles allow for more complete primer annealing and reaction.
[0325] 3. During the post-GEM incubation clean-up (step 3.2), 45.5 pl of elution solution I is used to elute material from SPRIselect beads. A total of 5 pl is used for GoT ChA library construction, and the remaining 40 pl is used for AT AC library construction as indicated in the standard protocol.
[0326] 4. To generate the GoT-ChA library, two additional PCRs were performed on the 5 pl set aside during step 3.2. The first PCR aims to amplify genotyping fragments before sample indexing and uses P5 and IDH2 N1 (or N2) primers with the following thermocycler program: 95 °C for 3 min; 15 cycles of 95 °C for 20 s, 65 °C for 30 s and
[0327] 72 °C for 20 s; followed by 72 °C for 5 min and ending with hold at 4 °C. After a 1.2x SPRIselect clean-up, biotinylated PCR product is bound and isolated using Dynabeads M- 280 Streptavidin magnetic beads (Thermo Fisher Scientific, 11206D). In brief, the beads are washed three times with lx sodium chloride sodium phosphate-EDTA buffer (SSPE, VWR, VWRV0810-4L), added to the purified PCR product and incubated at room temperature for 15 min. The beads are then washed twice with lx SSPE buffer and once with 10 mM Tris-HCl (pH 8.0) before resuspending in water. The bead-bound fragments are then amplified and sample indexed using P5 and RPI-X primers with the following thermocycler program: 95 °C for 3 min; 6-10 cycles of 95 °C for 20 s, 65 °C for 30 s and 72 °C for 20 s; followed by 72 °C for 5 min and ending with hold at 4 °C.
[0328] Final libraries were quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Q32854) and the High Sensitivity DNA chip (Agilent Technologies, 5067-4626) run on a Bioanalyzer 2100 system (Agilent Technologies) and sequenced on a NovaSeq 6000 or NovaSeq X system at the Weill Cornell Medicine Genomics Resources Core Facility with the following parameters: paired-end 100 or 150 cycles; read IN, 100 or 150 cycles; i7 index, 8 cycles; i5 index, 16 cycles; read 2N, 100 or 150 cycles. ATAC libraries were sequenced to a depth of 25,000-35,000 read pairs per cell and GoT-ChA libraries were sequenced to 5,000 read pairs per cell.
[0329] Cell sorting + Genotyping
[0330] Cryopreserved peripheral blood mononuclear cells from a CH donor with IDH2 R140Q mutation was thawed, and stained. Briefly, the cells were resuspended in staining buffer (BioLegend, 420201) and incubated with Human TruStain FcX (10 min at 4 °C; BioLegend, 422302) to block Fc receptor-mediated binding. Then, the cells were stained with CD8-BV650, CD4-FITC, CD19-PE, CD14-APC, CD56-BV786 (1:100 for up to 10A6 cells in a final volume of 100 pl) for 20 min at 4 °C, and DAPI (Invitrogen, DI 306). The samples were then sorted for DAPI- and single-marker positive cells (CD8 for CD8+ T cells, CD4 for CD4+ T cells, CD19 for B cells, CD14 for monocytes, CD56 for NK cells) using the BD FACSymphony™ S6 Cell Sorter. The genomic DNA of the sorted cells were extracted using kit per manufacturer’s instruction, and PCR- amplified using KAPA 2X mix and IDH2 F2 + IDH2 R2 primers. The linear PCR amplicon sequencing was performed by Plasmidsaurus using Oxford Nanopore Technology with custom analysis, and the genotype ratio was quantified by the number of reads with mapped to IDH2 WT and IDH2 R140Q.
[0331] Bulk Data analysis
[0332] Raw bulk DnD-seq FASTQ data were analyzed using FastQC for initial quality control. Nextera transposase sequence and homopolymer G were removed using CutAdapt with “CTGTCTCTTATACACATCTCCGAGCCCACGAGAC” for R1 and “CTGTCTCTTATACACATCTGACGCTGCCGACGA” for R2 and "G{ 12}" parameters whenever identified in the FastQC report. After trimming adapter and polyG sequences, processed FASTQ files were re-analyzed using FastQC and aligned to the Human reference genome (GRCh38) using BWA-MEM2. Read alignments were subsequently sorted and indexed using samtools. To call DnD edits in bulk datasets, bam files were analyzed as described in the following DnD-mediated single nucleotide variant calling section.
[0333] Genomic regions of 400 bp to 1 kb (indicated in the figures) centered at the peak summits were used to query the genome for TF motif discovery using MEME. DNA footprint analysis was done on the Bam files using the ATACseqQC factorFootprints function. To visualize the genomic tracks on IGV, bigwig files were generated using the DeepTools bamCoverage function with the -normalize Using BPM option set. Tornado plots (heat maps) were generated in DeepTools. ChlP-seq peak coordinates for CTCF, GATA1, GATA2, TALI, MAX, and H3K27me3 for K562 cells were downloaded from ENCODE (ENCODE Project Consortium 2012).
[0334] Single cell data analysis
[0335] Preprocessing and annotation of 10X scATAC-seq data
[0336] Raw scATAC-seq datasets were aligned to the Human reference genome (GRCh38) and quantified using CellRanger-ATAC (version 2.1.0). The resulting peak / cell matrix was filtered for low quality cells and normalized using Signac (version 1.13.0)[ref|. In detail, cells having 1) reads aligned to the blacklist region more than 0.05%, 2) less than 20% of reads aligned to peaks, 3) transcription start site enrichment score less than 3, 4) nucleosome signal higher than 4, 5) outliers having too many (average of peak counts + 2*standard deviation of peak counts) or less peaks (average of peak counts - l*standard deviation of peak counts) were filtered out. Multiplets were additionally annotated using AMULET and removed[ref] . The preprocessed data was normalized by latent semantic indexing analysis, projected to low dimensions using uniform manifold approximation (UMAP), and clustered using the smart local moving (SLM) algorithm. Gene activity was inferred to annotate markers of each cluster. Finally, motif activity score was computed using ChromVAR with the JASPAR2020 core database for Homo sapiens [refs].
[0337] For the cell line data, eight clusters were defined with a clustering resolution of 0.6, and annotated as one of two cell lines based on the marker gene activity. Cluster 7 which was found in both K562 and CA46 clusters were removed. The remaining number of cells was 5840 cells (CA46, n = 4108 and K562, n = 1732). For PBMC data, 18 and 16 clusters were identified with a clustering resolution of 0.8 in replicate 1 and 2, respectively. Clusters were annotated based on marker gene activity and cell type prediction using Azimuth[ref] . For replicate 1, two clusters were additionally removed because two different cell types were mixed (cluster 7, CD8 T cell and monocyte) or it was overall low quality and lacked markers (cluster 13). The total of 5413 cells in replicate 1 and 10,394 cells in replicate 2 remained.
[0338] Genotyping of IDH2 mutation using GoT-ChA
[0339] Raw FASTQ files were first analyzed using FastQC to examine overall data quality and mutant allele in the targeted loci[ref]. Based on a FastQC report, position of mutant allele, a read file with mutant allele, and primer sequence were identified. Input FASTQ files were then processed using the GoT-ChA analysis pipeline described by authorsfref]. In detail, Raw data was first split into smaller files and quality filtering was performed with the “FastqFiltering” function. The R3 FASTQ file was used for genotyping in both replicates with c(128:130) or c(37:39) as a mutation site for replicates 1 and 2, respectively. After quality control of raw data, mutation state is annotated with the “BatchMutationCalling” function based on the provided wild-type / mutant and primer sequences. For replicate 1, the primer sequence was set as “CTGGCTGTGTTGTTGCTTGGGGTTCAAATTCTGGTTGAAAGATGGCGGCTGCA GTGGGACCACTATTATCTCTGTCCTCACAGAGTTCAAGCTGAAGAAGATGTGGA AAAGTCCCAATGGAACTATC”, and “CCG” and “CAG” were set as wild-type and mutant sequences. For replicate 2, the primer sequence was “GATGGGCTCCCGGAAGACAGTCCCCCCCAGGATGTT”, while sequences for wildtype and mutant were set to “CCG” and “CTG” because reads were reversed complementary. Results from each split data were merged into a single barcode and read count matrix using “MergeMutationCalling”. Genotype of each cell was assigned by manually setting thresholds based on distribution of read counts. As a result, 49.12% of cells in replicate 1 (2659 out of 5413 cells) and 25.32% of cells in replicate 2 (2632 out of 10,394 cells) were genotyped.
[0340] DnD-mediated single nucleotide variant calling
[0341] To summarize genome edits introduced by DnD in scATAC-seq data, a bam file produced by CellRanger was split into each cell type based on barcode sequence and cell annotation using sinto. For bulk ATAC-seq data, bam files aligned to the genome using BWA-MEM2 were analyzed[ref] . After splitting bam files, following six steps were performed to analyze DnD-mediated genomic variants.
[0342] 1. First, each bam file was preprocessed to remove uninformative and low quality read alignments. Duplicated reads were marked and simultaneously filtered using “picard MarkDuplicates” with a REMOVE_DUPLICATES=true parameter. Read alignments with high mapping quality Phred score (>= 20), primary alignment, reads aligned to intact chromosomes, and those with properly aligned mates were retained using samtools (version 1.19).
[0343] 2. Next, all single nucleotide variants (SNVs) found in each filtered bam file were collected using “bcftools mpileup” with following parameters, -a FORMAT / AD,FORMAT / DP,INFO / AD -no-BAQ -min-MQ 1 -max-depth 8000. The pileup result subsequently converted into the vcf format reporting SNVs supported by at least two mutant reads from minimum three aligned reads using an in-house Python script.
[0344] 3. Germline mutations were then filtered based on loci and alleles from the gnomAD database and variant allele frequency higher than 10%. If available, custom databases were provided to additionally filter uninformative mutations.
[0345] 4. Preprocessed bam files from step 1 were analyzed using MACS2 to call peaks with -fBAM —nomodel parameters. Peaks were then filtered with blacklist region annotation using bedtools (version 2.31.1). Motif analysis was performed using H0MER2 to identify binding sites in peaks. 5. If ChlP-seq data is available, target peaks were identified by comparing ChlP-seq tracks using bedtools intersect function with -wa. The remaining peaks that did not overlap with ChlP-seq were annotated as background peaks. Motif position was located in target peaks using H0MER2 with -find <motiffile> . The motif score in the <motiffile> was manually set to 0 to include as many peaks as possible in the target peaks. Motif analysis was also performed for background peaks with an unmodified motif file to rescue peaks missed in the ChlP-seq tracks. When ChlP-seq is not available, only motif analysis was performed on all peaks identified in the previous step.
[0346] 6. Target peaks were resized to 200 bp (up / downstream lOObp from the motif position) and overlaid with C-to-T and G-to-A variants identified in step 3. When multiple motif positions were found, a position with the highest score was chosen. Background peaks were resized to 200 bp by taking + / - lOObp from the peak summit.
[0347] Example 2:
[0348] The symphony of gene expression consists of an orchestra of regulatory factors that are directed through multiple interconnected epigenetic signals, including chromatin accessibility and histone modifications, that affect transcription factor (TF) binding. These complex networks have been shown to be disrupted in cells during aging, disease or cancer. However, mechanistic profiling of these intricate pathways through single-cell analysis is mostly lacking, due to technical limitations of current methods for mapping DNArprotein interactions in single cells. This has resulted in a large gap in knowledge about where TFs or other chromatin factors bind to DNA, and how this binding is perturbed in pathological contexts in human tissues. To address this challenge, we have developed a versatile, high- throughput single-cell DNA footprinting method that tethers a species-specific antibodybinding nanobody to a cytosine base editing enzyme. Combined with single-cell ATAC- seq, this approach enables the profiling of even weak or transient factor binding to DNA in single-cells through identifying cytosine to uracil edits in target genomic regions following tagmentation and sequencing. Importantly, this Docking & Deamination followed by sequencing (D&D-seq) technique is easily incorporated into common single-cell multiomics workflows, allowing for multimodal analysis of gene regulation in single cells. We demonstrate the ability of D&D-seq to precisely profile CTCF and GATA1 binding in bulk as well as single-cell analyses of both cell lines and primary human cells. We further integrated D&D-seq with single-cell genotyping to assess the effect of IDH2 mutations on CTCF binding in human clonal hematopoiesis, identifying reduced CTCF binding in mutant cells. Together, the ability to directly measure TF or chromatin remodeler binding in vivo using high-throughput single-cell sequencing platforms empower novel discoveries relating to chromatin and transcriptional regulation in human cells across physiological and disease contexts, at unprecedented scale and resolution.
[0349] CUT&Tag-derived approaches, including NTT-seq and scCUT&Tag-Pro, are suitable for profiling stable and abundant chromatin binders, such as histones. These methods are based on Cleavage Under Targets and Tagmentation (CUT&Tag), an approach that generates a library from DNA fragments colocalizing with a chromatin factor of interest by fragmenting out DNA surrounding the antibody-bound chromatin factor. This process is facilitated by pA-Tn5 fusion protein, where protein A (pA) has an affinity to antibodies and Tn5 is a transposase that fragments and tags DNA with sequencing-ready adapters. As Tn5 binds to DNA with high affinity, pA-Tn5 staining and tagmentation are performed under high stringency conditions (e.g., 300-500mM NaCl) to prevent signal from open chromatin regions mediated by direct binding of Tn5 to DNA. However, these stringent conditions also limit the antibody-tethered activity of the pA-Tn5, resulting in inefficient tagmentation at the chromatin factor-bound sites. In addition, these high salt conditions can abolish true biological interactions, such as those between DNA and some TFs, and also disrupt interactions between antibodies that have lower affinity for their targets. As such, these methods are limited in their ability to profile weaker or scarcer DNA:protein interactions, including those involving TFs. To address this challenge, we pioneered a new method for mapping TF binding in single cells. As an alternative approach to tagmentation plus barcoding of fragments at the genomic location bound by the target protein, we reasoned that we could capture binding patterns of target proteins by tethering nanobodies to a base editing enzyme. Briefly, this enzyme is the fusion of the DddA split base editor with secondary nanobodies that can recognize rabbit, mouse, goat and rat immunoglobulins. The enzyme is capable of binding to protein- specific antibodies to catalyze cytosine deamination in the vicinity of the targeted genomic area. This deamination event results in the conversion of cytosine to thymidine on the genomic DNA, which can subsequently be identified through sequencing, providing a molecular footprint of the target of interest on the genomic DNA. After the binding event is recorded with a OU transition, the cells are processed with the regular scATAC-seq workflow, enabling the capture of accessible regions of the genome (Fig. la). Notably, during deaminase incubation, we are able to maintain the enzyme in an inactive state, avoiding non-specific deamination resulting from non-tethered random interaction of the enzyme with the genomic DNA. To achieve a switch-like control over enzyme activity we use a split enzyme, where the DddA protein is separated into two polypeptides (Fig. lb). The activity of the enzyme can be controlled by the addition of the C-terminal small peptide to the reaction, which promotes the dimerization of the two domains to reconstitute the split enzyme and its activity. We call this approach Docking & Deamination followed by sequencing (D&D-seq) for the profiling of DNA:protein interactions in regions of open chromatin, amenable to tagmentation ATAC-seq, where the majority of TF binding occurs.
[0350] To design the fusion enzyme, we selected engineered DddAl 1 that has increased activity and more versatile DNA editing context compared to the original DddA. We chose the split site with the longest N terminus, and linked the nanobodies with the N terminal protein (nb-DddA_NT) so that the C terminal peptide (DddA_CT) is 25 amino acids (Fig. lb), short enough to be synthesized. We truncated the last 5 aa of DddA_CT because it is a structurally disordered region (PDB 6u08), and was shown to be functionally redundant by Yin et al..
[0351] To first evaluate enzyme base editing specificity, we tested the deaminase activity of DddA across nucleotide contexts. Consistent with previous studies, we observed that nb- DddA preferentially deaminates cytosines that are in the TpC context, with little activity in other contexts or at methylated Cs (Fig. 1c). We next tested recombinant nb-DddA_NT activity using a deaminase assay, showing that nb-DddA_NT is not active, while nb- DddA_NT became active after DddA_CT was added (Fig. Id). Furthermore, we showed that the addition of Zn2+ increases the deaminase activity (Fig. Id). To assess enzyme function under various experimental conditions, we also tested the activity of nb-DddA under different temperatures and duration, determining that the activity was great at 37 °C and 50 °C compared to 30 °C and 4 °C. Together, these results show that our nanobody- tethered and activatable DddA enzyme reconstitutes the activity of the unmodified enzyme with high fidelity.
[0352] To evaluate nbDddA activity in vivo, we next tested nb-DddA in a bulk cell D&D- seq experiment targeting either CTCF, GATA1 or GATA2 transcription factors in K562 cells. We identified a greatly increased C>T and G>A signature in DNA compared to other mutations and compared to a no-deaminase ATAC control. To confirm that the deaminase signal is TF-site-specific and not random, we performed motif enrichment analysis using peaks called from reads with C>T or G>A mutations. Motif enrichment analysis by MEME or HOMER showed strong CTCF, GATA1 or GATA1 motif in the peaks (Fig. le-j). A deamination footprint analysis of D&D events in a 200 base pairs window around all the ENCODE-defined CTCF, GATA1 or GATA2 binding sites in the genome showed the expected bimodal distribution of C deamination events, where the binding sites for the targeted TF are at the center of the two modes, with low editing frequency in background accessible regions lacking binding sites (Fig. lf,hj). Furthermore, we used ENCODE CTCF or GATA1 ChlP-seq dataset in K562 cells to define on-target peaks and off-target peaks, and identified much higher deaminase activity in on-target peaks compared to the off-target peaks or the background (Fig. lk,l). These results support the efficient activity of nb-DddA in vivo and provide evidence that TF binding can be profiled with high specificity in vivo.
[0353] Example 3: Single-cell TF profiling with D&D-seq
[0354] For single-cell D&D-seq, isolated cells are first fixed and permeabilized to allow all the reagents to enter the cell. Then, the sample is incubated with an antibody specific to the DNA binding target of interest. After washing the antibody in excess, the sample is incubated with the D&D-deaminase specific to the antibody used to label the target protein. The nanobody domain of the D&D-deaminase specifically and stably binds with the antibody, localizing the enzyme in close proximity with the target. After washing the D&D- deaminase in excess, the sample can be processed with standard single-cell ATAC-seq methods. In ATAC-seq, genomic DNA is exposed to Tn5, a highly active transposase. Tn5 simultaneously fragments DNA, preferentially inserts into open chromatin sites, and adds sequencing primers (tagmentation). Open chromatin is identified from the sequenced DNA and data analysis can provide insight into gene regulation. In this way, we are able to profile accessible chromatin and at the same time record the protein binding events on the genomic DNA (Fig. la).
[0355] We next tested the ability of D&D-seq to record TF binding in single cells when incorporated into the lOx Genomics scATAC-seq workflow. We performed mixing experiments in cell lines, selecting two distinct cell lines (lymphocyte cell line CA46 and erythroblast cell line K562) and two TF targets (CTCF and GATA1) to allow us to assess the targeting specificity and potential signal contamination of D&D-seq. We profiled a total of n = 1,732 K562 cells and n = 4,108 CA46 cells (Fig. 2a), obtaining on average 5,437 fragments for CA46 and 5,263 fragments for K562 per cell (Fig. 2b), reflecting high quality tagmentation data. Cell lines were clearly separated by gene accessibility score of known marker genes (Fig. 2c). We projected cells into a low-dimensional space using latent semantic indexing (LSI) and uniform manifold approximation and projection (UMAP) using Seurat, and clustered cells using a weighted combination of the data modalities. There were two clear clusters corresponding to the two cell types showing good separation between K562 and CA46 cells, as defined by highly accessible markers for erythroid (Fig. 2d) and lymphoid cells (Fig. 2e), demonstrating that the D&D-seq reaction does not affect the overall performance of the sc AT AC.
[0356] Importantly, at the single-cell level, we were able to map protein binding events in all the cells analyzed, suggesting that the deamination reaction has a relatively high efficiency. Moreover, since the antibody staining step was performed before mixing the two cell lines, we were able to assess the presence of cross contamination potentially occurring during droplet encapsulation, barcoding, and library preparation reactions (Fig. 2f-i). Our results showed that the signals for CTCF and GATA1 are mutually exclusive, and are retained exclusively in the subpopulation stained with the respective specific antibody, suggesting that during deamination and ATAC-seq steps, the specificity of the signal is maintained at the cellular level.
[0357] To provide support for the specificity of D&D-seq, we performed de novo motif discovery analysis on the pseudo-bulk D&D-seq mutational peaks (Fig. 2j-m). We observed significant enrichment of the CTCF binding site (e-value 1.3e-171) in CA46 cells (Fig. 2j) and the GATA binding site (e-value 2.1e-89) in K562 cells (Fig. 21). We next orthogonally validated our findings using a motif-centric, rather than de novo, approach. Here, we used the total reads obtained from the single-cell experiment to calculate the frequency of D&D events in a 200 base pairs window around all the ENCODE-defined CTCF or GATA binding sites in the genome, and compared this with the frequency of C to T events observed in all the accessibility peaks measured in our experiment. This analysis revealed the presence of the expected bimodal distribution of C deamination events, where the binding sites for the protein of interest represent the center of the two modes (Fig. 2k, m, top). However, when the same analysis was performed for accessible regions lacking binding sites, we observed negligible edit event frequency and the absence of particular distribution patterns (Fig. 2k, m, bottom). These results demonstrate across two TFs that our D&D-seq approach faithfully recapitulates TF binding at the single-cell level. We next sought to evaluate the specificity of the deamination reaction at the site of target protein binding. As such, we collapsed all the reads assigned to a specific cluster, obtaining pseudo-bulk alignment map files. From these cell type-specific alignment files, we extracted all the reads where we identified C to U conversions and generated a D&D coverage plot. The peaks obtained with our in situ footprinting approach showed high concordance with bulk ChlP-seq reference data obtained from the ENCODE consortium (Fig. 2n,o), suggesting that base editing events occur only in the proximity of the targets of interest, with little evidence of off target activity. We quantified the similarity between our single-cell D&D-seq reads and those from ENCODE, showing high correlation between CTCF signals from each data type and between GATA1 signals for each data type, but no correlation between the CTCF and GATA1 signals (Fig. 2p). These data confirm the specificity of the scD&D-seq signal, indicating highly accurate recovery of TF binding in single cells.
[0358] Example 4: Single-cell in vivo TF profiles in primary human PBMCs
[0359] We tested our workflow for analysis of human samples in peripheral blood mononuclear cells (PBMCs). We selected CTCF as an initial target due to its ubiquitous presence on the genome, the high stability of the binding and the presence of a defined consensus sequence. For this experiment, immobilized PBMCs collected from healthy donors were crosslinked, permeabilized and stained with a CTCF-specific polyclonal antibody. Cells were then washed and incubated in the D&D-seq buffer that provides an activating peptide and the divalent cofactor necessary for the activation of the deamination reaction by the base editing enzyme, in order to record the presence of CTCF on the genomic DNA. After this step, the sample was processed with the lOx Genomics scATAC- seq kit according to the manufacturer instructions.
[0360] In this analysis, we obtained 5,413 high quality single cells, with an average number of fragments of 15,557 per cell. We projected cells into a low-dimensional space using LSI and UMAP, clustering cells using the chromatin accessibility data. We assigned cell states based on chromatin accessibility profiles of established marker genes and retrieved all the expected subpopulations for this tissue, demonstrating that the addition of the D&D reaction does not impact the overall quality of the chromatin accessibility assay. More importantly, we successfully mapped CTCF binding sites in this complex primary tissue. We extracted the reads with in situ C-to-U transition labeling and identified CTCF binding sites in close proximity to C to U deamination events. Example 5: Linking genotype to TF profiles and chromatin landscapes in single cells
[0361] In order to build a comprehensive toolkit to capture gene regulation in aging single cells, we reasoned that incorporating genotyping capability would enable the extension of D&D-seq to the study of somatic mosaicism. There is increasing recognition that tissues throughout the body contain many thousands of clonal expansions fueled by somatic driver mutations that confer a growth advantage, observed even in seemingly normal tissue at increasing levels with age. Indeed, recent data revealed that clonal expansions are nearly ubiquitous in hematopoietic stem cells of older individuals, resulting in highly prevalent CHIP by the age of 70. Modeling this human somatic evolution is challenging, as cell culture or murine models may not accurately reflect the evolutionary processes that occur in humans over decades. As such, primary human samples are critical for determining mechanisms of clonal outgrowth in human clonal mosaicism, which remain largely unknown. Given the mixture of mutated and wild-type cells within clonally mosaic tissues, single-cell multi-modal methods that capture both genotype and phenotype together are required to provide biological insights into mechanisms of clonal outgrowth in aging tissue. To address this challenge, our group previously developed Genotyping of Transcriptomes (GoT) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA) to enable high-throughput single-cell genotyping together with transcriptome profiling or chromatin accessibility profiles, respectively. We have used these methods for single-cell genotyping- to-phenotype mapping of non-malignant clonal hematopoiesis, uncovering the mechanism for how global disruption to DNA methylation, splicing or signal transduction pathways leads to cellular differentiation skews in mutant cells, directly in human samples. Leveraging this conceptual framework for linking single-cell genotype to molecular phenotype, we incorporated our D&D-seq approach into single-cell genotype-aware multi- omics, to connect somatic driver mutations to TF binding, and applied this workflow to profile human CHIP samples.
[0362] In order to develop genotype-aware D&D-seq to understand how somatic mutations can disrupt TF binding, adapted our GoT-ChA framework for single-cell chromatin accessibility profiling combined with somatic mutation genotyping to integrate into the D&D workflow. Briefly, in GoT-ChA, isolated nuclei are subjected to transposition of genomic DNA (gDNA) and loaded into microfluidics devices (40K cells needed for 10K target). During cell barcoding reactions, additional gene-specific primers are added to capture a locus of interest with an in-droplet PCR reaction, with a handle that is compatible with the lOx snATAC platform. The product is then split, with an aliquot used for an amplicon genotype library and the remaining 90% used for AT AC library construction. The ATAC and genotyping libraries can then be analytically integrated via shared cell barcodes, linking chromatin accessibility to targeted genotyping at the single-cell resolution. To incorporate genotyping into the D&D-seq framework, GoT-ChA primers are added during the ATAC step.
[0363] To test D&D-GoT-ChA, we profiled CTCF binding in PBMCs from a patient with CHIP carrying an IDH2R140Qmutation with a variant allele frequency of 0.15 identified by a targeted panel sequencing. Integration of GoT-ChA with D&D-seq allowed us to perform genotyping in individual cells, along with chromatin status and CTCF binding status. We first performed cell clustering based on chromatin accessibility data, agnostic to genotype, and projected genotypes onto the UMAP. We successfully genotyped 25.32% of single cells, identifying homozygous mutant, heterozygous and homozygous wild-type cells. There was uneven distribution of mutant cells across cell types, with most of the IDH2 mutant cells concentrated in the CD8 T cell cluster. We next assessed the CTCF binding by comparing the CTCF D&D signals between IDH2 wild type (WT), mutant (IDH2R140Q, MUT), and heterozygous (HET). Notably, we observed that CTCF binding signal was significantly decreased in mutant cells, a finding that is consistent with the decreased CTCF binding that has been reported in IDH-mutant glioma and acute myeloid leukemia, mediated through DNA hypermethylation at CTCF binding sites.
[0364] Additional data are shown in Fig. 11-Fig. 38.
[0365] Example 6: Evaluation of DnD-mediated genome edits
[0366] To compare DnD edit counts between target and background peaks, normalized edit counts and signal-to-noise ratio (SNR) were calculated (Fig. 3). First, SNVs counted in target and background peak groups were normalized by the total number of peaks and peak size and multiplied by a scale factor 100 to calculate the number of edits per 100 bp per peak. SNR was then calculated by dividing the DnD edit counts by the mean of non-DnD edit counts. Footprint analysis for DnD edits was performed by counting the number of DnD edits in each position from randomly sampled 200 target and background peaks. The random sampling was repeated for 10 times and mean and standard deviation were used for visualization. Example 7
[0367] Al. Gene expression programs are fundamentally shaped by highly coordinated epigenetic processes that are disrupted during aging. Regulatory control orchestrated by chromatin-binding factors, including transcription factors (TFs) and chromatin remodelers, underlies the gene expression programs responsible for maintaining cell identity, executing cellular functions and responding to environmental stimuli. These DNA:protein interactions are directed through epigenetic features1, such as histone modifications and DNA methylation (DNAme), which establish chromatin landscapes, modulating the binding of specific factors, and thereby sculpting the functional genome according to the needs of the cell. Importantly, epigenetic dysregulation has been linked to cellular dysfunction in aging, where aberrant chromatin landscapes transform the binding landscapes of TFs2, altering the normal biological processes of the cell3. Indeed, epigenetic alteration has been named as one of the hallmarks of aging4, and changing TF binding patterns have been identified as key mediators of cellular aging5. As such, understanding the complexity and dysregulation of the TF binding code across different cell types and cell states in older individuals remains a central challenge in aging biology.
[0368] A2. Methods for profiling TF binding or transient DNA: protein interactions in vivo are mostly limited to bulk analysis. While transcriptional readout through RNAseq identifies the active genes and pathways that contribute to cellular responses or phenotypes, understanding the regulatory factors that direct these expression programs is crucial for both biological insights as well as for the identification of factors that are dysregulated during aging. There are over 1,600 likely TFs encoded in the human genome, with DNA binding motif sequence information, obtained from various in vitro, bulk or bioinformatic analyses, established for -1,100 of these6. However, DNA sequence motif alone is insufficient for predicting TF binding6, which preferentially occurs in open chromatin7and is additionally shaped by TF-TF8as well as TF-nucleosome9interactions. Thus, profiling TF binding in native chromatin contexts is crucial for unraveling the normal regulation, and age-associated dysregulation, of gene expression. Chromatin immunoprecipitation with sequencing (ChlP-seq) and cleavage under targets & tagmentation (CUT&TAG) are powerful approaches for mapping genome-wide interactions of target proteins with DNA. These methods identify genomic regions bound by an individual protein of interest through sequencing short DNA fragments that are cross linked or tagmented. In addition, DamID10’11molecular footprinting techniques harnessing the DNA methyltransferase activity have been used to map TF binding. However, due to their various protocol requirements, these methods cannot be easily incorporated into available high-throughput single-cell workflows, limiting application to bulk analysis, or in single cells to the profiling of only the most tightly bound chromatin factors, such as histones.
[0369] A3. Single-cell methods to profile TF binding are needed to understand gene regulatory networks in aging cells. Bulk analysis can mask important biological signal that results in cell-to-cell heterogeneity. As such, single-cell analysis is required to define the regulatory landscape of the genome with sufficient precision to obtain functional insights into the complex gene expression networks. Moreover, to capture physiological effects of aging in human cells, single-cell methods that have the throughput and flexibility for application to primary human samples are needed. Finally, methods that can profile binding patterns of TFs at weaker affinity sites that could be aberrantly uncovered by epigenetic alterations in individual aging cells, or that profile less stable DNA interactors such as more transient chromatin remodelers, are crucial for resolving genomic regulatory networks as well as their resiliency or susceptibility to dysregulation. However, there are no existing single-cell methods that fulfill these criteria, creating a major gap in our ability to define TF binding or chromatin remodeling dynamics in individual primary cells, further hampering our ability to discover factors that are dysregulated in aging.
[0370] A4. Reconstructing genomic regulatory networks across modalities in single cells. Profiling of TF binding patterns in single cells in vivo has been mainly restricted to inferential approaches based on expression levels of key downstream TF target genes12 14or through motif analysis of ATAC-seq peaks15. While analysis of ATAC-seq peaks can be highly informative about the general chromatin landscape, for confident identification of specific TF binding sites, more direct methods are recommended16. Combining computational TF binding inference analysis with single-cell multiomics approaches that enable simultaneous capture of chromatin accessibility and gene expression within the same single cell allows for the characterization of genomic regulation across multiple molecular layers (from chromatin landscape to RNA). Incorporating direct TF binding measurements into such a single-cell framework would provide an unprecedented window into the exact mechanisms driving genome function, transcription networks and pathway regulation. Moreover, this multifaceted analysis spanning different molecular modalities as a readout of direct TF binding would accelerate the identification of key factors that mediate normal cellular activity, as well as those that are disrupted during aging. A5. Genotype-aware single-cell multiomics mechanistically profile somatic mutations that arise with aging. Cell-to-cell variability encompasses not only transcriptional and epigenetic variation, but also genetic variation. Specifically, somatic evolution introduces genetic diversity across tissues in humans during aging17 2(i. Notably, in age-related clonal hematopoiesis of indeterminant potential (CHIP), there is frequent selection of driver mutations21 23that pleio tropic ally disrupt DNAme, chromatin accessibility and histone modifications (for example, somatic mutations in DNMT3AU, TET2.~S' ' or ASXLl '2"), leading to dysregulation of gene expression across epigenetic layers. Understanding the phenotypic effects of these aging-associated somatic mutations requires single-cell multi-modality capture of genotype and phenotypic readouts. Our group has performed single-cell genotyping-to-phenotype mapping of CHIP, uncovering the mechanism for how global disruption to DNA methylation24,34or splicing35leads to cellular differentiation skews in mutant cells, directly in human samples. We further developed Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA) for high-throughput joint analysis of genotype and chromatin accessibility in single cells, which we used to uncover cell type- specific, cell-intrinsic inflammatory phenotypes reflecting chromatin landscape changes that were driven by JAK2 mutations in myelofibrosis patients15. Here, we inferred TF binding from motif enrichment at open chromatin sites, but to comprehensively understand the drivers of epigenetic rewiring and regulatory alterations in aging cells, we need to identify the mechanistic effects of somatic mutations on TF binding directly. As such, high-throughput, single-cell multi-omics methods that connect genotypes to chromatin accessibility as well as to regulatory readouts that capture actual DNA:protein interactions, including direct TF or chromatin remodeling complex binding, would provide unprecedented mechanistic insight into the molecular biology of aging.
[0371] A6. A single-cell epigenetics toolkit to define the aging genome regulatory landscape. To unravel genome regulation networks in aging cells, we need to accurately map DNA:protein interactions with single-cell precision and sufficient power for application to human samples. Here, we aim to develop a suite of tools that directly profile TF or chromatin factor binding along with chromatin accessibility in single cells, based on tethering a base-editing deaminase to nanobodies targeting a DNA binding protein of interest combined with single-cell ATAC-seq. Upon genome-wide antibody binding to the TF, the nanobody targets the tethered base editor enzyme to the antibody, resulting in deamination of cytosine to uracil at regions of TF binding, leaving a permanent genomic signature that can be identified through downstream single-cell sequencing of the tagmented loci. Facile incorporation into existing droplet-based single-cell sequencing workflows further allows for the addition of other molecular modalities onto this backbone technology, which we call D&D-seq (Docking and Deamination followed by sequencing). To provide a comprehensive toolkit for defining the epigenetic features that program genome regulation and age-related dysregulation (Fig. 4), we first develop and optimize D&D-seq for the single-cell interrogation of chromatin landscapes and direct TF binding that is additionally compatible with methods that allow for the simultaneous measurement of downstream (e.g., gene expression, protein levels) as well as upstream (e.g., histone modifications, genotype) effectors, in primary human cells, and additionally incorporate multiplex capability to measure binding of multiple TFs in the same cell (Aim 1). We then combine D&D-seq with our method for single-cell profiling of histone modifications36, to define how TF or chromatin remodeler binding changes as a function of active chromatin states regulated through histone marks, as well as to directly profile repressive DNA binding factors located in inaccessible regions (Aim 2). To capture the downstream effects of regulatory changes in chromatin landscape and TF binding on gene expression, we combine D&D-seq with multimodal single-cell technologies to profile genomic regulation across chromatin and RNA layers, allowing us to reconstruct complete signaling networks with a functional readout for TF binding. Finally, to mechanistically interrogate the consequences of age-associated somatic mutations on TF binding and chromatin architecture, we harness our advances in single-cell genotype-aware multi-omics methods15for integration with D&D-seq to provide a new technology for defining the functional effects of somatic variation on genomic regulation (Aim 3). Together, these powerful new methods provide an unprecedented window into genomic and transcriptomic regulation at the single-cell level, with in vivo TF binding maps generated from single cells allowing researchers to ask questions and perform analyses that were impossible before with existing approaches, including how TF binding changes in across cell types during aging, what epigenetic features overlap with TF binding, how binding motifs are differentially utilized in different cellular contexts and how chromatin landscapes and TF binding interplay to effect transcription level.
[0372] Here we propose to develop transformative methods for application to the study of human aging at the single-cell level, delivering comprehensive molecular profiles that collectively define regulatory networks in individual cells. Specifically, the ability to directly measure TF binding in vivo using high-throughput single-cell sequencing platforms empower novel discoveries relating to chromatin and transcriptional regulation, as well as dysregulation driven by stresses such as aging, at unprecedented scale and resolution.
[0373] Moreover, our project is conceptually innovative, providing a framework for the analysis of upstream and downstream mediators of chromatin regulation through TF binding, unifying within the same single cell signal from chromatin accessibility, TF binding and histone modification or gene expression, as well as genotype. Thus, our tools can be used in order to holistically define the functional state of a cell, through producing composite profiles that go beyond single modalities (e.g., transcriptome). These advances are propelled by significant technological innovation, which allows us to combine multiple single-cell profiling modalities together for the first time. These technologies are designed to have broad applicability across human samples, imparting the ability to comprehensively produce in vivo TF binding maps from cells throughout the body in physiologically relevant contexts. Currently, there are no existing methods that provide high-throughput single-cell TF binding profiles, or that additionally combine multimodal or genotype-aware profiling, for analysis of primary samples. Our tools that we aim to develop here thus fill a major gap in the field of molecular aging, with the potential to unravel complex regulatory networks of cells and how they aberrantly change in aging cells, delivering:
[0374] • A foundational technology harnessing a novel nanobody -base editor fusion combined with ATAC-seq to map DNA:protein interactions along with chromatin accessibility in single cells (D&D-seq). This highly innovative approach is (i) amenable to incorporation with other single-cell multiomics sequencing techniques, (ii) can be applied at high throughput for the analysis of primary human samples and (iii) can be adapted for multiplexed profiling of more than one TF or chromatin factor. This allows for the analysis of single-cell in vivo TF binding in primary human tissue with functional molecular readouts for the first time.
[0375] • Cutting-edge single-cell multi-omics assays to link TF binding patterns and chromatin accessibility landscapes with histone marks, gene expression and / or protein levels in the same cells, to explore how chromatin state regulation impacts TF binding, and in turn how changes in TF binding affect transcriptional output. These methods enable mapping the effects of TF binding to comprehensively define the regulatory networks that constitute normal or aberrant cell states across an unprecedented number of molecular features. • Technologies that combine single-cell genotyping with chromatin landscape and TF binding profiling, allowing for single-cell genotype-to-phenotype mapping to mechanistically decipher the effect of somatic mutations on TF binding and chromatin accessibility. Thus, the interplay between genetic and epigenetic variation, and its impact on the regulation of gene expression can be determined in human cells.
[0376] Mapping TF binding in single-cells with D&D-seq. We recently developed Nanobody Tethered Tagmentation followed by sequencing (NTT-seq36) for multimodal single-cell epigenetic profiling to better understand epigenetic remodeling dynamics in human primary cells at single-cell resolution. This method is based on Cleavage Under Targets and Tagmentation (CUT&Tag37), an approach that generates a library from DNA fragments colocalizing with a chromatin factor of interest by fragmenting out DNA surrounding the antibody-bound chromatin factor. This process is facilitated by pA-Tn5 fusion protein, where protein A (pA) has an affinity to antibodies and Tn5 is a transposase that fragments and tags DNA with sequencing-ready adapters. As Tn5 binds to DNA with high affinity, pA-Tn5 staining and tagmentation are performed under high stringency conditions (e.g., 300-500mM NaCl) to prevent signal from open chromatin regions mediated by direct binding of Tn5 to DNA. However, these stringent conditions also limit the antibody-tethered activity of the pA-Tn5, resulting in inefficient tagmentation at the chromatin factor-bound sites. In addition, these high salt conditions can abolish true biological interactions, such as those between DNA and some TFs, and also disrupt interactions between antibodies that have lower affinity for their targets. As such, CUT&Tag-derived approaches, including NTT-seq and scCUT&Tag-Pro38, are suitable for profiling stable and abundant chromatin binders, such as histones. However, these methods are limited in their ability to profile weaker or scarcer DNA:protein interactions, including those involving TFs.
[0377] To address this challenge, we pioneered an entirely new method for mapping TF binding in single cells. As an alternative approach to tagmentation plus barcoding of fragments at the genomic location bound by the target protein, we reasoned that we could capture binding patterns of target proteins by tethering nanobodies to a base editing enzyme. Briefly, this enzyme is the fusion of the DddA39,40split base editor with secondary nanobodies that can recognize rabbit, mouse, goat or rat immunoglobulins (Fig. 5). The enzyme is capable of binding to protein- specific antibodies to catalyze cytosine deamination in the vicinity of the targeted genomic area. This deamination event results in the conversion of cytosine to uracil on the genomic DNA, which can subsequently be identified through sequencing, providing a molecular footprint of the target of interest on the genomic DNA. Unlike the DamID approach10,11that requires additional DNAme sequencing with further enzymatic or chemical processing, TF footprinting using our base editing strategy only requires a genomic DNA sequencing readout, making it compatible with common single-cell genomics platforms. After the binding event is recorded with a C>U transition, the cells are processed with the regular scATAC-seq workflow, enabling the capture of accessible regions of the genome (Fig. 6a). Notably, during deaminase incubation, we are able to maintain the enzyme in an inactive state, avoiding non-specific deamination resulting from non-tethered random interaction of the enzyme with the genomic DNA (Fig. 5). To achieve a switch-like control over enzyme activity we use a split enzyme, where the DddA protein is separated into two polypeptides. The activity of the enzyme can be controlled by the addition of the C-terminal small peptide to the reaction, which promotes the dimerization of the two domains to reconstitute the split enzyme and its activity. We call this approach Docking & Deamination followed by sequencing (D&D- seq) for the profiling of DNA:protein interactions in regions of open chromatin, amenable to tagmentation ATAC-seq, where the majority of TF binding occurs41.
[0378] We first tested the ability of D&D-seq to record TF binding in single cells when incorporated into the lOx Genomics scATAC-seq workflow. We performed mixing experiments in cell lines, selecting two distinct cell lines (lymphocyte cell line CA46 and erythroblast cell line K562) and two TF targets (CTCF and GATA1) to allow us to assess the targeting specificity and potential signal contamination of D&D-seq. We profiled a total of n = 1,732 K562 cells and n = 4,108 CA46 cells (Fig. 6b), obtaining on average 5,437 fragments for CA46 and 5,263 fragments for K562 per cell (Fig. 6c), reflecting high quality tagmentation data. We projected cells into a low-dimensional space using latent semantic indexing (LSI) and uniform manifold approximation and projection (UMAP) using Seurat, and clustered cells using a weighted combination of the data modalities. There were two clear clusters corresponding to the two cell types showing good separation between K562 and CA46 cells (Fig. 6b), as defined by highly accessible markers for erythroid and lymphoid cells (Fig. 6d), demonstrating that the D&D-seq reaction does not affect the overall performance of the scATAC. We next sought to evaluate the specificity of the deamination reaction at the site of target protein binding. As such, we collapsed all the reads assigned to a specific cluster, obtaining pseudo-bulk alignment map files. From these cell type-specific alignment files, we extracted all the reads where we identified C to U conversions and generated a D&D coverage plot. The peaks obtained with our in situ footprinting approach showed high concordance with bulk ChlP-seq reference data obtained from the ENCODE consortium41(Fig. 6e), suggesting that base editing events occur only in the proximity of the targets of interest, with little evidence of off target activity.
[0379] To provide further support for the specificity of D&D-seq, we performed de novo motif discovery analysis on the pseudo-bulk D&D-seq mutational peaks (Fig. 6f-6i). We observed significant enrichment of the CTCF binding site (e-value 1.3e-171) in CA46 cells (Fig. 6f) and the GATA binding site (e-value 2.1e-89) in K562 cells (Fig. 6h). We next orthogonally validated our findings using a motif-centric, rather than de novo, approach. Here, we used the total reads obtained from the single-cell experiment to calculate the frequency of D&D events in a 200 base pairs window around all the ENCODE-defined41CTCF or GATA binding sites in the genome, and compared this with the frequency of C to U events observed in all the accessibility peaks measured in our experiment. This analysis revealed the presence of the expected bimodal distribution of C deamination events, where the binding sites for the protein of interest represent the center of the two modes (Fig. 6g, i, top). However, when the same analysis was performed for accessible regions lacking binding sites, we observed negligible edit event frequency and the absence of particular distribution patterns (Fig. 6g, i, bottom). Importantly, at the single-cell level, we were able to map protein binding events in all the cells analyzed, suggesting that the deamination reaction has a relatively high efficiency. Moreover, since the antibody staining step was performed before mixing the two cell lines, we were able to assess the presence of cross contamination potentially occurring during droplet encapsulation, barcoding, and library preparation reactions (Fig. 6j-m). Our results showed that the signals for CTCF and GATA1 are mutually exclusive, and are retained exclusively in the subpopulation stained with the respective specific antibody, suggesting that during deamination and ATAC-seq steps, the specificity of the signal is maintained at the cellular level.
[0380] Together, these preliminary results demonstrate that D&D-seq can be used to map the binding of key TFs with high specificity and good efficiency, paving the way for expanded applications to primary cells and incorporation of other single-cell molecular modalities that could be used to chart dysregulation of gene expression networks in single cells for mechanistic insight into molecular aging. DI. Overall goal. Here we aim to develop high-throughput multimodal analysis tools that map the direct binding of TFs or chromatin remodeling complexes in single cells in vivo. These advances fill a major gap in the field, as current approaches for accurately mapping single-cell TF binding are limited, and there are no existing methods that can be integrated with common high-throughput single-cell sequencing platforms. The simultaneous mapping of chromatin landscapes, TF binding, transcription and genotype in the same cell set the stage for new discoveries about the mechanisms of gene regulation in aging human cells. Our novel and broad-reaching technologies further empower the field to interrogate how genomic regulatory networks are altered during aging at unprecedented levels of resolution, with the potential to lead to important discoveries about the factors that contribute to cellular phenotypes observed with aging, potentially helping to shape possible intervention or treatment strategies for the mitigation of disease-associated effects. Ultimately, our method is able to be used generate in vivo single-cell binding maps for key TFs that regulate pathways that are activated or disrupted during aging (e.g., ROS response, mTOR signaling, autophagy or inflammation pathways)42,43across cell types and contexts, incorporating additional functional information about chromatin landscape state and downstream gene expression, as well as co-factor binding or genotype, to mechanistically define genomic regulation within single cells for the first time.
[0381] D2. Rigor and Reproducibility. To obtain robust and unbiased results: 1) Sequencing results are deposited in relevant NIH repositories (see Data Management Plan); 2) All protocols and biostatistical tools are provided open-source to enable replication; 3) Cell mixing experiments are used to determine the reproducibility of our D&D-seq method across tested targets. 4) Both sexes are equally represented in the analysis.
[0382] D3. Aim 1: Single-cell profiling of chromatin and transcription factor binding along with chromatin accessibility using D&D-seq. Rationale: Transcription factors are central coordinators controlling gene expression, with the binding of TFs to regulatory genomic regions representing an essential step required for transcription of most genes. TFs exhibit differential binding across cell types and cell states, where multiple factors influence the strength and location of TF binding, including DNA sequence, chromatin accessibility and binding of other TFs, which can be dysregulated in aging cells. Given this complexity, it is crucial to examine TF binding in single cells in vivo, from human samples, in order to mechanistically understand transcription regulation during aging. Here, we first optimize D&D-seq for application to primary human samples, ensuring that there is sufficient scale and power to confidently map TF binding together with chromatin accessibility from single cells. We then expand the D&D-seq framework to optimize multiplexed D&D-seq to simultaneously profile multiple TFs, in order to assess the effects of co-factor or competitor binding on TFs in single cells.
[0383] Aim la. Expand D&D-seq to primary human cells. Our preliminary work using D&D-seq has shown the feasibility of the footprinting approach in cell lines (Fig. 6). Here we aim to expand D&D-seq for application to primary human samples, for high-throughput profiling of TF binding and chromatin accessibility in single cells. Such capability is crucial for the study of physiological conditions, including aging, which may be poorly modeled with cell lines. Moreover, to broaden our single-cell toolkit for the interrogation of genome regulation at the level of TF binding and chromatin landscape in different cell types, we aim to demonstrate that our method is able to profile a range of target TFs.
[0384] In proof-of-principle experiments, we tested our workflow for analysis of human samples in peripheral blood mononuclear cells (PBMCs). We selected CTCF as an initial target due to its ubiquitous presence on the genome, the high stability of the binding and the presence of a defined consensus sequence44. For this experiment, immobilized PBMCs collected from healthy donors were crosslinked, permeabilized and stained with a CTCF- specific polyclonal antibody. Cells were then washed and incubated in the D&D-seq buffer that provides an activating peptide and the divalent cofactor necessary for the activation of the deamination reaction by the base editing enzyme, in order to record the presence of CTCF on the genomic DNA. After this step, the sample was processed with the lOx Genomics scATAC-seq kit according to the manufacturer instructions.
[0385] In this analysis, we obtained 5,413 high quality single cells, with an average number of fragments of 15,557 per cell (Fig. 7a). We projected cells into a low-dimensional space using LSI and UMAP, clustering cells using the chromatin accessibility data (Fig. 7b). We assigned cell states based on chromatin accessibility profiles of established marker genes and retrieved all the expected subpopulations for this tissue, demonstrating that the addition of the D&D reaction does not impact the overall quality of the chromatin accessibility assay. More importantly, we successfully mapped CTCF binding sites in this complex primary tissue. We extracted the reads with in situ C-to-U transition labeling and identified CTCF binding sites in close proximity to C to U deamination events (Fig. 7c).
[0386] Encouraged by these results, we now aim to build upon our preliminary analysis to expand the validation of our D&D workflow to different DNA-binding proteins, including other key transcription factors (e.g., GATA1) and chromatin remodeling factors (e.g., SWI / SNF members such as BRG1), to establish an adaptable protocol for scalable application to primary human samples. We profile a range of TFs with different binding affinities establish the optimal protocols for profiling high-affinity versus low-affinity TFs using human PBMC samples. We identify D&D reads (C to U changes) to map TF binding and perform de novo motif discovery analysis as well as analysis of the canonical motif for each TF, comparing the binding profiles to published bulk ChlP-seq data from matched cell types41. We use the PBMC ATAC-seq data to generate pseudotime hematopoietic differentiation trajectories, which are disrupted in aging45, and quantify how the binding of GATA1, a key erythroid lineage factor, changes over pseudotime, as performed previously for single-cell histone profiling36, for biological validation of our TF binding maps. The approach developed here can eventually be used to help unravel the TF binding landscape across different cells, allowing for comparisons of molecular aging across primary tissues.
[0387] Aim lb. Expand D&D-seq to capture cell surface protein levels and to generate binding maps for multiple DNA-binding proteins. For increased power to distinguish cell types that allows for better cell type mapping to guide analysis, we take advantage of our adaptable systems and further integrate a protein modality into the D&D-seq framework through a barcoded antibody-derived tag (ADT) approach from ASAP-seq46that is an adaptation of CITE-seq (cellular indexing of transcriptomes and epitopes by sequencing)47. Here, we perform cell permeabilization and deamination prior to ASAP-seq processing, which combines ATAC-seq with CITE-seq. We perform cell mixing experiments using our established lymphocyte cell line CA46 and erythroblast cell line K562, profiling CTCF and GATA1 with D&D-seq, and profile an established panel of hematopoietic ADTs46. We assess quality metrics for each modality, and optimize the protocol for maximum D&D read recovery while maintaining sufficient ADT counts to distinguish cell types. We complement these wet lab advances with optimized computational pipelines for bioinformatic analysis that combines these molecular layers, leveraging our experience in single-cell multiomics15,48.
[0388] Individual TFs have can have different dependencies on specific combinations of cofactors that can affect binding affinity and function49. To enable interrogation of cofactor influence on in vivo TF binding in single cells, we develop multiplexed D&D-seq through addition of another distinct base editing enzyme into the workflow. Briefly, we create a fusion enzyme consisting of adenine base editor Tad8e50with enzymatically dead DddA, tethered to a nanobody that recognizes a different species from the cytosine base editing DddA nanobody tethered enzyme (Fig. 8). In this workflow, cells are stained with two primary antibodies targeting two TFs or chromatin factors of interest, each from distinct species (e.g., rabbit and mouse). After washing, the cells are incubated with two D&D- deaminases (cytosine base editor DddA or adenine base editor Tad8e fused to catalytically dead DddA) with tethered nanobodies that are specific to the antibody used to label the target proteins. The addition of the activating peptide starts the footprinting reaction, resulting in cytosine to uracil or adenine to inosine edits in the sequencing reads. Cells then undergo scATAC-seq (or can be incorporated with other single-cell sequencing platforms). The OU and A>I edits are then called for each of the TF binding sites. Building off our preliminary data, we optimize our protocol for multiplexed D&D-seq by profiling both GATA1 and CTCF binding together in erythroblast cell line K562, for which we already have single TF binding (GATA1; Fig. 6) profiles. This allows us to assess the efficiency of the multiplexed protocol in vivo, and optimize our analytical pipeline for generating in vivo single-cell TF binding maps for multiple factors. We then use this system to assess cofactor binding, profiling GATA1 and TALI51in K562 cells, optimizing our bioinformatic pipeline for the analysis of TF-TF binding, comparing the binding profiles of single D&D- seq experiments for each TF to the multiplexed experiment profiling both. In further benchmarking endeavors, we profile different components of the BAF chromatin remodeling complex52(e.g., SMARCB1 and ARID1A versus SMARCB1 and ARID1B), to evaluate the ability of multiplexed D&D-seq to resolve multi-subunit chromatin-binding complexes, and distinguish the in vivo binding profiles of functionally distinct complexes in single-cells.
[0389] Statistical considerations. All statistical analyses are performed with R (version 4.2.1). Categorical variables are compared using the Pearson Chi-square test or Fisher Exact test as appropriate, and continuous variables are compared using non-parametric tests for single-cell analyses. Multivariable analysis is performed with ANOVA and generalized linear models, and linear mixed models to control for potential confounders (random effects from subjects).
[0390] Pitfalls, alternatives and expected outcomes. Although binding matrix sparsity, a common problem in single-cell approaches, could be an issue, for each individual cell we were able to map from 10 to 75 binding events with D&D-seq in initial experiments. We anticipate that protocol optimization and more refined bioinformatic analysis further improve the efficiency of the workflow, potentially enabling the simultaneous mapping of hundreds of protein binding events along with chromatin accessibility in thousands of primary cells. There could be variability between different antibodies to recognize target proteins bound to chromatin. For each target TF or chromatin remodeling factor, we test different antibodies to identify those with optimal performance in the D&D-seq workflow. We expect that these optimized technologies are suitable for successful application to different human tissues and enable exploration of the impact of aging on genomic regulatory networks in cells directly from clinical samples, which can be collected and processed according to protocols that are compatible with single-cell analysis, allowing broad application to a high number of cells from diverse tissues. This generate in vivo single-cell binding maps for TFs (e.g., AP-1 in aging lymphocytes53; FOXO TFs across different aging tissues54), which can be tracked across pseudotime trajectories reconstructed based on single-cell ATAC data in aged samples. In vivo TF maps further help to unravel the dysregulation of age-related pathways, including DNA damage response, inflammation, senescence and autophagy. Multiplexed D&D-seq could be leveraged to profile the regulation of self-renewal and quiescent pathways in aged hematopoietic stem cells, for example through co-mapping of YY1 and cohesion complex factor SMC355. Further mechanistic insights can be obtained through de novo motif discovery across cell types and contexts, assessment of the relationship between TF binding and open chromatin, as well as dissection of co-factor or multi- subunit binding for TF signaling networks and chromatin remodeling factors, to fundamentally define the rules of gene regulation in aging cells.
[0391] Performance Measures and milestones for Aim 1. For application to primary human samples, based on barcode quality control, we aim to obtain 5,000 high-quality cells per experiment, with a minimum cell-specific ATAC peak transcription start site (TSS) enrichment score of 5 and a minimum number of unique fragments of 3,500 to ensure sufficient library complexity. We aim to optimize the protocol to obtain >500 D&D counts genome- wide for high-occupancy TFs (e.g., CTCF) from primary cells. As there are no available single-cell TF binding profile methods that can be combined with the lOx Genomics platform, for benchmarking, all D&D-seq peaks are pseudobulked and compared to available bulk ENCODE ChlP-seq data41for the DNA-binding protein of interest, as in Fig. 3E. We benchmark D&D genomic deamination footprinting performance for each target protein assayed against the non-targeted IgG control that should not induce base editing. For multiplexed experiments, we compare the single TF binding profiles to the multiplexed binding profiles, aiming to attain the same number of D&D reads in the multiplexed experiments as for the single D&D-seq context.
[0392] D5. Aim 2: Develop a method for simultaneously mapping TF binding and histone marks to assess how active or repressive chromatin states influence TF binding in single cells. Histone modifications regulate chromatin states and transcription, which are disrupted in aging2,4. Profiling histone marks in single cells can reveal the presence of poised and active enhancers, promoters, repressed and heterochromatic regions. Importantly, chromatin remodelers act to modify chromatin and histone modification to allow TFs to reach their destinations on the DNA strand. Understanding the processes of these cellular construction workers in single cells could help shed mechanistic light on the dysregulation of chromatin landscapes in aging cells. Currently, there are no methods that allow for the single-cell mapping of TF or chromatin remodeler binding together with histone profiling, resulting in a lack of knowledge about the interplay between chromatin states and TF activity in normal or aging cells. Here, we combine our D&D-seq method with our NTT-seq approach36in order to simultaneously assess TF binding and chromatin states defined by histone marks for insights into how epigenetic states impact TF binding in aging single cells. We leverage D&D-seq genomic footprinting to profile DNA-binding proteins in regions of heterochromatin, substituting the AT AC-based tagmentation with NTT-seq to profile repressive chromatin marks, providing an unprecedented window into pioneer factor activity, as well as the activity of chromatin remodeling complexes in active and repressed chromatin regions, in aging cells.
[0393] Aim 2a. Leverage D&D to map TF binding within heterochromatin in single cells, further incorporated with NNT-seq histone binding profiles. Our D&D-seq approach that combines molecular footprinting with tagmentation via AT AC is well- suited for generating in vivo single-cell TF binding maps across cell types and tissues, but is limited assessing TF binding to regions of open chromatin. As it is known that most TF binding for both activators and repressors occurs in regions of open chromatin7,41, D&D-seq can capture most TF:DNA interactions. However, there is also important regulatory binding of TFs in heterochromatin regions, where pioneer factors7,56can initialize the opening of chromatin and subsequent binding of downstream regulatory factors. To overcome this limitation, we speculate that by combining the D&D reaction with our NTT-seq36technology for profiling histone modifications in single cells, we are able to first record the position of repressive protein targets of interest on the genome by deamination, and then drive the tagmentation on repressed chromatin through profiling H3K27me3 with NTT-seq, allowing us to expand our technology to map heterochromatic DNA binding proteins, including pioneer factors. Thus, to enable the mapping of TF binding across genomic regions marked by either active or repressive histone marks, we propose to adapt our D&D workflow in combination with NTT-seq36to provide a detailed map of TFs or chromatin remodeling factor occupancy across the genome with accompanying histone post- translational modification profiles in the same single cell. We test the feasibility of simultaneous mapping of TFs or chromatin remodeling factors and histone post- translational modification at single-cell resolution by performing cell mixing experiments, replacing the accessible chromatin tagmentation reaction (as outlined in Fig. 6) with nanobody-tethered tagmentation. Briefly, isolated cells are fixed and permeabilized to allow in situ reactions. Then, the sample is incubated with an antibody specific to the DNA binding target of interest in and antibodies against histone modifications. After washing the antibody in excess, the sample is incubated with the D&D-deaminase specific to the antibody used to label the target protein. The addition of the activating peptide start the footprinting reaction. At this point, the buffer is exchanged with a series of washes to increase the stringency of the reaction and allow the incubation with nb-Tn5. This stringency avoid the binding of Tn5 to accessible regions and selectively tagmenting the genomic regions stained with histone-specific antibodies. Our previous work using NTT- seq shows that we can retrieve high-quality single-cell multimodal chromatin profiling by processing the post-tagmented cell with the lOx Genomics scATAC kit. A weighted combination of the two histone modalities is used to cluster the cells in the two subpopulations based on chromatin features specific to the cell of origin. The cell line incubated with the D&D negative IgG control should present very minimal C to U transition, representing spontaneous deamination events or PCR errors.
[0394] For our optimization of our protocol, we perform cell-mixing experiments where we simultaneously profile the binding of GATA1 in erythroblast K562 cells and lymphoblast CA46 cells. K562 cells are stained with the target- specific cocktail, while CA46 cells are stained with non-targeting IgG and used as a negative control to model signal background and to validate the specificity of the D&D reaction. We profile H3K27me3 for inactive and H3K27Ac for active chromatin. We optimize conditions to maximize recovery of reads for each modality, and develop bioinformatic pipelines that intersect TF binding reads (base edited regions) with specific histone modifications (NTT signal) in single cells. We further test how active and repressive marks affect the binding of TFs that have been optimized in Aim 1, as well as test pioneer factor binding (e.g., OCT4, PU.1). After cell line optimization, we apply D&D + NTT-seq to PBMCs, and analyze hematopoietic lineage TF binding together with H3K27me3 and H3K27Ac, to establish a system for the interrogation of TF binding in hematopoiesis during aging. We annotate cell clusters using label transfer with annotated scATAC-seq datasets57,58, using the H3K27ac reads with manual annotation of active and repressive histone marks at key marker genes for each cell type, as we have done previously36(Fig. 9). We then compare the binding of early (e.g., H0XB4) and late (e.g., XBP1) hematopoietic lineage TFs across genomic bins over pseudotime, to dynamically compare binding patterns in the context of open and repressed chromatin across hematopoietic differentiation, to optimize the protocol and bioinformatic analyses for future application of mapping TFs in the aging blood, which has disrupted hematopoiesis45.
[0395] Aim 2b. Combine D&D-seq with NTT-seq to profile chromatin remodelers and histone marks. Several chromatin remodeling factors function as repressors, with the ability to induce chromatin compaction to shut down the transcription of nearby genes. These factors, such as polycomb repressive complex factors (e.g., PRC259), are located in regions of the genome that are not able to be assayed with single-cell ATAC-based strategies due to being inaccessible and refractory to tagmentation. To assess the kinetics of chromatin remodeling, we profile SWI / SNF complex members (ARIDA1, BRG1) in the context of both active and repressive marks. We also profile repressive complexes (PRC1 and PRC2) along with the specific histone marks that are catalyzed by that complex. Specifically, as in Aim 2a, we perform cell-mixing experiments for protocol optimization, initially targeting D&D nanobodies to EZH2 (member of the PRC2 complex that has reported to be somatically mutated in clonal hematopoiesis and some cancers60), marking the genomic region with C to U edits, followed by NTT-seq profiling of H3K27me3 to drive tagmentation of repressive regions. We further test the PRC1 complex (BMI1), followed by NTT-seq profiling of H2AK119ubl; and BRG1, followed by NTT-seq profiling of H3K27ac and H3K27me3. As before, we stain K562 cells with target- specific antibodies, while negative-control CA46 cells are stained with non-targeting IgG to validate the specificity of the D&D reaction. We optimize our bioinformatics pipeline for integrating binding signals across factors and histone marks, using our previous experience in mapping repressive marks on chromatin in single cells36. Statistical considerations. All statistical analyses are performed as in Aim 1, with clustering based on weighted nearest neighbor space combining different histone marks (e.g., H3K27ac and H3K27me3), as in36.
[0396] Pitfalls, alternatives and expected outcomes. We have previously demonstrated that NTT-seq can be used to simultaneously profile multiple (up to 3), non-overlapping histone modifications to profile open chromatin or active transcription in the context of epigenetic background36. We aim to integrate D&D-seq with multiplexed histone profiling, but if we fail to obtain high-quality profiles, we focus on single histone modifications. As potentially more transient chromatin remodelers may have lower D&D signal than stably bound TFs, we take advantage of the multi-subunit nature of these complexes, and use a cocktail of antibodies against multiple subunits of the same complex, to increase the signal- to-noise ratio. We expect that methods are used to define the in vivo interactions between TFs or chromatin remodelers and histones in single cells. These methods are able to assess how TF binding is affected directly by repressive or activating histone marks. Additionally, by overcoming the limitation that D&D combined with ATAC-seq has for analyzing open chromatin only, adaptation of D&D-seq for profiling DNA binding proteins within heterochromatin can shed light on how pioneer factors interact with closed chromatin in vivo, or what role TFs have in establishing or maintaining heterochromatin regions61. Furthermore, these multimodal assays allow for the reconstruction of the kinetics of chromatin modifiers catalyzing histone modification (e.g., PRC2 binding and H3K27me3 marks), and the assessment of TF binding dynamics across pseudotime from human samples, with matched histone profiles36.
[0397] Performance Measures and milestones for Aim 2. As for Aim 1, we aim to obtain 5,000 high-quality cells per experiment, with a minimum cell-specific ATAC peak TSS enrichment score of 5 and a minimum number of unique fragments of 3,500 to ensure sufficient library complexity. We aim to obtain >500 D&D counts genome-wide for GATA1 in K562 cells, profiled together with H3K27me3 and H3K27Ac, before moving to PBMCs. As there are no available single-cell TF binding profile methods that can be combined with single-cell histone modification profiling, for benchmarking, all D&D-seq peaks and NTT-seq peaks are pseudobulked and compared to available bulk ENCODE ChlP-seq data41, as in Fig. 3E. For integration of D&D-seq with NTT-seq, we benchmark D&D genomic deamination footprinting performance for each target protein assayed against the non-targeted IgG control that should not induce base editing, plus or minus NTT-seq processing.
[0398] D5. Aim 3: Functionally defining protein-DNA interactions through integration of multi-omics profiling in single cells. Rationale: Given the ability of D&D-seq to generate in vivo single-cell TF binding maps, we further aim to extend this technology to additionally capture downstream gene expression programs, as well as upstream somatic genetic variants, to deliver powerful single cells tools capable of dissecting the mechanisms of gene regulation at the level of the single cell. As such, we aim to combine our ability to simultaneously map TF binding and chromatin accessibility in single cells with transcriptome profiling, to capture transcription regulation across different layers, leveraging the versatility of our approach for incorporation with available single-cell sequencing approaches through integrating D&D-seq with the lOx Genomics Multiome platform. Furthermore, we incorporate our genotype-aware single-cell multi-omics framework, GoT-ChA15, with D&D-seq to enable the analysis of single-cell TF binding in cells harboring driver mutations directly from human samples, mapping genotype to singlecell TF binding phenotype in age-related clonal mosaicism (such as CHIP).
[0399] Aim 3a. Simultaneous measurement of chromatin accessibility, protein-DNA interaction, and transcriptome at single-cell level. We have previously used single-cell multiomics to measure the dynamics of gene expression across molecular features, including domains of open regulatory chromatin (DORCs), gene expression and protein levels across pseudotime trajectories of single cells from human myelofibrosis samples15, allowing for the assessment of regulatory dynamics across molecular features. To investigate the complex interplay between epigenetic and transcriptomic regulation in the context of single-cell in vivo TF binding, we develop a method to simultaneously measure chromatin accessibility, DNA-protein interaction, and transcription at the single-cell level. The D&D workflow was designed to be fully compatible with droplet microfluidic products that are commercially available. Therefore, here we aim to evaluate the performance of the lOx Genomics Chromium Single Cell Multiome ATAC + Gene Expression in combination with D&D. Although the two workflows are perfectly compatible, the D&D transcription factor footprinting approach consists of several additional steps that ultimately could interfere with RNA retention and integrity. Therefore, we aim to evaluate the feasibility of this approach by performing a cell-mixing experiment, similar to what is presented in our preliminary data (Section C). Briefly, isolated cells are first fixed and permeabilized to allow in situ reactions. Then, the sample is incubated with an antibody specific to the DNA binding target of interest. After washing the antibody in excess, the sample is incubated with the D&D-deaminase specific to the antibody used to label the target protein. The footprinting reaction is started by the addition of the activating peptide. At this point the sample can enter the lOx Genomics Multiome workflow, which consists of tagmentation of open chromatin regions, cell droplet encapsulation, and barcoding.
[0400] Extensive bioinformatic analysis is performed to evaluate the quality of the components of our multiomic approach. Since we already know that open chromatin is not affected by the D&D reaction, we focus in particular on determining quality metrics to verify the quality of the transcriptome. As in our previous work, we leverage our cell mixing control experiments profiling CTCF and GATA 1 in lymphocyte cell line CA46 and erythroblast cell line K562 (as in Fig. 6) to quantitatively assess key metrics for each modality, including sensitivity (molecules / cell), specificity (% reads mapping to a single cell type), and doublets rate, during technology development. As the D&D and ATAC portions of the technology have been optimized in Aim 1, we focus on attaining high- quality RNA metrics during this benchmarking stage, and aim to profile additional targets that have been already been optimized for D&D-seq alone.
[0401] Aim 3b: Profiling effects of somatic mutation on DNA-protein interactions in single cells. In order to build a comprehensive toolkit to capture gene regulation in aging single cells, we reasoned that incorporating genotyping capability would enable the extension of D&D-seq to the study of somatic mosaicism. There is increasing recognition that tissues throughout the body contain many thousands of clonal expansions fueled by somatic driver mutations that confer a growth advantage, observed even in seemingly normal tissue at increasing levels with age. Indeed, recent data revealed that clonal expansions are nearly ubiquitous in hematopoietic stem cells of older individuals, resulting in highly prevalent CHIP by the age of 7062. Modeling this human somatic evolution is challenging, as cell culture or murine models may not accurately reflect the evolutionary processes that occur in humans over decades. As such, primary human samples are critical for determining mechanisms of clonal outgrowth in human clonal mosaicism, which remain largely unknown. Given the mixture of mutated and wild-type cells within clonally mosaic tissues, single-cell multi-modal methods that capture both genotype and phenotype together are required to provide biological insights into mechanisms of clonal outgrowth in aging tissue. To address this challenge, our group previously developed Genotyping of Transcriptomes (GoT63) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA15) to enable high-throughput single-cell genotyping together with transcriptome profiling or chromatin accessibility profiles, respectively. We used these methods for single-cell genotyping-to-phenotype mapping of non-malignant clonal hematopoiesis, uncovering the mechanism for how global disruption to DNA methylation24,34, splicing35or signal transduction pathways15leads to cellular differentiation skews in mutant cells, directly in human samples. Leveraging this conceptual framework for linking single-cell genotype to molecular phenotype, we now aim to incorporate our D&D-seq approach into single-cell genotype-aware multi-omics, to connect age-related somatic driver mutations to TF binding. To this end, we integrate GoT-ChA genotyping primers into the D&D-seq protocol for targeted or multiplexed profiling of somatically mutated driver genes, optimizing our workflow through genotyping of human CHIP mutations.
[0402] In order to develop genotype-aware D&D-seq to understand how somatic mutations can disrupt TF binding, we aim to adapt our GoT-ChA15framework for single-cell chromatin accessibility profiling combined with somatic mutation genotyping to integrate into the D&D workflow. Briefly, in GoT-ChA, isolated nuclei are subjected to transposition of genomic DNA (gDNA) and loaded into microfluidics devices (40K cells needed for 10K target). During cell barcoding reactions, additional gene-specific primers are added to capture a locus of interest with an in-droplet PCR reaction, with a handle that is compatible with the lOx snATAC platform. The product is then split, with an aliquot used for an amplicon genotype library and the remaining 90% used for AT AC library construction. The ATAC and genotyping libraries can then be analytically integrated via shared cell barcodes, linking chromatin accessibility to targeted genotyping at the single-cell resolution. To incorporate genotyping into the D&D-seq framework, GoT-ChA primers can be added during the ATAC step (Fig. 10a).
[0403] For a pilot experimental application of GoT-ChA D&D-seq, we profiled CTCF binding in PBMCs from a patient with clonal hematopoiesis carrying an IDH2R140Qmutation with a variant allele frequency of 0.15 identified by a targeted panel sequencing. Integration of GoT-ChA with D&D-seq allowed us to perform genotyping in individual cells, along with chromatin status and CTCF binding status. We first performed cell clustering based on chromatin accessibility data, agnostic to genotype, and projected genotypes onto the UMAP (Fig. 10b). We successfully genotyped 25.32% of single cells, identifying homozygous mutant, heterozygous and homozygous wild-type cells. There was uneven distribution of mutant cells across cell types, with most of the IDH2 mutant cells concentrated in the CD8 T cell cluster (Fig. 10b). We next assessed the CTCF binding by comparing the CTCF DnD signals between IDH2 wild type (WT), mutant (1DH2R14OQ, MUT), and heterozygous (HET). Notably, we observed that CTCF binding signal was significantly decreased in mutant cells (Fig. 10c), a finding that is consistent with the decreased CTCF binding that has been reported in IDH-mutant glioma and acute myeloid leukemia, mediated through DNA hypermethylation at CTCF binding sites64,65, demonstrating the potential for biological discoveries enabled by this analysis.
[0404] We optimize this method by profiling the binding of key lineage-defining TFs (e.g., GATA1, FLI1) in CHIP samples harboring mutations in epigenetic modifiers (e.g., DNMT3A, TET2 or ASXL1). We additionally identify and profile samples with myeloproliferative neoplasm-associated somatic mutations in TFs, including RUNX166. We compare the D&D counts between mutant and wild-type cells and perform de novo motif discovery as well as known canonical motif enrichment analysis (as in Fig. 7c) to establish an analytical pipeline that considers level of genotyping efficiency and D&D read count. The ability to obtain single-cell TF binding profiles together with genotyping allow to pursue further this mechanistic link showing direct alteration in TF binding preferences. Thus, this approach allow us to directly assess how CHIP-associated mutations disrupt TF regulatory landscapes.
[0405] Statistical considerations. General statistical considerations are the same as for Aim 1. Based on our previous work, ~5K cells / sample in GoT-ChA experiments is sufficient to enable the comparison between mutant vs wild-type cells, considering ~7-10 main progenitor clusters, >5% mutant cell frequency, and genotyping efficiency of >50%. Given the ability to compare mutant vs wild-type cells within patient samples (controlling for patient-specific and technical variabilities), this sample size is sufficient to statistically substantiate our results63.
[0406] Pitfalls, alternatives and expected outcomes. Single-cell genotyping efficiency can vary across loci (for example, high GC content can lead to lower genotyping efficiency)15. While we have shown that GoT-ChA is capable of capturing both alleles in cells, incomplete capture is anticipated, which would affect the detection of heterozygous loci. This leads to mis-classification of mutated cells as wild type, due to allelic drop-out of the mutated allele. When possible, we optimize primers to include a neighboring germline polymorphic site if accompanying germline genotyping is available, to allow accurate phasing for allelic capture determination. Moreover, we have extensive experience addressing this challenge from our previous work15,24,63. This experience has shown that statistical inference can be optimized to mitigate this limitation, for example through amplicon library downsampling, and that despite the decrease in effect size due to mutated- > wild type mis-assignment, clear genotype-phenotype inference is feasible. We leverage our rich experience in these analyses to address the potential confounder of inter-individual variability. For this, we leverage the direct comparison of mutated and wild-type cells within the same individual, obviating sample- and individual- specific confounders. Integration across individuals then leverages linear mixed modeling and downsampling to ensure the learning of generalizable and mutation- specific genotype-phenotype inferences. We expect that these tools can be used to interrogate the dynamics of TF binding and transcriptional outputs for individual genes across cell types, with the potential broad application across systems for new insights into normal or disrupted cellular function during aging. We further expect that our genotype-aware methods for simultaneously assessing TF binding and chromatin accessibility can be used to define mechanisms of clonal outgrowth, manifesting as differential motif usage or TF binding that directs mutant cell-specific gene expression, driven by age-related somatic mutations, across cell types. This approach may help to identify mutant cell- specific vulnerabilities that can be exploited to eliminate pathogenic clones in somatic diseases48, presenting a powerful new tool for single-cell genomics in medicine.
[0407] Performance Measures and milestones for Aim 3. We optimize D&D-GoT-ChA for a subset of GoT-ChA targets that are commonly mutated in CHIP using cell line admixtures. For genotyping, we measure on-target rates and optimize or re-design primers until each assay has >40% on-target rate (defined as correct read structure). Using cellmixing experiments, we aim to produce optimized GoT-ChA-D&D-seq assays for 3 loci at >35% genotyping efficiency and 90% accuracy, with >90% specificity for mapping scATAC-seq fragments to the appropriate cell line. Based on barcode quality control, we aim to obtain 1,000 high-quality cells per experiment, with a minimum cell-specific AT AC peak TSS enrichment score of 5 and a minimum number of unique fragments of 3,500 to ensure sufficient library complexity. After achieving these technical milestones, we aim to profile primary CHIP samples with optimized genotyping primers and processing pipelines.
[0408] Example 8: Methods and Materials for Examples 9-14 Methods
[0409] Cell culture
[0410] Human K562 (ATCC, CCL-243) and CA46 (ATCC, CRL-1648) cell lines were maintained according to standard procedures in RPMI-1640 (Thermo Fisher Scientific, 11- 875-119) with 10% FBS (Thermo Fisher Scientific, 10-437-028) at 37 °C with 5% CO2. Cell lines in culture were screened biweekly for mycoplasma contamination using the MycoAlert PLUS Mycoplasma Detection Kit (Lonza, LT07-703).
[0411] Primary cell acquisition and processing
[0412] Frozen PBMCs used for scD&D-seq were thawed into DMEM with 10% FBS, spun down at 4 °C for 5 minutes at 400g and washed twice with PBS with 2% BSA. Live cells from the PBMCs were enriched by the Dead Cell Removal Kit (Miltenyi, 130-090-101) per the manufacturer's instructions.
[0413] Cloning of nb-DddA constructs
[0414] Previously published sequences coding for the DddAll enzyme were split at position 1397, and the C terminus of DddAl 1 (1290-1397) was synthesized as a gene fragment (Integrated DNA Technologies (IDT), Supplementary Note) flanked by restriction enzyme sites EcoRI and Spel. The gene fragments were digested with EcoRI and Spel for 1 hour at 37 °C, ligated for 1 hour at room temperature with pTXBLalnbOc-Tn5 (Addgene, 184285), pTXBl-nbMmKappa-Tn5 (Addgene, 184286), or 3xFlag-pA-Tn5-Fl (Addgene, 124601) digested with the same enzymes. The ligated product was transformed into competent cells (New England Biolabs, C2987) per the manufacturer’s instruction. The final plasmid products (pTXBl-nb-DddA_NT) were confirmed by Sanger sequencing (Eton Bio) or Whole Plasmid Sequencing using Oxford Nanopore Technology with custom analysis and annotation (Plasmidsaurus). nb-DddA production
[0415] The pTXBl-nb-DddA_NT vectors were transformed into BL21(DE3)-competent Escherichia coli cells (NEB, C2527), and nb-DddA_NT was produced via intein purification with an affinity chitin-binding tag24,59. First, transformed bacterial colonies were grown overnight in 5 ml Luria broth (LB). Next day, 5 ml of LB culture was added to 400 ml and grown at 37 °C to optical density (OD600) = 0.6. nb-DddA expression was induced with 0.5 rnM isopropyl-B-D-thiogalactopyranoside (IPTG) and 50 pM ZnCh at 30 °C for 4 hours. After induction, cells were pelleted and then frozen at -80 °C overnight. Cells were then lysed by sonication in 30 ml HEGX (20 mM HEPES-KOH pH 7.5, 0.8 M NaCl, 10% glycerol, 0.2% Triton X-100) with a protease inhibitor cocktail (Roche, 04693132001). The lysate was pelleted at 10,000g for 20 minutes at 4 °C. The supernatant was transferred to a new tube, and 600 pl of neutralized 10% polyethyleneimine (Sigma- Aldrich, P3143) was added drop wise to the bacterial extract, gently mixed and centrifuged at 12,000g for 40 minutes at 4 °C to precipitate DNA. The supernatant was loaded on a 2-ml chitin column (NEB, S665 IS), followed by washing with 12 ml of HEGX. Then, 3 ml of HEGX containing 100 mM DTT was added to the column with incubation for 48 hours at 4 °C to allow cleavage of nb-DddA from the intein tag. After incubation, 2 ml HEGX was added to elute nb-DddA directly into a 10-kDa molecular weight cutoff (MWCO) spin column (Thermo Scientific, 88527). Protein was dialyzed three times using 15 ml of 2x dialysis buffer (100 HEPES-KOH pH 7.2, 0.2 M NaCl, 0.2 mM EDTA, 2 mM DTT, 20% glycerol) and concentrated to 0.5-1 ml by centrifugation at 5,000 g. The protein concentrate was transferred to a new tube, mixed with an equal volume of 100% glycerol, and stored at -20 °C. After purification, the protein was denatured at 95 °C for 5 minutes, analyzed by SDS-PAGE gel (BIO RAD, 4561085) and imaged by BIORAD ChemiDoc Touch Imaging System (Fig. 43b).
[0416] Deamination assay
[0417] DNA substrates, including lambda phage DNA (NEB, N3011) or 5' 6-FAM-labeled 30-mer dsDNA oligonucleotides (IDT, Supplementary Note) were used for testing nb- DddA deamination activity. Deamination reactions of 250-1000 ng DNA were performed in 50 pl of deamination buffer (40 mM Tris-HCl pH 7.4, 50 mM KC1, 1 mM MgC12, 1 mM dithiothreitol (DTT), 20 pM ZnC12) with 50 pM nb-DddA_NT and 100 pM DddA_CT (Eton Bio, HPLC-purified, purity >95%, Supplementary Note). The reactions were incubated at 37 °C for 1 hour unless otherwise indicated. For lambda phage DNA, the products were then purified with 1.6X Ampure beads (Beckman Coulter, A63881), treated with 1 ul of T7 endonuclease (NEB, M0302) in 30 ul T7 buffer, further incubated at 37 °C for 15 minutes per the manufacturer’s instructions, and analyzed by an agarose gel (Invitrogen, A42135). For 6-FAM-labeled dsDNA oligonucleotides, 1 ul USER enzyme (NEB, M5508) was added to the deaminase-treated products, incubated at 37°C for 1 hour, denatured at 95°C for 5 minutes, and analyzed by 15% polyacrylamide TBE-Urea gel (BIORAD, 4566055). The cleaved DNA fragments were imaged by BIORAD ChemiDoc Touch Imaging System. Antibodies
[0418] Antibodies used were CTCF (1:100, Active Motif, 61932), GATA1 (1:100, Abeam, abll852), and GATA2 (1:100, Invitrogen, PAI-100).
[0419] D&D-seq
[0420] Cell fixation, Permeabilization, and Antibody binding
[0421] For all the following centrifugations, cells were centrifuged in a swing-bucket centrifuge at 300g for 5 minutes before fixation and 600g for 5 minutes after fixation. About 200K to 1 million cells were collected and resuspended in 400 pl PBS. Then, 16% methanol-free formaldehyde (Thermo Fisher Scientific, PI28906) was added for fixation (final concentration 0.1%) at room temperature for 5 minutes. Glycine (final concentration 125 mM) was added to stop the cross-linking, followed by a wash with 1 ml of PBS. Fixed cells were permeabilized in 200 pl permeabilization buffer (20 mM Tris-HCl pH 7.4, 150 mM NaCl, 3 mM MgCl2, 0.1% NP40, 0.1% Tween-20, 1% BSA, lx protease inhibitors) on ice (7 minutes for cell lines, 5 minutes for primary cells), followed by a wash with 200 pl cold wash buffer (20 mM HEPES pH 7.6, 150 mM NaCl, 0.5 mM spermidine, 1% BSA, lx protease inhibitors). After cell counting, 100K-400K cells were transferred to PCR tubes and resuspended in 100 pl antibody buffer (20 mM HEPES pH 7.6, 150 mM NaCl, 2 mM EDTA, 0.5 mM spermidine, 1% BSA, lx protease inhibitors), with 1 pl (1:100 dilution) of the antibodies. The cells were incubated at 4 °C with slow rotation overnight.
[0422] D&D binding and activation
[0423] The cells were washed once with 200 ul wash buffer, and another time with 200 ul D&D binding buffer (Supplementary Note). The cells were resuspended in 50 ul D&D activation buffer (Supplementary Note) with 100 uM DddA_NT, and incubated at room temperature on a rotator for 1 hour. The cells were then washed with 200 ul D&D binding buffer, resuspended in 50 ul D&D activation buffer (Supplementary Note), and incubated at 37 °C for 1 hour for deamination.
[0424] Bulk ATAC
[0425] Tn5 adaptors were purchased from IDT. Adaptors (100 uM) (Supplementary Note) were annealed in TE buffer to form mosaic-end, double- stranded (MEDS) oligos by incubating at 95 °C for 5 minutes and then cooling at 0.2 °C per second to 12 °C. MEDS-A and MEDA-B were mixed 1:1, and 2 pl was transferred to a new tube and mixed with 18 pl of TnY (in-house) enzyme after 1 hour at room temperature to allow for transposome assembly60. After D&D activation, the cells were washed once with 200 ul lx Tris-TD buffer and resuspended in 50 ul lx Tris-TD buffer with 9 ul loaded TnY (TnY volume depends on its concentration and needs to be titrated to get the optimal tagmentation). To initiate tagmentation, the reaction was incubated at 37 °C for 1 hour. To extract DNA, 1 ul 10% SDS, 2 ul proteinase K (NEB, P8107S), and 3 ul 0.5M EDTA were added to the reactions, which were then incubated at 55 °C for 1 hour, followed by column-based DNA purification per the manufacturer’s instruction (Zymo, D5205). For library preparation PCR, we used uracil-tolerant DNA polymerase (NEBNext Q5U Master Mix, NEB M0597S) and Nextera-compatible indexing primers (Supplementary Note) to amplify the purified DNA. To increase the efficiency of the initial gap fill-in step, we spiked in 1 ul of non-hot-start polymerase Bst 3.0 (NEB, M0374). The PCR steps are: 72°C for 5 min; 98°C for 30 s; 12 cycles of 98°C for 10 s, 55°C for 30 s and 72 °C for 30 s; followed by 72 °C for 5 min.
[0426] Bulk PTA
[0427] Bulk PTA was performed using the ResolveDNA Whole Genome Amplification Kit (BioSkryb Genomics, 100136) per the manufacturer’s instructions. Briefly, after D&D activation, cell lysis was performed to denature DNA, followed by PTA amplification and DNA Cleanup. WGS libraries were generated using KAPA HiFi HotStart Library Amp Kit (Roche, 07958960001). Libraries were sequenced to 3-5x genome coverage (paired-end lOObp).
[0428] Single-cell ATAC
[0429] After D&D activation, the cells were resuspended in lx nuclei buffer and counted using trypan blue and a Countess II FL Automated Cell Counter. For the remaining steps, we follow single-cell ATAC-seq per lOx protocol (version CG000209 Rev F, lOx Genomics) with the following modifications.
[0430] 1. During the GEM generation and barcoding reaction (step 2.1), 2 ul of uracil- tolerant DNA polymerase (NEBNext Q5U Master Mix, NEB M0597S) was added to the barcoding reaction mixture to facilitate the first few rounds of PCR, where we expect to have some uridine in the reaction.
[0431] Sequencing
[0432] The sequencing libraries were sequenced on a NovaSeq6000, NovaSeqX Plus, or NextSeq2000 with dual indexed, paired-end 100 or 150-bp settings. Specifically, i5: 8 bp (16 bp for single-cell sequencing), i7: 8 bp, readl: 100 or 150 bp, read2: 100 or 150 bp. Genotyping of Targeted loci with Chromatin Accessibility (GoT-ChA)
[0433] We performed GoT-ChA according to the previously published paper with minor modifications21. Briefly, we performed whole cell fixation, permeabilization, antibody incubation, D&D binding and activation as aforementioned bulk D&D-seq. After D&D activation, the cells were resuspended in lx diluted nucleus buffer (lOx Genomics) and counted using trypan blue and a Countess II FL Automated Cell Counter. Afterwards, the cells were processed according to the Chromium Next GEM Single Cell ATAC Solution user guide (version CG000209 Rev F, lOx Genomics) with the following modifications:
[0434] 1 . During the GEM generation and barcoding reaction (step 2.1), 1 pl of 22.5 pM GoT-ChA primer mix was added to the barcoding reaction mixture. The primers used are IDH2 locus-specific primers IDH2_R14O_F1 (or IDH2_R140_F2) and IDH2_R14O_R1 (or IDH2_R140_R2) (Supplementary Note). These primers allow for exponential amplification of the GoT-ChA fragments relative to the linear amplification of ATAC fragments. In addition, to facilitate the first few rounds of PCR, where we expect to have some uridine in the reaction, we spiked in 2 ul of uracil-tolerant DNA polymerase (NEBNext Q5U Master Mix, NEB M0597S) into the barcoding reaction mixture.
[0435] 2. During the post-GEM incubation clean-up (step 3.2), 45.5 pl of elution solution I is used to elute material from SPRIselect beads. A total of 5 pl is used for GoT- ChA library construction, and the remaining 40 pl is used for ATAC library construction as indicated in the standard protocol.
[0436] 3. To generate the GoT-ChA library, two additional PCRs were performed on the 5 pl set aside during step 3.2. The first PCR aims to amplify genotyping fragments before sample indexing and uses P5 and IDH2_R14O_N1 (or IDH2_R140_N2) primers (Supplementary Note) with the following thermocycler program: 95 °C for 3 min; 15 cycles of 95 °C for 20 s, 65 °C for 30 s and 72 °C for 20 s; followed by 72 °C for 5 min and ending with hold at 4 °C. After a 1.2x SPRIselect clean-up, biotinylated PCR product is bound and isolated using Dynabeads M-280 Streptavidin magnetic beads (Thermo Fisher Scientific, 11206D). In brief, the beads are washed three times with lx sodium chloride sodium phosphate-EDTA buffer (SSPE, VWR, VWRV0810-4L), added to the purified PCR product and incubated at room temperature for 15 minutes. The beads are then washed twice with lx SSPE buffer and once with 10 mM Tris-HCl (pH 8.0) before resuspending in water. The bead-bound fragments are then amplified and sample indexed using P5 and RPI- X primers (Supplementary Note) with the following thermocycler program: 95 °C for 3 min; 6-10 cycles of 95 °C for 20 s, 65 °C for 30 s and 72 °C for 20 s; followed by 72 °C for 5 min and ending with hold at 4 °C.
[0437] Final libraries were quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Q32854) and the High Sensitivity DNA chip (Agilent Technologies, 5067-4626) run on a Bioanalyzer 2100 system (Agilent Technologies) and sequenced on a NovaSeq 6000 or NovaSeq X system at the Weill Cornell Medicine Genomics Resources Core Facility with the following parameters: paired-end 100 or 150 cycles; read IN, 100 or 150 cycles; i7 index, 8 cycles; i5 index, 16 cycles; read 2N, 100 or 150 cycles. ATAC libraries were sequenced to a depth of 25,000-35,000 read pairs per cell and GoT-ChA libraries were sequenced to 5,000 read pairs per cell. A list of the primer sequences used in this study is provided in Supplementary Note.
[0438] Cell sorting + Genotyping
[0439] Cryopreserved peripheral blood mononuclear cells from a CHIP donor with IDH2R140Q mutation were thawed and stained. Briefly, the cells were resuspended in staining buffer (BioLegend, 420201) and incubated with Human TruStain FcX (10 min at
[0440] 4 °C; BioLegend, 422302) to block Fc receptor-mediated binding. Then, the cells were stained with CD8-BV650 (Biolegend, 344730), CD4-FITC (Biolegend, 344604), CD19-PE (BD biosciences, 561741), CD14-APC (Thermo Fisher, 47-0149-41), CD56-BV786 (BD biosciences, 564058) (1 : 100 for up to 106 cells in a final volume of 100 pl) for 20 minutes at 4 °C, and DAPI (Invitrogen, D1306). The samples were then sorted for DAPI- and single-marker positive cells (CD8 for CD8+ T cells, CD4 for CD4+ T cells, CD19 for B cells, CD14 for monocytes, CD56 for NK cells) using the BD FACSymphony™ S6 Cell Sorter (Supplementary Note). The genomic DNA of the sorted cells was extracted using Puregene cell kit (QIAGEN, 158043) per the manufacturer’s instruction, and PCR- amplified using KAPA 2X mix (Roche, 07958927001) and IDH2_R140_F2 + IDH2_R140_R2 (Supplementary Note) primers. The linear PCR amplicon sequencing was performed by Plasmidsaurus using Oxford Nanopore Technology with custom analysis, and the genotype ratio was quantified by the number of reads mapped to IDH2 wild type and IDH2R140Q.
[0441] Bulk Data analysis
[0442] Preprocessing of bulk ATAC-seq data
[0443] Raw bulk D&D-seq FASTQ data were analyzed using FastQC for initial quality control61. Nextera transposase sequence and homopolymer G were removed using CutAdapt with “CTGTCTCTTATACACATCTCCGAGCCCACGAGAC” for R1 and “CTGTCTCTTATACACATCTGACGCTGCCGACGA” for R2 and "G{ 12}" parameters whenever identified in the FastQC report62. After trimming adapter and polyG sequences, processed FASTQ files were re-analyzed using FastQC and aligned to the Human reference genome (GRCh38) using BWA-MEM263. Read alignments were subsequently sorted and indexed using samtools. To call D&D edits in bulk datasets, bam files were analyzed as described in the following Extraction of D&D signal section.
[0444] For TF motif discovery, genomic regions of 400 bp to 1 kb (indicated in the figures) centered at the peak summits were used to query the genome using MEME64. To visualize the genomic tracks on IGV, bigwig files were generated using the DeepTools bamCoverage function with the -normalizeUsing BPM option set65. ChlP-seq peak coordinates for CTCF, GATA1 and GATA2 for K562 cells were downloaded from ENCODE20.
[0445] Copy number alteration analysis
[0446] Bam file was transformed to bed format using bedtools bamtobed with default parameters. Copy number prediction was performed using Ginkgo (https: / / doi.org / 10.1038 / nmeth.3578). Briefly, genome data was built using the hg38 human reference genome, and a config file for the analysis was generated with following modifications from the provided config file, binMeth=variable_500000_101_bwa and chosen_genome=hg38. Copy number plot was generated by running analyze.sh script with the bed and config files.
[0447] Single-cell data analysis
[0448] Preprocessing and annotation of lOx scATAC-seq data
[0449] Raw scATAC-seq datasets were aligned to the Human reference genome (GRCh38) and quantified using CellRanger-ATAC (version 2.1.0). The resulting peak / cell matrix was filtered for low quality cells and normalized using Signac (version 1.13.0)66. In detail, cells having 1) reads aligned to the blacklist region more than 0.05%, 2) less than 20% of reads aligned to peaks, 3) transcription start site enrichment score less than 3, 4) nucleosome signal higher than 4, 5) outliers having too many (average of peak counts + 2*standard deviation of peak counts) or fewer peaks (average of peak counts - l*standard deviation of peak counts) were filtered out (Fig. 44a, 45a, 46b). Multiplets were additionally annotated using AMULET and removed67. The preprocessed data were normalized by latent semantic indexing analysis, projected to low dimensions using uniform manifold approximation (UMAP), and clustered using the smart local moving (SLM) algorithm. Gene activity based on chromatin accessibility was inferred in order to annotate markers of each cluster. Finally, motif activity score was computed using ChromVAR with the JASPAR2020 core database {'ov Homo sapiens6'69.
[0450] For the cell line data, eight clusters were defined with a clustering resolution of 0.6, and annotated as one of two cell lines based on the marker gene activity. Cluster 7, which was found in both K562 and CA46 clusters, was removed. The remaining number of cells was 5,840 cells (CA46, n = 4,108 and K562, n = 1,732). For PBMC data, 18 and 16 clusters were identified with a clustering resolution of 0.8 in replicate 1 and 2, respectively. Clusters were annotated based on marker gene activity and cell type prediction using Azimuth with a Human PBMC reference dataset70,71. For replicate 1, two clusters were additionally removed because two different cell types were mixed (cluster 7, CD8 T cell and monocyte) or it had overall low quality and lacked markers (cluster 13). The total of 5,413 cells in replicate 1 and 10,394 cells in replicate 2 remained.
[0451] Genotyping of IDH2 mutation using GoT-ChA
[0452] Raw FASTQ files were first analyzed using FastQC to examine overall data quality and mutant allele at the targeted locus61. Based on a FastQC report, position of mutant allele, a read file with mutant allele, and primer sequence were identified. Input FASTQ files were then processed using the GoT-ChA analysis pipeline (available at github.com / landau-lab / Gotcha)21. In detail, raw data were first split into smaller files and quality filtering was performed with the “FastqFiltering” function. The R1 FASTQ file for replicate 1 and R3 FASTQ file for replicate 2 were used for genotyping with c(37:39) as a mutation site. After quality control of raw data, mutation state is annotated with the “BatchMutationCalling” function based on the provided wild-type / mutant and primer sequences. For both replicates, the primer sequence was set as “GATGGGCTCCCGGAAGACAGTCCCCCCCAGGATGTT”, and “CCG” and “CTG” were set as wild-type and mutant sequences, respectively. Results from each split data were merged into a single barcode and read count matrix using “MergeMutationCalling”. Genotype of each cell was assigned by manually setting thresholds based on distribution of log-transformed read counts (Fig. 46d).
[0453] Extraction of D&D signal
[0454] To summarize genome edits introduced by D&D in scATAC-seq data, a bam file produced by CellRanger was split into each cell type based on barcode sequence and cell annotation using sinto. For bulk ATAC-seq data, bam files aligned to the genome using BWA-MEM2 were analyzed63. After splitting bam files, the following five steps were performed to analyze D&D-mediated genomic variants (Fig. 47c). All steps are integrated in the D&D analytic pipeline.
[0455] 1) First, each bam file was preprocessed to remove uninformative and low quality read alignments. Duplicated reads were marked and simultaneously filtered using “picard MarkDuplicates” with a REMOVE_DUPLICATES-true parameter. Read alignments with high mapping quality Phred score (>= 20), primary alignment, reads aligned to intact chromosomes, and those with properly aligned mates were retained using samtools (version 1.19).
[0456] 2) Next, all single nucleotide variants (SNVs) found in each filtered bam file were collected using “bcftools mpileup” with following parameters, -a FORMAT / AD,FORMAT / DP,INFO / AD -no-BAQ -min-MQ 1 -max-depth 8000. The pileup result subsequently converted into the vcf format reporting SNVs supported by at least two reads supporting variants from minimum three aligned reads. These thresholds can be adjusted by users based on their data. Alternatively, users can use pysam with multithread options to collect variants from the bam file.
[0457] 3) Germline mutations were then filtered based on loci and alleles from the gnomAD database and variant allele frequency higher than 10%. After filtering out germlines and high VAF variants, the remaining C-to-T edits were highly enriched with D&D edits and depleted of residual somatic mutations (Fig. 47d, 47e). If available, custom databases in vcf format can be provided to additionally filter uninformative mutations. In this step, we generate two vcf files after filtering germlines; one for all possible single variants in the sample and the other specifically for C-to-T or G-to-A variants which is used to identify edited peaks in the following step.
[0458] 4) Preprocessed bam files from step 1 were analyzed using MACS2 (version 2.2.9.1) to call peaks with -fBAM —nomodel parameters. Peaks were then filtered with blacklist region annotation using bedtools (version 2.31.1)72. Motif analysis was performed using MEME Simple Enrichment Analysis (SEA)73with HOmo sapiens Comprehensive MOdel Collection (HOCOMOCO)31vll core motif set to identify binding sites in peaks. Optionally, users can perform motif analysis using different motif databases supported by MEME, HOMER2 or a reference bed file generated by ChlP-seq.
[0459] 5) Called peaks are separated into two groups based on the presence of the motif sequence. Peaks harboring motifs of interest or overlapping with ChlP-seq reference tracks were classified as target peak candidates. Target peak candidates were resized to 200 bp (up / downstream 100 bp from the motif center or peak summit with ChlP-seq. The size can be adjusted by users) and overlaid with a vcf file which annotates C-to-T and G-to-A variants from step 3 using bedtools intersect. Overlapping peaks with motif and D&D edits were annotated as target peaks. When multiple motif positions were found in a single peak, a position with the highest motif score was chosen. Peaks without binding motifs of interest are classified as background peaks and were resized to 200 bp by taking + / - 100 bp from the peak summit. Background peaks were also overlaid with D&D edits and compared with target peaks in footprinting analysis. We chose a window size of 200 bp because D&D edits in the majority of TFs were enriched around + / -20-50 bp from the motif sequence and started to be depleted. As presented in the Fig. 47f GATA1 binding site of the TALI promoter, D&D edits are depleted within the TF motif and enriched near the up- / downstream of the motif sequence which was not found in regular ATAC-seq.
[0460] To visualize D&D signals, QNAMES (read id) of reads with D&D edits were extracted from alignment files (bam) using pysam. Briefly, D&D variants and positions are searched in individual reads and QNAMES of reads with matching variant status and position were extracted and saved in a text file. Edited reads were then extracted from the read alignment using samtools with view -N <qnames.txt> . D&D bam was then converted into bigwig using bamCoverage from the DeepTools package. Bigwig files were visualized using the trackplot package in R. Extraction of D&D reads is also integrated in the D&D edit calling pipeline.
[0461] Evaluation of D&D signal
[0462] To compare D&D edit counts between target and background regions, edit counts per peak were summarized and signal-to-noise ratio (SNR) were calculated by dividing the number of C-to-T and G-to-A by the number of other variants. First, edits per peak were calculated by dividing SNV counts by the number of peaks in resized target and background peak regions. We assume that the frequencies of non-D&D edit (edits other than C-to-T) are consistent across the target and background regions, so this can be used to normalize the D&D edit counts. In this manner, SNR was then calculated by dividing the D&D edit counts by the mean of non-D&D edit counts.
[0463] Footprint analysis for D&D edits was performed by counting the number of D&D edits in each base pair from randomly sampled target and background peaks, +-100 bp from the center of the motifs. The random sampling was repeated for 10 times and mean and standard deviation were used for visualization. For both the cell mixing experiment and primary blood cells, we used 200 randomly selected peaks. The number of subsampled peaks can be set by the user.
[0464] Benchmark analysis with ultra-low-input cleavage under targets and release using nuclease (uliCUT&RUN)
[0465] Raw uliCUT&RUN data for CTCF and negative controls with no primary antibody were downloaded from GEO (GSE111121)33. The quality of raw FASTQ files was assessed using FastQC61and the adapter sequence in both reads was trimmed using CutAdapt62with -a AGATCGGAAGAG -A AGATCGGAAGAG -m 21 parameters. Trimmed reads were aligned to the mouse reference genome (mmlO) using BWA-MEM2 with 10 as the minimum seed length (-k 10)6i. Read alignments with high mapping quality Phred score (>= 20), primary alignment, reads aligned to intact chromosomes, and those with properly aligned mates were retained from the resulting bam files using samtools (version 1.19). The processed bed file of CTCF ChlP-seq was downloaded from an independent study which used the same cell line (E14 mouse embryonic stem cells, GSE11431)74. Because the bed file was aligned to mm8, the genomic coordination was updated to mmlO using LiftOver75, sorted and merged for overlapping coordination using bedtools72, and modified to SAF format to use it as a custom reference for alignment. Processed bam files were aligned to the SAF using FeatureCounts (Subread package version 2.0.4) with -p —countReadPairs -F SAF parameters76. Fraction of fragments in peaks was calculated by dividing the assigned fragment counts by the total fragment counts.
[0466] The total of 4,105 CA46 cells with D&D edits from the cell mixing experiment were sampled to 10, 50, 500, and 5,000 cells with replacement and repeated for 100 times. K562 cells with D&D edits (n = 1,290) were downsampled to 10, 50, 500, and 1,000 cells in the same manner. To compare D&D to uliCUT&RUN, all reads harboring C-to-T or G- to-A tentative D&D edits with matched downsampled cell barcodes were collected. The fraction of fragments in peaks was calculated as described above using K562 CTCF ChlP- seq from ENCODE.
[0467] Evaluation of CTCF binding in CD8 T cells
[0468] In our integrated PBMC dataset, CD8 and CD4 T cells formed a continuous cluster without a clear separation (Fig. 42a). Because the IDH2 mutation in these samples was specifically enriched in the CD8 T cell population (Fig. 42b-d), we re-clustered CD4 and CD8 T cell populations to extract high confident CD8 T cells and removed clusters expressing CD4 T cell or non-T cell markers and mixed CD4 and CD8 T cells. The remaining CD8 T cell cluster (n = 7,546) consisted of 2,132 (534 WT, 329 MUT, and 1,269 not genotyped) cells from replicate 1 and 5,414 (853 WT, 486 MUT, and 4,075 not genotyped) cells from replicate 2. These cells were further grouped into MUT-enriched (n = 6,437 cells) and WT-enriched (n = 1,109 cells) based on the clustering and genotype proportions (Fig. 42h). To compare differential CTCF binding between MUT and WT CD8 T cells, we leveraged the linear mixture model (LMM) from the GoT-ChA analytic pipeline which considers sample- specific batch effect and variability in cell numbers. LMM analysis was applied to 1,102 CTCF peaks harboring more than 5 D&D edits and found in more than 5 cells. We used a permissive false discovery rate (FDR) threshold of FDR < 0.25, which identified 104 CTCF peaks with differential CTCF bindings between MUT and WT (Fig. 42i).
[0469] Co-accessible peaks were identified using Cicero46. First, ATAC peaks in two replicates were merged using GenomicRanges::reduce and new count matrices were generated using Signac workflow for both replicates. Count data were loaded in R using the Monocle 3 workflow. Batch effect was normalized by aligning two replicates using align_cds with preprocess_method = "LSI", alignment_group = "Dataset” and reduce_dimension with preprocess_method = "Aligned" parameters. Cicero object was generated with batch corrected CellDataSet and UMAP coordinates from Monocle 3 and co-accessible peaks were inferred using the run_cicero function as described in the Cicero workflow (available at cole-trapnell-lab.github.io / cicero-release / ).
[0470] Hi-C contact map predictions
[0471] To generate Hi-C contact map predictions, we used the pseudobulk ATAC reads and D&D reads for each cell type as the input for C.origami35. The bam files of ATAC-seq and D&D reads were first converted to bigwig files using DeepTools with the command bamCoverage — normalizeUsing RPKM —binSize 1 —bam $bamfile -o $bigwig65. The bigwig file for D&D reads were then normalized against the bigwig of the ATAC and log2- transformed with a pseudocount of 1 using the command bigwigCompare —binSize 1 -bl $D&D.bigwig -b2 $ATAC.bigwig -o $D&D.normalized.bigwig. The ATAC bigwig files and the normalized D&D bigwig files were used as the inputs for inference of HiC contact maps using C.origami. For predictions generated without incorporating CTCF binding information, the ATAC bigwig files, along with a bigwig file containing zeros for all regions, were used as input. Code Availability:
[0472] Python and R scripts used in this study to analyze D&D-seq data are available on GitHub (available at github.com / sangholl30 / DnD / ). Detailed parameter settings and thresholds used in the analyses are described in the Methods. All analyses were performed using Python and R in a mamba virtual environment. Detailed software versions are also described in the Methods.
[0473] Data Availability:
[0474] Raw data and processed data files generated from cell lines are available at Gene Expression Omnibus (GEO) under accession number GSE296495. Patient raw sequencing data containing genomic sequences generated in this study are deposited at the European Genome-Phenome Archive under accession number (PRJEB89720). The GRCh38 reference genome was used for alignment of single-cell ATAC-seq data (refdata- cellranger-atac-GRCh38- 1.2.0), freely available from the lOx Genomics website (available at support.10xgenomics.com).
[0475] ChlP-seq datasets of CTCF (ENCFF111JKR), GATA1 (ENCFF844WTT), GATA2 (ENCFF997NUA) in K562 cell, human primary CD4 T cell (ENCSR470KCE), CD8 T cell (ENCSR116AKQ), B cell (ENCSR075NRV), monocyte (ENCSR162KZY), and NK cell (ENCSR856TKC); ATAC-seq dataset of K562 cell (ENCFF077FBI); Hi-C datasets of human primary CD4 T cell (ENCSR335JYP), CD8 T cell (ENCSR321BHC), B cell (ENCSR847RHU), monocyte (ENCSR236EYO), and NK cell (ENCSR971CJS) are from ENCODE.
[0476] Supplementary Note
[0477] DddAll (1290-1397) gblock GGATGTGAATTCGGTGGCGGTGGCTCTGGCGGTGGTGGGAGTGGAGGTGGGGG ATCAGGAGGAGGCGGTTCCCATATGGGATCTTACGCTTTAGGACCGTATCAGAT CAGTGCCCCTCAACTGCCAGCGTACAACGGTCAAACCGTTGGAACGTTTTACTA TGTGAATGATGCGGGGGGTCTTGAGAGCAAAGTATTTATTAGCGGAGGCCCCA CTCCCTACCCAAATTATGTAAGCGCCGGACACGTTGAGGGTCAGTCCGCACTTT TTATGCGCGACAACGGAATCAGTGAGGGGCTGGTGTTCCATAACAATCCAAAA GGGACTTGTGGATTCTGCGTTAATATGATTGAAACCCTTCTGCCAGAAAACGCT AAGATGACAGTGGTGCCCCCAGAAGGGTGCATCACGGGAGATGCACTAGTTGC CCT (SEQ ID NO: 21) DddAll_NT (1290-1397) amino acid
[0478] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPYPNYVSAGH
[0479] VEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVVPPEG (SEQ ID NO: 9)
[0480] DddAll_CT (1398-1422) amino acid
[0481] AIPVKRGATGETKVFIGNSNSPKSP (SEQ ID NO: 10)
[0482] Deamination assay oligo s
[0483] / 56-FAM / denotes 5' 6-FAM (Fluorescein)
[0484] / iMe-dC / denotes 5mC
[0485] AT AC adaptors
[0486] ME_19_Phos
[0487] / 5Phos / CTGTCTCTTATACACATCT (SEQ ID NO: 29)
[0488] ME-A
[0489] Unmethylated: / 5Phos / TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO: 30)
[0490] Methylated: T / iMe-dC / GT / iMe-dC / GG / iMe-dC / AG / iMe-dC / GT / iMe- dC / AGATGTGTATAAGAGA / iMe-dC / AG (SEQ ID NO: 31)
[0491] ME-B
[0492] Unmethylated:
[0493] / 5Phos / GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 32) Methylated: GT / iMe-dC / T / iMe-dC / GTGGG / iMe-dC / T / iMe- dC / GGAGATGTGTATAAGAGA / iMe-dC / AG (SEQ ID NO: 33)
[0494] MEDS-A
[0495] Anneal ME- A and ME19_Phos
[0496] MEDS-B
[0497] Anneal ME-B and ME19_Phos
[0498] ATAC indexing primers (Nextera-compatible)
[0499] S5XX
[0500] AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC (SEQ ID NO: 34)
[0501] N7XX
[0502] CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG
[0503] (SEQ ID NO: 35)
[0504] N denotes a 8 -bp barcoding index.
[0505] GoTChA primers
[0506] GoTChA_Forward_primer: a locus- specific primer with a Nextera R1 handle (TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG, 20-22 bp locus -specific sequence) (SEQ ID NO: 36)
[0507] GoTChA_Reverse_primer: A locus-specific primer with a IS2 handle for multiplexed GoTChA (optional for single-target GoTChA) (AGCAAGTGAGAAGCATCGTGTC + 20-22 bp locus-specific sequence) (SEQ ID NO: 37)
[0508] GoTChA_Nested_primer: a nested, biotinylated, locus-specific primer with a partial
[0509] TruSeq small RNA read 2 handle
[0510] ( / 5BiosG / CCTTGGCACCCGAGAATTCCA + 20-22 bp locus-specific sequence) (SEQ ID
[0511] NO: 38)
[0512] / 5BiosG / denotes 5' Biotin
[0513] IDH2_R14O_F1 (NexteraRl)
[0514] TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGCAGAGCCCACACATTTG*C*A
[0515] *C (SEQ ID NO: 39)
[0516] * denotes phosphorothioate bonds
[0517] IDH2_R14O_R1 (IS2)
[0518] AGCAAGTGAGAAGCATCGTGTCTGTGGCCTTGTACTGCA*G*A*G (SEQ ID NO: 40)
[0519] IDH2_R14O_N1 (ADT)
[0520] / 5BiosG / CCTTGGCACCCGAGAATTCCAGATGGGCTCCCGGAAGACAG
[0521] (SEQ ID NO: 41)
[0522] IDH2_R140_F2 (NexteraRl)
[0523] TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGATGGGCTCCCGGAAG
[0524] A*C*A*G (SEQ ID NO: 42)
[0525] IDH2_R140_R2 (IS2)
[0526] AGCAAGTGAGAAGCATCGTGTCCAGAGCCCACACATTTG*C*A*C (SEQ ID NO: 43)
[0527] IDH2_R140_N2 (ADT)
[0528] / 5Biosg / CCTTGGCACCCGAGAATTCCACTGGCTGTGTTGTTGCTTGG (SEQ
[0529] ID NO: 44) P5 (binds to the P5 Illumina sequencing handle) AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO: 45)
[0530] RPI-X: indexing primers that bind to the partial TruSeq small RNA read 2 handle and adds a sample index and P7 Illumina sequencing handle (CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCCTTGGCA CCCGAGAATTCCA, where N denotes a user-defined sample index) (SEQ ID NO: 46)
[0531] Buffers
[0532] Permeabilization Buffer (1 ml)
[0533] Wash buffer (10 ml)
[0534] Antibody buffer (1 ml)
[0535] • Wash buffer 1 ml + 0.5 M EDTA 4 ul, chill on ice 10X D&D Binding buffer (10 ml)
[0536] D&D Binding buffer (10 ml)
[0537] D&D Activation buffer
[0538] • Deamination buffer 1 ml + 1 mM ZnC1220 ul (final 20 M)
[0539] • Final component o 40 mM Tris-HCl o 50 mM KC1 o 1 mM MgC12 o 1 mM DTT o 20 pM ZnC12 o 0.5 mM Spermidine o 1% BSA
[0540] 5x Tagmentation buffer (TD-Tris buffer) (10 ml)
[0541] Example 9: Summary of Docking & Deamination Followed by Sequencing (D&D-seq) Technique
[0542] Gene expression is coordinated by a multitude of transcription factors (TFs), whose binding to the genome is directed through multiple interconnected epigenetic signals, including chromatin accessibility and histone modifications. These complex networks have been shown to be disrupted during aging, disease, and cancer. However, profiling these networks across diverse cell types and states has been limited due to the technical constraints of existing methods for mapping DNA:Protein interactions in single cells. As a result, a critical gap remains in understanding where TFs or other chromatin remodelers bind to DNA and how these interactions are perturbed in pathological contexts. To address this challenge, we developed a transformative single-cell immuno-tethering DNA:Protein mapping technology. By coupling a species-specific antibody-binding nanobody to a cytosine base editing enzyme, this approach enables profiling of even weak or transient factor binding to DNA, a task that was previously unachievable in single cells. Thus, our Docking & Deamination followed by sequencing (D&D-seq) technique induces cytosine- to-uracil edits in genomic regions bound by the target protein, offering a novel means to capture DNA:Protein interactions with unprecedented resolution. Importantly, this technique can be seamlessly incorporated into common single-cell multiomics workflows, enabling multimodal analysis of gene regulation in single cells. In this study, we used ATAC-seq and scATAC-seq to define open chromatin regions as a functional readout for transcription factors activity. In addition, we coupled D&D-seq with whole genome sequencing (WGS), to capture CTCF binding events in both active and inactive chromatin compartments.
[0543] We tested the ability of D&D-seq to record 6 different TFs and 1 chromatin remodeling factor binding in bulk, we then validated D&D-seq at the single-cell level by profiling CTCF and GATA family members, obtaining high specificity and efficiency, with clear identification of TF footprint and signal retention in the targeted cell subpopulations. Furthermore, the deamination reaction showed minimal off-target activity, with high concordance to bulk ChlP-seq reference data. Applied to primary human peripheral blood mononuclear cells (PBMCs), D&D-seq successfully identified CTCF binding sites and enabled integration with advanced machine -learning algorithms for predicting 3D chromatin structure. Furthermore, we integrated D&D-seq with single-cell genotyping to assess the impact of IDH2 mutations on CTCF binding in a human clonal hematopoiesis sample, uncovering altered binding and chromatin co-accessibility patterns in mutant cells. Altogether, D&D-seq represents an important technological advance enabling the direct mapping of TF or chromatin remodeler binding to the DNA in primary human samples, opening new avenues for understanding chromatin and transcriptional regulation in health and disease.
[0544] Profiling of TF binding patterns in single cells in primary samples has been mainly restricted to inferential approaches based on expression levels of key downstream TF target genes10 12or through motif analysis of ATAC-seq peaks13,14. While analysis of ATAC-seq peaks can be highly informative about the general chromatin landscape, for confident identification of specific TF binding sites, more direct methods are recommended14,15. Combining computational TF binding inference analysis with single-cell multi-omic approaches that enable simultaneous capture of chromatin accessibility and gene expression within the same single cell allows for the characterization of genomic regulation across multiple molecular layers (from chromatin landscape to RNA)16. Incorporating direct TF binding measurements into such a single-cell framework would provide an unprecedented window into the exact mechanisms driving genome function, transcription networks and pathway regulation. Moreover, this multifaceted analysis spanning different molecular modalities as a readout of direct TF binding would accelerate the identification of key factors that mediate normal cellular activity, as well as those that are disrupted during disease.
[0545] To address these limitations, we present a new method to accurately map DNA:Protein interactions in low-input samples or single cells that is amenable for direct application to primary human samples. To directly profile TF or chromatin factor binding in single cells, we fused the base-editing deaminase DddA that catalyzes C-to-U singlenucleotide changes to secondary nanobodies targeting an antibody-labeled DNA binding protein of interest, amenable to integration into single-cell genome accessibility workflows such as microfluidic-based scATAC-seq, combinatorial barcoding Paired-seq17or Share- seq16, or whole-genome sequencing approaches, such as DLP+18or Primary Template Amplification19.
[0546] As a proof of concept, we incorporated D&D into the lOx Genomics sc AT AC- seq workflow, enabling the capture of accessible regions of the genome where most gene regulation occurs5,20. Upon genome-wide antibody binding to the TF, nanobody-tethered base editor enzyme deaminates proximal cytosines to uracil, leaving a permanent genomic signature at regions of TF binding that can be identified through downstream single-cell sequencing of the tagmented loci. Compatibility with existing droplet-based single-cell sequencing workflows allow for broad use of Docking and Deamination followed by sequencing (D&D-seq) across systems. We demonstrated the sensitivity and specificity of the D&D enzyme in a deamination assay using DNA oligos or phage DNA, and perform bulk experiments to obtain genome- wide TF binding profiles showing concordance with ChlP-seq data. We tested the robustness of single-cell D&D-seq through cell- line mixing experiments, identifying canonical GATA1 binding patterns in erythroid cells. We further applied D&D-seq to primary human peripheral blood mononuclear cells (PBMCs), recovering high enrichment of CTCF binding motifs coinciding with D&D edits. In addition, utilizing single-cell D&D and ATAC data, the D&D-C. Origami pipeline successfully predicted 3D chromatin structure in human primary samples across different cell subtypes. Finally, we integrated single-cell genotyping21into the D&D-seq framework to simultaneously capture genotypes together with TF binding and chromatin accessibility from the same single cell, profiling CTCF binding in an / D / / 2- mutant clonal hematopoiesis of indeterminate potential (CHIP) sample. We envisioned that the experimental tools and computational analyses developed here could be broadly applied for the interrogation of the interplay between TF or chromatin remodeler binding and chromatin landscapes, as well as defining the impact of somatic mutations on TF binding in single cells, directly from primary samples. The versatility of D&D-seq allows for further integration with existing or newly developed single-cell multiomics methods to expand the profiling of downstream modalities including gene expression, opening new avenues for the study of gene regulation in both health and disease.
[0547] Example 10: Docking & Deamination (D&D-seq) allows for mapping Protei DNA binding via molecular foot-printing CUT&Tag-derived approaches22’23, including NTT-seq24and scCUT&Tag-Pro25, are suitable for profiling stable and abundant chromatin binders, such as histones, in single cells. These methods are based on CUT&Tag26, an approach that generates a library from DNA fragments colocalizing with a chromatin factor of interest by fragmenting out DNA surrounding the antibody-bound chromatin factor. This process is facilitated by pA-Tn5 fusion protein, where protein A (pA) has an affinity to antibodies and Tn5 is a transposase that fragments and tags DNA with sequencing-ready adapters. As Tn5 binds to DNA with high affinity, pA-Tn5 staining and tagmentation are performed under high stringency conditions (e.g., 300-500 mM NaCl) to prevent signal from open chromatin regions mediated by direct binding of Tn5 to DNA. However, these stringent conditions also limit the antibody-tethering activity of the pA-Tn5, resulting in inefficient tagmentation at the chromatin factor-bound sites. In addition, high-salt conditions disrupt weak interactions, such as those between DNA and TFs27, or prevent binding of antibodies with low affinity for their targets. As such, these methods are limited in their ability to profile weak DNA:Protein interactions, including those involving TFs and their DNA binding sites.
[0548] To address this challenge, we developed D&D-seq, a new method for mapping TF binding in single cells. As an alternative approach to tagmentation plus barcoding of fragments at the genomic location bound by the target protein, we reasoned that binding patterns of target proteins can be captured by tethering a base-editing enzyme to a target protein in a precise manner. Briefly, this enzyme is the fusion of the double- stranded DNA deaminase DddA28,29with secondary nanobodies, or other IgG binding protein such as protein A / G, that can recognize FC domains of immunoglobulins, including rabbit, mouse, goat or rat (referred to as nb-DddA). The fusion protein is thus capable of binding to primary antibodies and catalyzing cytosine deamination in the vicinity of the site that is bound by the targeted TF or chromatin remodeler. This deamination event results in the conversion of cytosine to uracil on the genomic DNA, which can be then identified through sequencing, providing a molecular footprint of the target of interest on the genomic DNA. To showcase the versatility of this approach, we chose to integrate it with the standard ATAC-seq workflow, which provides high-resolution insights into accessible regions of the genome, where most transcription factors bind5, making ATAC-seq an ideal method for capturing active genomic regions while simultaneously recording TF molecular footprints (Fig. 39a). To design the fusion enzyme, we selected engineered DddAl 1 that has increased activity and more versatile DNA editing context compared to the original DddA29. To achieve a switch-like control over enzyme activity, we use a split enzyme design, where the deaminase is separated into two polypeptides (Fig. 39b, Fig. 43a, 43b). This allows for the DddA enzyme to be maintained in an inactive state upon DNA binding, avoiding nonspecific deamination due to random interaction of the DddA enzyme with the genomic DNA. The activity of the enzyme can be controlled by the addition of the C-terminal small peptide to the reaction, which through heterodimerization reconstitutes the split enzyme and its catalytic activity28-30. We chose the split site with the longest N terminus and linked the nanobodies with the N terminal protein (nb-DddA_NT) so that the C terminal peptide (DddA_CT) is 25 amino acids long (Fig. 39b), short enough to be synthesized. We truncated the last 5 amino acids of DddA_CT in the original construct because it is a structurally disordered region (PDB 6u08), and was shown to be functionally redundant by Yin et al30.
[0549] To first evaluate enzyme base editing specificity, we assessed the deaminase activity of DddA across nucleotide contexts. We tested recombinant nb-DddA_NT activity using a deamination assay. This assay showed that nb-DddA_NT is only active when DddA_CT is present (Fig. 39c, Fig. 43c, 43d). Furthermore, we showed that the addition of Zn2+ increases the deaminase activity (Fig. 39c, Fig. 43c, 43d). Consistent with previous studies29,30, we observed that nb-DddA preferentially deaminates cytosines that are in the TpC context (Fig. 39d, Fig. 43e). To assess enzymatic activity, we tested the activity of nb- DddA at different temperatures and incubation times (Fig. 43f, 43g, 43h), determining that the activity was greater at 37 °C and 50°C compared to 30°C and 4°C. Together, these results show that our D&D-deaminase recapitulates the base-editing activity of the unmodified enzyme with high fidelity.
[0550] For bulk D&D-seq, isolated cells are first fixed and permeabilized to allow all the reagents to enter the cell. Then, the sample is incubated with a primary antibody specific to the targeted DNA-binding protein. After washing the unbound antibody, the sample is incubated with nb-DddA_NT. This allows for the deaminase to localize in close proximity with the target. After washing the nb-DddA_NT in excess, we activate the deaminase by incubating the samples with DddA_CT and Zn2+. Followed by washing, the sample can be processed with standard ATAC-seq methods. During the ATAC-seq protocol, genomic DNA is exposed to a highly active transposase (Tn5 or TnY). The transposase simultaneously fragments DNA, preferentially inserts into open chromatin sites, adding the adaptors compatible with downstream amplification and sequencing. Open chromatin is then identified from the sequenced DNA and data analysis can provide insight into gene regulation using the multimodal data. In this way, we profile accessible chromatin and at the same time record the protein binding events on accessible genomic DNA (Fig. 39a).
[0551] To evaluate nb-DddA activity in cells, we next tested nb-DddA in a bulk D&D-seq experiment targeting either CTCF, GATA1 or GATA2 transcription factors in K562 cells. We identified an increased C>T (or G>A) mutation signature in DNA compared to other mutations or to a no-deaminase ATAC control (Fig. 39e, 39f, 39g, Fig. 43i, 43j, 43k). We further characterized the dinucleotide and trinucleotide context in all mutated sites and found that nb-DddA deaminates not only in a TC context, but also in CC and GC contexts (Fig. 431, 43m). We reasoned that this is due to high local concentration of DddA tethered around the targeted TFs. A deamination footprint analysis of edit events in a 200 base-pair window around all the HOCOMOCO-defined31CTCF, GATA1 or GATA2 binding sites in the genome showed the expected bimodal distribution of the edit events, where the binding sites for the targeted TF are at the center of the two modes, with low editing frequency in background accessible regions lacking binding sites (Fig. 39h, 39i, 39j). To confirm that the deaminase signal is TF-site-specific, we performed motif enrichment analysis using peaks called from reads with DddA-specific edits. Motif enrichment analysis showed strong CTCF and GATA motifs in the respective peaks (Fig. 39k, 391, 39m). By ranking the motifs identified by de novo discovery tools, we observed that the CTCF and GATA motifs were consistently enriched in the top motif identified. They were only preceded by the Tn5 consensus, which is expected due to open chromatin tagmentation. These results further confirm the high specificity of the D&D reaction (Fig. 39n, 39o, 39p).
[0552] Furthermore, we used ENCODE20CTCF or GATA1 ChlP-seq data from K562 cells to visualize on-target and off-target peaks, and identified significantly higher deaminase activity in on-target peaks compared to the off-target peaks or to the background (Fig. 39q). Motivated by these results, we explored the possibility of profiling transcription factor binding sites using D&D-seq beyond those regions of accessible chromatin. To identify all possible binding events, even those that occur outside open chromatin, we coupled D&D- seq with whole-genome sequencing. Briefly, the protocol followed the same approach as in our previous experiments, except that after the CTCF-specific deamination reaction, we processed the sample with the Primary Template Amplification (PTA) protocol to amplify and obtain a sequencing library from the entire genome. The analysis of the entire genome confirmed the expected karyotype of K562 cells, with the specific copy number alterations previously observed in this cell line (Fig. 39r). This indicates that the D&D workflow does not affect the quality of whole genome amplification. More importantly, we were able to identify 2045 peaks outside open chromatin regions, expanding the applicability of our protocol to more complex scenarios where binding occurs in regions that don’t overlap AT AC assays. Indeed, by extracting reads from the entire genome library that exhibited D&D signal, we were able to generate CTCF peak tracks that were comparable to high- quality ENCODE ChlP-seq, regardless of the accessibility of the peaks. To validate that D&D is not restricted to open chromatin, we examined histone post-translational modification patterns of non-ATAC D&D peaks. As expected, we found a significant subset of these inaccessible peaks overlapped with heterochromatic marks, such as H3K9me3. This is consistent with prior reports showing that CTCF binds inactive regions of the chromatin to help insulate domains and prevent the spread of heterochromatin (citation in slack: available at world wide web sciencedirect.com / science / article / pii / S2589004223001839) (Fig. 39s). To further validate our findings, we performed deamination footprint analysis of edit events within a 200-base- pair window encompassing all known CTCF binding sites, dividing them into accessible and non-accessible regions. Remarkably, the footprints in both accessible and non- accessible regions were significantly higher than those observed in CTCF-negative regions (Fig. 39t).
[0553] Having validated our protocol on TFs with well-characterized and well-defined binding motifs, we next sought to evaluate the performance of D&D-seq on chromatin remodeling factors, which typically exhibit more transient interactions and lack a specific consensus DNA-binding sequence. As a representative example, we focused on p300, a histone acetyltransferase and transcriptional co-activator that localizes to active enhancers and plays a central role in chromatin remodeling. Despite the dynamic and indirect nature of p300-DNA interactions, D&D-seq successfully captured robust deamination footprints around p300-bound regions (Fig. 43n, 43o). Notably, the observed footprints were broader than those detected for sequence-specific transcription factors such as CTCF, consistent with the less sharply defined binding architecture of chromatin remodeling complexes. Using high-quality ENCODE p300 ChlP-seq data from K562 cells as a reference, we observed strong enrichment of deamination signal across ChlP-defined peaks, demonstrating the ability of D&D-seq to detect diffuse occupancy patterns associated with non-DNA-binding regulatory factors (Fig. 43o). As an additional layer of validation, we performed a de novo motif analysis on the D&D-seq-positive peaks identified at p300- bound regions (Fig. 43p). Given that p300 does not bind DNA directly but is recruited through interactions with sequence- specific transcription coactivators, we specifically searched for motifs corresponding to known p300 co-occupiers. Remarkably, all enriched motifs matched previously reported recognition sequences of transcription factors that physically or functionally interact with p300 (Fig. 43q). This finding reinforces the utility of D&D-seq for profiling chromatin remodeling factors by capturing their genomic recruitment sites through indirect binding partners.
[0554] Together, these results establish D&D-seq as a versatile method for profiling transcription factor binding, yielding high-resolution and highly specific footprints for CTCF, GATA1, and GATA2, with additional evidence of performance for JUN, SP1, and TALI (Fig. 43r, 43s). By coupling D&D-seq with whole-genome sequencing, we extended its applicability beyond open chromatin regions, enabling the identification of binding events in inaccessible genomic contexts. Finally, we demonstrated the feasibility of using D&D-seq to map chromatin remodeling factors such as p300, capturing their broader and less sequence-specific binding patterns.
[0555] Example 11 D&D-seq profiles single-cell DNAtProtein interaction
[0556] We next tested the ability of D&D-seq to record TF binding in single cells when incorporated into the lOx Genomics single-cell (sc) ATAC-seq workflow. We performed cell line mixing experiments, selecting two distinct cell lines (lymphocyte cell line CA46 and erythroblast cell line K562) and two TF targets (CTCF and GATA1) to allow us to assess the targeting specificity and potential cross-contamination of signal in D&D-seq.
[0557] The two cell lines were individually crosslinked, permeabilized, and stained with a CTCF antibody for CA46 and a GAT Al antibody for K562. Cells were then washed, equally mixed and incubated in the D&D-seq buffer containing the activating C-terminal peptide of DddA and the Zn2+ cofactor necessary for the activation of the deamination reaction. In this manner, we were able to record the presence of CTCF or GATA1 on the genomic DNA of each of the tested cell lines. After this step, the sample was processed with custom or commercially available single-cell AT AC workflows with minimal modifications to the original protocol.
[0558] I l l We profiled K562 (n = 1,732 cells) and CA46 (n = 4,108 cells) (Fig. 40a), obtaining 5,263 + / - 2,633 (mean + / - standard deviation) fragments per cell for K562 and 5,437 + / - 2,588 (mean + / - standard deviation) fragments per cell for CA46 (Fig. 40b), reflecting high quality tagmentation data (Fig. 44a). Cell lines were identified by gene accessibility score of known marker genes (Fig. 40c). We projected cells into a low-dimensional space using latent semantic indexing (LSI) and uniform manifold approximation and projection (UMAP) using Seurat, and clustered cells using a shared nearest neighbor (SNN) approach. There were two clear clusters corresponding to the two cell types showing good separation between CA46 and K562 cells, as defined by highly accessible markers for lymphoid (Fig. 40d) and erythroid cells (Fig. 40e), demonstrating that the D&D-seq reaction does not affect the overall performance of the scATAC.
[0559] Importantly, at the single-cell level, we were able to map protein binding events in all the cells analyzed, suggesting that the deamination reaction has a relatively high efficiency and specificity (Fig. 44b, Methods). While the absolute number of edits per cell is relatively low (Fig. 40f, 40g, 40h, 40i), this is comparable to established tagmentationbased single-cell transcription factor binding inference methods. Furthermore, the data can be effectively integrated through pseudobulk or meta-cell analysis, enabling robust downstream interpretations. Since the antibody staining step was performed before mixing the two cell lines, we were able to assess the presence of cross-contamination potentially occurring during droplet encapsulation, barcoding, and library preparation reactions. Our results showed that the signals for CTCF and GATA1 were mutually exclusive and were retained exclusively in the subpopulation stained with the respective specific antibody, demonstrating that the specificity of the signal was maintained at the cellular level (Fig. 40f, 40g, 40h, 40i).
[0560] To provide support for the specificity of D&D-seq, we performed de novo motif discovery analysis on the pseudobulk D&D-seq mutational peaks (Fig. 40j, 40k, 401, 40m). We observed significant enrichment of the CTCF binding site (e-value 1 ,3e- 171) in CA46 cells (Fig. 40j) and the GATA1 binding site (e-value 2.1e-89) in K562 cells (Fig. 401). We next orthogonally validated our findings using a motif-centric, rather than de novo, approach. Here, we used the total reads obtained from the single-cell experiment to calculate the frequency of D&D events in a 200 base-pair window around all the HOCOMOCO-defined31CTCF or GATA binding sites in the genome, and compared this with the frequency of C-to-T events observed in all the accessible peaks measured in our experiment. As per the bulk experiment, the single-cell pseudobulk analysis revealed the presence of the expected bimodal distribution of the deamination events (Fig. 40k, 40m, top). However, when the same analysis was performed for accessible regions lacking known binding sites, we observed negligible edit event frequency and the absence of clearly-defined distribution patterns (Fig. 40k, 40m, bottom). Moreover, when we repeated this analysis but for the non-targeted factor for each cell type, we observed no GATA1 binding signal in CA46 cells and no CTCF binding signal in K562 cells (Fig. 44c). These results demonstrated across two TFs that our D&D-seq approach faithfully recapitulated TF binding at the single-cell level.
[0561] We next sought to evaluate the specificity of the deamination reaction at the site of target protein binding. To do so, we collapsed the reads from a specific cluster to obtain pseudobulk alignments. We next extracted the reads where we identified C-to-T conversion events and generated a D&D coverage plot. The peaks obtained with our in situ footprinting approach showed high concordance with bulk ChlP-seq reference data obtained from the ENCODE consortium20(Fig. 40n), suggesting that base editing events occur only in the proximity of the targets of interest, with minimal off-target activity. We quantified the similarity between our single-cell D&D-seq reads and those from ENCODE, showing high correlation between CTCF signals from each data type and between GATA1 signals for each data type, but low correlation across the CTCF and GATA1 signals (Fig. 40o). We note that the correlation values we observed between our single-cell D&D-seq reads and those from ENCODE were comparable to the results obtained using well-established technologies on histones, such as CUT&Tag on H3K27me3 or H3K4me326, which is often used as a benchmarking target due to its high abundance and high-quality results (Fig. 44d).
[0562] The lack of technologies with similar capabilities, especially for single-cell applications, limits direct benchmarking of D&D-seq with the current state-of-the-art methods for profiling DNA:Protein interactions. Despite this gap, attempts have been made to profile the chromatin occupancy of non-histone proteins in low input samples by modifying well-established bulk protocols. Among those, ultra-low input cleavage under targets and release using nuclease (uliCUT&RUN) is a variant of CUT&RUN, with key modifications to reduce background signal, increase output and decrease the amount of starting material required to generate protein occupancy profiles from mammalian cells32. This protocol was used to profile the genomic locations of the insulator protein CTCF from populations of mouse embryonic stem cells (mESCs) ranging in number from 500,000 to IO33. In order to provide further benchmarking for D&D-seq, we iteratively randomly subsampled (n=100) the single-cell CA46 subpopulation, stained with the CTCF antibody, to generate matching datasets with the same number of cells analyzed by uliCUT&RUN. As negative control, the same number of K562 cells, which represent cells that were not stained with CTCF antibody, were subsampled to match the uliCUT&RUN negative control, cells that were not stained with antibody. To evaluate the specificity of the two protocols, we calculated the fraction of the reads in the peaks (FRiP), using a reference peak annotation obtained from a high-quality ChlP-seq. A high fraction of reads in peaks indicates that the majority of the reads are located in the regions of interest, and that the experiment has a high signal-to-noise ratio. On the other hand, a low fraction of reads in peaks may indicate that the majority of the reads are located in non-specific regions, and that the experiment has a low specificity. D&D-seq data showed that the majority of the base-edited reads map to regions where CTCF was previously mapped, consistently reaching -60% FRiP across analyses. Moreover, the signal-to-noise ratio was largely unaffected by the number of the cells analyzed, showing similar values for the 5,000-, 500-, 50-, and 10-cell sample sizes, with mean FRiP equal to 56.75%, 56.78%, 56.83%, and 57.00% respectively (Fig. 40p, top). As expected, the cell line that was not stained with the CTCF antibody showed very low values of FRiP in all the input conditions, indicating that the majority of the edited reads were not specific for CTCF. Importantly, at least for this specific metric, D&D-seq outperforms uliCUT&RUN, which scored 13.69%, 13.69%, 13.89%, and 14.04% FRiP for the 5000-, 500-, 50-, and 10-cell sample sizes, reflecting a majority of reads mapped to regions of the genome that are not bound by CTCF (Fig. 40p, bottom).
[0563] While D&D-seq demonstrates high specificity in capturing transcription factor binding sites even at low cell numbers, global detection of aggregated footprint signals can be impacted by the sparsity of scATAC-seq data and base editing signal. For instance, read support for D&D edit peaks diminishes when fewer than 100 cells are sampled, as seen for both CTCF and GATA1 (Fig. 44e, 44f). To systematically assess the minimal input required for reliable footprint detection, we performed down-sampling analyses on the CA46 and K562 datasets, ranging from 50 to 4,000 cells and 50 to 1,500 cells respectively, with 100 replicates per condition (Fig. 44g, 44h). These benchmarks revealed that robust aggregate footprint signals consistently emerge with sample sizes of 250 cells or greater, whereas signal detection becomes limited below this threshold due to a reduced number of confidently called peaks (Fig. 44g, 44h, 44i). These findings provide practical guidance for experimental design and highlight the sensitivity of D&D-seq in capturing regulatory footprints at modest cell numbers.
[0564] In conclusion, these data confirm the specificity of the scD&D-seq signal, indicating that our method performs at a level comparable to the current state-of-the-art approaches. However, beyond achieving similar accuracy, D&D-seq offers significant advantages, as it is capable of simultaneously capturing a wider range of biological features at single-cell resolution, thus enabling a deeper exploration of complex regulatory mechanisms that were previously inaccessible.
[0565] Example 12: Single-cell DNA:Protein interaction profiling in primary human cells
[0566] We tested D&D-seq for analysis of primary human PBMCs. We selected CTCF as an initial target due to its ubiquitous presence on the genome in all cell types and cell states, and the presence of a defined consensus DNA binding sequence34. For this experiment, mobilized PBMCs collected from healthy donors were crosslinked, permeabilized, and stained with a CTCF-specific polyclonal antibody. The sample was then processed with the D&D reaction, followed by standard lOx Genomics scATAC-seq with minimal modification (see Methods in Example 8).
[0567] We obtained high-quality single cells (n = 5,358 cells), with 15,534 + / - 5,934 (mean + / - standard deviation) fragments per cell. We projected cells into a low-dimensional space using LSI and UMAP, clustering cells using the chromatin accessibility data (Fig. 41a). We assigned cell states based on chromatin accessibility profiles of established marker genes and retrieved all the expected subpopulations, demonstrating that the addition of the D&D reaction does not impact the overall quality of the chromatin accessibility assay (Fig. 45a, 45b). Critically, we mapped CTCF binding sites in this complex primary tissue (Fig. 41b, Fig. 45c). We extracted the reads with in situ C-to-T transition labeling and identified CTCF binding sites in close proximity to C-to-T deamination events (Fig. 41c, 41d). These results demonstrate the direct applicability of D&D-seq in primary human samples.
[0568] We next hypothesized that the joint availability of chromatin accessibility and CTCF binding information from the same single cells could support machine-leaming- based approaches originally developed to predict 3D chromatin structure from bulk data, potentially extending their application to single-cell genomics. In particular, we explored the use of the C.Origami pipeline, which integrates chromatin accessibility with CTCF ChlP-seq data to train a deep neural network for de novo prediction of cell-type-specific 3D genome organization. Leveraging our D&D-seq protocol, we generated CTCF binding profiles that served as input features for C.Origami, together with matched single-cell ATAC-seq data, to produce pseudobulk Hi-C-like contact maps. Although C.Origami was originally optimized for bulk genomic datasets, we were encouraged to observe that its application to single-cell-derived input data (scATAC-seq and scD&D-seq) generated contact matrices that qualitatively resembled publicly available Hi-C maps. Notably, both short- and long-range chromatin interactions were recovered in human primary CD8 T cells (Fig. 41e). To further explore the potential of this approach, we used C.Origami to predict 3D chromatin structure across 100 random 2-Mb genomic regions in major immune cell types including B cells and monocytes (Fig. 41f, Fig. 45d, 45e, 45f, 45g), comparing results from CTCF ChlP-seq and D&D-seq inputs. Predictions generated without CTCF input served as a negative control. We observed that predictions using scD&D-seq data yielded comparable performance to those using ChlP-seq, suggesting the feasibility of this strategy (Fig. 41f, Fig. 45e, 45g). However, we note that these findings are preliminary and that further optimization of algorithms specifically tailored for single-cell data may be necessary to improve prediction accuracy and robustness. Our results thus highlight a promising opportunity to adapt and refine bulk-trained models like C.Origami for singlecell applications, potentially enabling more precise inference of chromatin architecture during dynamic cellular processes such as differentiation and aging.
[0569] Example 13: Linking genotype to DNA:Protein interaction profiles and chromatin landscapes in single cells
[0570] Given the compatibility of the D&D workflow with high-throughput single-cell platforms such as lOx Genomics, it is possible to incorporate this versatile backbone technology into other single-cell multiomics frameworks that capture multiple modalities. To this end, we reasoned that integrating single-cell genotyping capability would enable the extension of D&D-seq to study the effects of somatic variants on gene regulation. Given the mixture of mutant and wild-type cells within clonally mosaic tissues, single-cell multimodal methods that capture both genotype and phenotype together are required to provide biological insights into mechanisms of clonal outgrowth in normal, aging or diseased tissues36. To address this challenge, we previously developed Genotyping of Transcriptomes (GoT)37and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA)21to enable high-throughput single-cell genotyping together with transcriptome profiling or chromatin accessibility profiles. Leveraging this conceptual framework for linking single-cell genotype to molecular phenotype, we incorporated our D&D-seq approach into single-cell genotype- aware multi-omics, to link somatic mutations to changes in TF binding, and applied this workflow to profile a human CHIP sample.
[0571] Briefly, in GoT-ChA, isolated cells are subjected to transposition of genomic DNA (gDNA) and loaded into microfluidics devices (Fig. 46a, steps 1-4). During cell barcoding reactions, additional gene-specific primers are added to capture a locus of interest with an in-droplet PCR reaction, with a handle that is compatible with the lOx scATAC platform (Fig. 46a, step 4). The product is then split, with an aliquot used for an amplicon genotype library and the remaining 90% used for scATAC library construction and sequencing (Fig. 46a, steps 5-6). The scATAC and genotyping libraries can then be analytically integrated via shared cell barcodes, linking chromatin accessibility to targeted genotyping at singlecell resolution. To incorporate genotyping into the D&D-seq framework, GoT-ChA primers are added during the scATAC in-drop PCR step.
[0572] To test D&D-GoT-ChA, we profiled CTCF binding in PBMCs from a patient with CHIP carrying an IDH2RI40Qmutation. The mutation was identified by targeted panel sequencing and had a variant allele frequency of 0.15. Integration of GoT-ChA with D&D- seq allowed us to build a multiomic dataset that includes the genotype of individual cells, along with accessible chromatin status and CTCF binding pattern.
[0573] The simultaneous implementation of targeted genotyping and D&D-deamination did not affect the quality of the chromatin accessibility data, which we used to resolve the tissue into highly granular cell populations, identifying all the major cell subtypes expected for this sample (Fig. 42a). The experiment was performed in two technical replicates, and after quality check and batch correction, the data from the individual experiments were merged into a single dataset of 15,807 high-quality cells. The resulting dataset showed appropriate QC metrics, with a mean of 7,592 fragments per cell (+ / - 3,971, standard deviation), enrichment around transcriptional start sites, and expected library size distribution (Fig. 46b). Accessibility in proximity of key marker genes was used to guide cell annotation, allowing us to identify all the major cell types (Fig. 46c).
[0574] The genotyping library obtained by GoT-ChA targeted amplification was sequenced and processed following our previously published pipeline (available at github.com / landau- lab / Gotcha)21. We successfully genotyped 33.59% of single cells (5,291 out of 15,807 cells), identifying IDH2RI40Q-muta.nt and wild-type cells. Genotyping data were assigned to individual cells and projected into the accessibility latent space to visualize the distribution of wild-type and mutant cells in the tissue (Fig. 42b). This analysis revealed an enrichment of mutant cells in the CD8 T cell subcluster, showing that the IDH2 mutation was almost exclusively found in this specific cell type (Fig. 42c, Fig. 46d). To validate this finding, peripheral blood cells from the same donor were immunolabeled and split into the main subpopulation by FACS (Methods, Supplementary Note). Genomic DNA was isolated from natural killer (NK), monocytes, CD8 T cells, CD4 T cells, and B cells and processed for standard genotyping by bulk nanopore sequencing. The bulk genotyping data matched the single-cell genotyping data obtained with D&D-GoT-ChA (Fig. 42d), further validating the precision of our single-cell approach.
[0575] Next, we analyzed the CTCF binding pattern across cells. The CTCF molecular footprinting signal was present in all the single-cell clusters as expected from this factor (Fig. 42e, 42f, Fig. 46e). We identified 5,631 CTCF positive peaks, and the motif enrichment analysis of the footprinted reads again revealed the specificity of our deamination assay, showing a significant enrichment of CTCF binding sites in the proximity of editing events (Fig. 42g, Fig. 46f, 46g). To evaluate whether D&D-seq can resolve cell type-specific CTCF binding, we performed differential analysis of CTCF edit signals across PBMCs main cell subpopulations. This analysis successfully identified a set of differentially enriched CTCF binding sites that were specific for each individual cell type (Fig. 46h). Notably, several monocyte-specific peaks were observed at promoters of genes such as MAFB, GNAI2, and NECTIN2, all of which were selectively expressed in the monocyte lineage. For example, MAFB encodes a bZIP transcription factor critical for myeloid differentiation, and CTCF binding at its promoter may reflect lineage- specific chromatin organization, potentially through the formation of regulatory loops.
[0576] IDH-mutant cells have high levels of 2-hydroxyglutarate (2-HG) that interfere with the TET family of 5 '-methylcytosine hydroxylases38. TET enzymes catalyze a key step in the removal of DNA methylation39. Previous in vitro studies have shown that IDH-mutant cells exhibit hyper-methylation at CTCF binding sites, compromising the binding of this methylation- sensitive insulator protein40,41. Reduced CTCF binding is associated with loss of insulation between topological domains and aberrant gene activation40, although the extent and functional significance of IDH-mutant mediated alterations of epigenetic states remain unclear in vivo. Our D&D technology allows us, for the first time, to address this question by analyzing differences in CTCF binding between IDH wild-type and mutant cells in primary human cells. As the vast majority of mutant cells were detected in CD8 T cells, we focused our analyses on this specific cell type (Fig. 42b, 42h).
[0577] First, we observed that CD8 T cells form two distinct clusters characterized by the differential presence of wild-type and mutant cells. Therefore, we reassigned the identity of these two clusters as CD8 wild-type-enriched and CD8 mutant-enriched clusters to assess IDH2-driven phenotypic changes (Fig. 42h, Fig. 46i). We assessed CTCF binding in CD8 T cells by comparing the CTCF D&D binding signals (C>T edit read counts) between IDH2 wild type (WT) and mutant (IDH2R140Q, MUT). Notably, the vast majority of statistically significant binding sites differentially bound by CTCF exhibit a decrease in the CTCF binding signal between mutant and wild-type cells (Fig. 42i, Fig. 46i ), which corroborates previous reports suggesting that CTCF binding is reduced in / DH- mutant glioma and acute myeloid leukemia, mediated by DNA hypermethylation at CTCF binding sites40,41.
[0578] For example, the binding site in close proximity to the GIT1 gene showed the strongest reduction in binding. GPCR-kinase interacting protein 1 (GIT1) is a scaffold protein that interacts with proteins such as RAC1, PAK, and paxillin to regulate actin cytoskeleton remodeling42. This is crucial for T cell migration, synapse formation, and activation43,44. Furthermore, this protein has been implicated in modulating T cell-mediated inflammatory responses45. To further explore the functional consequences of CTCF loss in this genomic domain, we took advantage of our ability to simultaneously profile and integrate differential CTCF binding, chromatin accessibility, and genotyping to identify changes in cis-regulatory DNA interactions affecting the GIT1 topologically associated domain (TAD) that may be a consequence of IDH2 mutation.
[0579] To evaluate the potential alterations in 3D genomic interactions between mutant and wild-type cells, we once again utilized the C. O...
Claims
What is claimed is:
1. A fusion protein comprising a base-editing enzyme and a protein that binds to a fragment crystallizable region (Fc region) of an immunoglobulin.
2. The fusion protein of claim 2, wherein the Fc region is an Fc region from IgGl, IgG2, IgG3, or IgG4.
3. The fusion protein of claim 1 or 2, wherein the protein that binds to a Fc region of an immunoglobulin comprises an antibody or antigen-binding fragment thereof, a protein G, and / or a protein A.
4. The fusion protein of any one of claims 1-3, wherein the protein that binds to a Fc region of an antibody is or comprises an antibody or antigen-binding fragment thereof.
5. The fusion protein of claim 4, wherein antigen -binding fragment thereof is selected from a Fab, a single chain variable region (scFv), a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody.
6. The fusion protein of claim 5, wherein the antigen-binding fragment thereof is a nanobody.
7. The fusion protein of claim 6, wherein the nanobody recognizes a rabbit, mouse, goat or rat immunoglobulin.
8. A fusion protein comprising a base-editing enzyme and a protein that binds to a DNA binding protein.
9. The fusion protein of claim 8, wherein the protein that binds to a DNA binding protein is an antibody or antigen-binding fragment thereof.
10. The fusion protein of claim 9, wherein antigen-binding fragment thereof is selected from a Fab, a single chain variable region (scFv), a nanobody, a Fv, a F(ab’)2, a Fab’, a dsFv, a sc(Fv)2, a nanobody, or a diabody.
11. The fusion protein of claim 10, wherein antigen-binding fragment thereof is a nanobody.
12. The fusion protein of any one of claims 1-11, wherein the nanobody comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 1-4.
13. The fusion protein of claim 12, wherein the nanobody comprises the amino acid sequence of any one of SEQ ID NOs: 1-4.
14. The fusion protein of any one of claims 1-3 wherein the protein that binds to a Fc region of an immunoglobulin comprises a protein G, optionally wherein the protein is a protein G or a fusion protein comprising a protein G.
15. The fusion protein of claim 14, wherein the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to an amino acid sequence of SEQ ID NO: 16 or 20.
16. The fusion protein of claim 15, wherein the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence of SEQ ID NO: 16 or 20.
17. The fusion protein of any one of claims 1-3, wherein the protein that binds to a Fc region of an immunoglobulin comprises a protein A, optionally wherein the protein is a protein A or a fusion protein comprising a protein A.
18. The fusion protein of claim 17, wherein the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence at least 95% homology to an amino acid sequence of SEQ ID NO: 5 or 20.
19. The fusion protein of claim 18, wherein the protein that binds to a Fc region of an immunoglobulin comprises an amino acid sequence of SEQ ID NO: 5 or 20.
20. The fusion protein of any one of claims 1-19, wherein the base-editing enzyme is a base-editing deaminase.
21. The fusion protein of claim 20, wherein the base-editing deaminase deaminates cytosine to uracil (C>U).
22. The fusion protein of claim 21, wherein the base-editing enzyme is selected from DddA, CbDaOl, AcDaOl, non-DddA-like OU dsDNA deaminases, or DddA-like.
23. The fusion protein of claim 22, wherein the DddA base-editing enzyme is DddAl 1.
24. The fusion protein of claim 22, wherein the DddA-like base-editing enzyme is MGYPDa829 or LbsDaOl.
25. The fusion protein of any one of claims 1-24, wherein the base-editing enzyme comprises an amino acid sequence at least 95% homology to any one of SEQ ID NOs: 6-8, optionally wherein the base-editing enzyme comprises an amino acid sequence selected from SEQ ID NOs: 6-8.
26. The fusion protein of any one of claims 1-25, wherein the base-editing enzyme is a non- split full length base editor.
27. The fusion protein of any one of claims 1-25, wherein the base-editing enzyme is a split base editor.
28. The fusion protein of claim 27, wherein the base-editing enzyme is a DddA split base editor.
29. The fusion protein of claim 28, wherein the DddA split base editor comprises an N- terminal portion of DddA (DddA_NT).
30. The fusion protein of claim 29, wherein the N-terminal portion of DddA (DddA_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 9.
31. The fusion protein of claim 30, wherein the N-terminal portion of DddA (DddA_NT) comprises the amino acid sequence of SEQ ID NO: 9.
32. The fusion protein of any one of claims 28-31, wherein the DddA split base editor is activated by the addition of a C-terminal portion of DddA (DddA_CT).
33. The fusion protein of claim 32, wherein the C-terminal peptide comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 10 or 17.
34. The fusion protein of claim 33, wherein the C-terminal peptide comprises the amino acid sequence of SEQ ID NO: 10 or 17.
35. The fusion protein of claim 34, wherein the base-editing enzyme is a MGYPDa829 split base editor.
36. The fusion protein of claim 35, wherein the MGYPDa829 split base editor comprises an N-terminal portion of MGYPDa829 (M829_NT).
37. The fusion protein of claim 36, wherein the N-terminal portion of MGYPDa829 (M829_NT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 11.
38. The fusion protein of claim 37, wherein the N-terminal portion of MGYPDa829 (M829_NT) comprises the amino acid sequence of SEQ ID NO: 11.
39. The fusion protein of any one of claims 35-38, wherein the MGYPDa829 split base editor is activated by the addition of a C-terminal portion of MGYPDa829 (M829_CT).
40. The fusion protein of claim 39, wherein the C-terminal portion of MGYPDa829 (M829_CT) comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 12 or 19.
41. The fusion protein of claim 40, wherein the C-terminal portion of MGYPDa829 (M829_CT) comprises the amino acid sequence of SEQ ID NO: 12 or 19.
42. The fusion protein of claim 20, wherein the base-editing deaminase deaminates A to I (AM).
43. The fusion protein of claim 42, wherein the AM deaminase is (1) hADA; or (2) a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant.
44. The fusion protein of claim 43, wherein the adenine base editor is fused to the N- terminus of the DddA variant.
45. The fusion protein of claim 43 or 44, wherein the adenine base editor is TadA, optionally wherein the adenine base editor is TadA8e.
46. The fusion protein of any one of claims 43-45, wherein the DddA variant is DddA E1347A.
47. The fusion protein of any one of claims 42-46, wherein the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13 or 22, optionally wherein the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13 or 22.
48. The fusion protein of any one of claims 42-47, wherein the AM deaminase is a nonsplit full length AM deaminase.
49. The fusion protein of any one of claims 42-47, wherein the AM deaminase is a split AM deaminase.
50. The fusion protein of claim 49, wherein the AM deaminase split base editor comprises TadA8e fused to an N-terminal portion of DddA E1347A.
51. The fusion protein of claim 50, wherein the N-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14.
52. The fusion protein of claim 51, wherein the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14.
53. The fusion protein of any one of claims 49-52, wherein the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A.
54. The fusion protein of claim 53, wherein the C-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18.
55. The fusion protein of claim 54, wherein the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18.
56. The fusion protein of any one of claims 1-55, wherein the fusion protein is capable of binding to a DNA binding protein.
57. The fusion protein of any one of claims 1-56, wherein the fusion protein is capable of binding to a DNA binding protein- specific antibody.
58. The fusion protein of any one of claims 1-57, wherein the fusion protein is capable of catalyzing cytosine (C) or adenosine (A) deamination in the vicinity of the DNA binding protein-targeted genomic area.
59. The fusion protein of any one of claims 1-58, wherein the DNA-binding protein binds to a DNA directly.
60. The fusion protein of any one of claims 1-58, wherein the DNA-binding protein binds to a DNA indirectly.
61. The fusion protein of any one of claims 1-60, wherein the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
62. A nucleic acid encoding the fusion protein of any one of claims 1-61.
63. A vector comprising the nucleic acid of claim 62.
64. A kit comprising the fusion protein of any one of claims 1-61.
65. The kit of claim 64, wherein the kit further comprises a primary antibody that binds the DNA-binding protein.
66. The kit of claim 64 or 65, wherein the kit further comprises the C-terminal portion of DddA (DddA_CT) or the C-terminal portion of MGYPDa829 (MGYPDa829_CT).
67. The kit of any one of claims 64-66, wherein the kit further comprises a divalent cofactor.
68. The kit of claim 67, wherein the divalent cofactor is Zn2+.
69. A method of identifying a binding event of a DNA-binding protein in a cell or a population of cells, comprising:(a) contacting the cell or the population of cells with a fusion protein of any one of claims 1-48;(b) incubating the cell or the population of cells at a condition for a period of time; and(c) detecting base changes induced by the fusion protein by sequencing, thereby identifying the binding event.
70. The method of claim 69, wherein the DNA-binding protein binds to a DNA directly.
71. The method of claim 69, wherein the DNA-binding protein binds to a DNA indirectly.
72. The method of any one of claims 69-71, wherein the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
73. The method of any one of claims 69-72, wherein the method profiles all binding events of the DNA-binding protein in the cell or the population of cells.
74. The method of any one of claims 69-73, wherein the population of cells comprises at least 100 cells, at least 150 cells, at least 200 cells, or at least 250 cells.
75. The method of any one of claims 69-74, wherein the cell is a primary cell, optionally wherein the cell is a primary human cell, e.g., primary human PBMCs.
76. The method of any one of claims 69-74, wherein the cell is a cell line.
77. The method of any one of claims 69-76, wherein the cell or the population of cells are contacted with an antibody specific to the DNA-binding protein prior to step (a).
78. The method of any one of claims 69-77, wherein the cell or the population of cells are fixed and permeabilized prior to step (a).
79. The method of any one of claims 69-78, wherein step (b) occurs in the presence of the C-terminal portion of DddA (DddA_CT).
80. The method of any one of claims 69-79, wherein step (b) occurs in the presence of a divalent cofactor.
81. The method of claim 80, wherein step (b) occurs in the presence of Zn2+.
82. The method of any one of claims 69-81, wherein step (b) occurs at a temperature from 37°C to 50 °C.
83. The method of any one of claims 69-82, wherein the sequencing at step (c) is next generation sequencing.
84. The method of any one of claims 69-82, wherein the sequencing at step (c) is single cell sequencing.
85. The method of any one of claims 69-82, wherein the sequencing at step (c) is tagmentation-based single cell sequencing.
86. The method of any one of claims 69-82, wherein the sequencing at step (c) is Assay for Transposase- Accessible Chromatin using sequencing (ATAC-seq), single-cell (sc) ATAC-seq, Nanobody Tethered Tagmentation followed by sequencing (NTT-seq), CUT&Tag-related seq, or Dogma-seq.
87. The method of claim 86, wherein the sequencing at step (c) comprises Tn5 or TnY tagmentation.
88. The method of claim 86, wherein the sequencing at step (c) uses a lOx Genomics single-cell (sc) ATAC-seq kit.
89. The method of any one of claims 69-82, wherein the sequencing at step (c) is a droplet-based single cell sequencing method.
90. The method of any one of claims 69-82, wherein the sequencing at step (c) is a split and pool single-cell sequencing method.
91. The method of any one of claims 69-82, wherein the sequencing at step (c) is a spatially resolved chromatin profiling method.
92. The method of claim 91, wherein the sequencing at step (c) is Slide-seq or dBIT- seq.
93. The method of any one of claims 69-82, wherein the sequencing at step (c) is wholegenome sequencing.
94. The method of any one of claims 69-82, wherein the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing.
95. The method of claim 93, wherein the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PT A) protocol.
96. The method of any one of claims 69-95, wherein the method can be used to predict 3D chromatin structure in the cell or the population of cells.
97. A method of identifying a binding event of a DNA-binding protein in a cell or a population of cells, comprising:(a) fixing and permeabilizing the cell or the population of cells;(b) contacting the cell or the population of cells with a primary antibody that specifically binds the DNA-binding protein;(c) contacting the cell or the population of cells with a fusion protein comprising (1) a protein that binds the Fc region of the primary antibody, optionally wherein the protein is a secondary nanobody, a protein A, or a protein G; and (2) an N-terminal portion of a baseediting enzyme, optionally wherein the base-editing enzyme is selection from DddAl 1, MGYPDa829, LbsDaOl, or TadA8e-DddA E1347A;(d) adding a C-terminal portion of the base-editing enzyme and a divalent cofactor;(e) performing Tagmentation, optionally wherein the Tagmentation comprises Tn5 or TnY tagmentation;(f) extracting genomic DNA from the cell and constructing a library for sequence; and(g) sequencing the genomic DNA library to detect base changes induced by the fusion protein, thereby identifying the binding event.
98. The method of any one of claims 69-97, wherein the method further comprises conducting a genotyping method.
99. The method of claim 98, wherein the genotyping method is a high-throughput single-cell genotyping method.
100. The method of claim 98 or 99, wherein the genotyping method is Genotyping of Transcriptomes (GoT) and Genotyping of Targeted Loci and Chromatin Accessibility (GoT-ChA).
101. The method of any one of claims 69-100, wherein the method further comprises conducting cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), NTT-seq, ATAC with select antigen profiling by sequencing (ASAP-seq), simultaneous high-throughput ATAC and RNA expression with sequencing (SHARE-seq), Paired-Tag, and / or combined assay of transcriptome and enriched chromatin binding (CoTECH).
102. A system or kit comprising:(a) a first fusion protein comprising a first nanobody and a first deaminase; and(b) a second fusion protein comprising a second nanobody and a second deaminase, wherein (1) the first deaminase and the second deaminase have different motif preferences; and (2) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
103. The system or kit of claim 102, wherein the first deaminase prefers CpG context, and the second deaminase prefers TpC context.
104. The system or kit of claim 102 or 103, wherein the first deaminase is CbDaOl or AcDaOl; and the second deaminase is DddA or DddA-like.
105. A system or a kit comprising:(a) a first fusion protein comprising a first nanobody and a AM deaminase; and(b) a second fusion protein comprising a second nanobody and a C>U deaminase, wherein the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species.
106. The system or kit of claim 105, wherein the AM deaminase is (1) hADA, or (2) a fusion protein comprising an adenine base editor and an enzymatically dead DddA variant.
107. The system or kit of claim 106, wherein the adenine base editor is fused to the N- terminal of the DddA variant.
108. The system or kit of claim 106 or 107, wherein the adenine base editor is TadA8e.
109. The system or kit of any one of claims 106-108, wherein the DddA variant is DddA E1347A.
110. The system or kit of any one of claims 105-108, wherein the AM deaminase comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 13 or 22, optionally wherein the AM deaminase comprises the amino acid sequence of SEQ ID NO: 13 or 22.
111. The system or kit of any one of claims 105- 110, wherein the AM deaminase is a non-split full-length AM deaminase.
112. The system or kit of any one of claims 105-110, wherein the AM deaminase is a split AM deaminase.
113. The system or kit of claim 112, wherein the AM deaminase split base editor comprises TadA8e fused to an N-terminal portion of DddA E1347A.
114. The system or kit of claim 113, wherein the N-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 14.
115. The system or kit of claim 114, wherein the N-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 14.
116. The system or kit of any one of claims 112-115, wherein the AM deaminase split base editor is activated by the addition of a C-terminal portion of DddA E1347A.
117. The system or kit of claim 116, wherein the C-terminal portion of DddA E1347A comprises an amino acid sequence at least 95% homology to the amino acid sequence of SEQ ID NO: 15 or 18.
118. The system or kit of claim 117, wherein the C-terminal portion of DddA E1347A comprises the amino acid sequence of SEQ ID NO: 15 or 18.
119. The system or kit of any one of claims 102-118 further comprising primary antibodies to the two different DNA-binding proteins.
120. The system or kit of any one of claims 102-119 further comprising the C-terminal portion of DddA El 347 A.
121. The system or kit of any one of claims 102-120, wherein the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
122. The system or kit of claim 121, wherein the two DNA-binding proteins are cofactors, or different components from a DNA-binding protein complex.
123. A method of identifying binding events of two DNA-binding proteins in a cell or a population of cells, comprising:(a) contacting the cell or the population of cells with the system or kit of any one of claims 102-122;(b) incubate the cell or the population of cells at a condition for a period of time; and(c) detecting base changes induced by the system or kit by sequencing, thereby identifying the binding events of two DNA-binding proteins.
124. The method of claim 123, wherein the DNA-binding protein is a transcription factor or a chromatin remodeling factor.
125. The method of claim 123 or 124, wherein the method profiles all binding events of the two DNA-binding proteins in the cell or the population of cells.
126. The method of any one of claims 102-125, wherein the cell is a primary cell or a cell line.
127. The method of any one of claims 102-126, wherein the cell or the population of cells are contacted with two primary antibodies specific to the two DNA-binding proteins, respectively, prior to step (a).
128. The method of any one of claims 102-127, wherein the cell or the population of cells are fixed and permeabilized prior to step (a).
129. The method of any one of claims 102-128, wherein the sequencing at step (c) is single cell sequencing.
130. The method of any one of claims 102-128, wherein the sequencing at step (c) is tagmentation-based single cell sequencing.
131. The method of any one of claims 102-128, wherein the sequencing at step (c) is ATAC-seq, NTT-seq, CUT&Tag-related seq, or Dogma-seq.
132. The method of any one of claims 102-128, wherein the sequencing at step (c) comprises Tn5 or TnY tagmentation.
133. The method of claim 132, wherein the sequencing at step (c) uses a lOx Genomics scATAC-seq kit.
134. The method of any one of claims 102-128, wherein the sequencing at step (c) is a droplet-based single cell sequencing method.
135. The method of any one of claims 102-128, wherein the sequencing at step (c) is a split and pool single-cell method.
136. The method of any one of claims 102-128, wherein the sequencing at step (c) is a spatially resolved chromatin profiling method.
137. The method of claim 136, wherein the sequencing at step (c) is Slide-seq or dBIT- seq.
138. The method of any one of claims 102-128, wherein the sequencing at step (c) is whole genome sequencing.
139. The method of any one of claims 102-128, wherein the sequencing at step (c) comprises extracting genomic DNA from the cell and constructing a library for sequencing.
140. The method of claim 138, wherein the sequencing at step (c) comprises amplifying and obtaining a sequencing library from the entire genome using a Primary Template Amplification (PT A) protocol.
141. A method of identifying binding events of two DNA-binding proteins in a cell or a population of cells, comprising:(a) fixing and permeabilizing the cell or the population of cells;(b) contacting the cell or the population of cells with two primary antibodies specific to the two DNA-binding proteins, respectively;(c) contacting the cell or the population of cells with (1) a first fusion protein comprising a first nanobody and a first deaminase; and (2) a second fusion protein comprising a second nanobody and a second deaminase, wherein (i) the first deaminase and the second deaminase have different motif preferences or the first deaminase is a A>I deaminase and the second deaminase is a OU deaminase; and (ii) the first nanobody and the second nanobody bind to two different DNA-binding proteins, or the first nanobody and the second nanobody are secondary nanobodies that recognize immunoglobulins from two different species; and(d) sequencing genomic DNA of the cell or the population of cells to detect base changes induced by the first and the second fusion proteins, respectively, thereby identifying the binding events of the two DNA-binding proteins.
142. The method of claim 141, wherein the first deaminase and / or the second deaminase is a split enzyme, and the step (c) further comprises contacting the cell with a C-terminal portion of the first deaminase to activate the first deaminase and / or with a C-terminal portion of the second deaminase to activate the second deaminase.