Predicting 3D genome architecture from single-molecule sequencing data

WO2026169727A1PCT designated stage Publication Date: 2026-08-13CZ BIOHUB SF LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-08-13

Smart Images

  • Figure US2026013880_13082026_PF_FP_ABST
    Figure US2026013880_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a computer-implemented method for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions or a plurality of chromatin features from a single-molecule, long-read sequencing assay. The method comprising the steps of receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay and generating, via one or more processors, a visual representation using an artificial neural network model trained to predict haplotype- specific 3D chromatin structure information. Generating the visual representation includes steps of processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions, and processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions.
Need to check novelty before this filing date? Find Prior Art

Description

33167 / 70819PREDICTING 3D GENOME ARCHITECTURE FROM SINGLE-MOLECULE SEQUENCING DATACROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 754,197, filed on February 5, 2025, the entire contents of each of which are fully incorporated herein by reference.BACKGROUND OF THE INVENTION

[0002] Given the differences in scale between the linear genome length and cell nucleus size, all organisms have evolved a set of mechanisms to organize their genome in 3D space (Dekker and Mirny 2024). There is ever-increasing evidence linking 3D genome organization to fundamental cellular processes and various diseases (Sun et al. 2024; Giorgio et al. 2015; Krumm and Duan 2019; Sarni et al. 2020; Dubois et al. 2022; Dileep et al. 2023; Xu et al. 2022; Zheng and Xie 2019). To explore 3D genome structure, various short-read-based methods like Hi-C and Micro-C have been developed. Similarly, techniques like ChlP-seq, ATAC-seq, and bisulfite sequencing, have been developed to probe orthogonal genomic regulatory features. However, methods like Hi-C can be expensive, time-consuming, and technically challenging. A complete understanding of the genomic regulatory machinery in different contexts necessitates the use of multiple techniques, exacerbating experimental costs. These short-read approaches are additionally limited in their inability to resolve individual haplotypes, repetitive genomic regions, and the combinatorial interactions between regulatory machinery.

[0003] Recently, single-molecule sequencing methods such as Fiber-seq (Stergachis et al. 2020), SMAC-seq (Shipony et al. 2020), and DiMeLo-seq (Altemose et al. 2022) have emerged, addressing some of the limitations of traditional short-read assays for probing chromatin accessibility and protein-DNA interactions. For example, Fiber-seq has been used to examine the chromatin state in previously unmappable regions of the genome, such as centromeres and telomeres, generate haplotype-specific maps of genome regulation, and dissect disease mechanisms (Dubocanin et al. 2023; Dubocanin et al. 2022; Grasberger et al. 2024; Vollger et al. 2024; Vollger et al. 2023).

[0004] In parallel, recent advances in machine learning have enabled the development of frameworks to predict 3D genome organization directly from sequence data (Fudenberg et al. 2020; Schwessinger et al.2020; Zhou 2022). While these models can accurately reconstruct 3D genome architecture, they struggle to generalize across cell types, restricting their utility in research and clinical contexts. The C. Origami framework was developed to integrate chromatin accessibility and CTCF binding information alongside sequence data to reconstruct de novo Hi-C contact maps, closely matching experimental Hi-C results (Tan et al. 2023). However, this approach still requires multiple experiments and cannot provide haplotype-phased contact information.

[0005] Additionally, concerns have been raised that sequence-to-function models might learn biologically implausible patterns from training data, further limiting the generalizability of sequence-based models33167 / 70819(Kathail et al. 2024; Tang et al. 2023; Sasse et al. 2023; Huang et al. 2023). Recently, there has been progress in developing sequence-free models using scATAC-seq that may outperform traditional, sequence-based models (Gao et al. 2024).SUMMARY OF THE INVENTION

[0006] Disclosed herein is a computer-implemented method for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions from a single-molecule, long-read sequencing assay. In some embodiments, the method comprising the steps of receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generating, via one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin structure information.

[0007] In some embodiments, generating a visual representation comprises processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions, and processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions.

[0008] In some embodiments, the computer-implemented method may further comprise conducting the single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information.

[0009] Also disclosed herein is a computer-implemented method for concurrently predicting and visualizing a data set of a plurality of chromatin features from a single-molecule, long-read sequencing assay.

[0010] In some embodiments, the method comprising the steps of receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generating, via one or more processors, a visual representation using an artificial neural network model trained to predict a plurality of chromatin features.

[0011] In some embodiments, generating the visual representation comprises processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of a plurality of chromatin features, and processing, via one or more processors, the data set of a plurality of chromatin features to generate the visual representation of the data set of a plurality of chromatin features.

[0012] In some embodiments, the computer-implemented method further comprising conducting the single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information.

[0013] In some embodiments, the computer-implemented method includes where the genomic DNA-protein contact information comprises one or more or all of: chromatin accessibility information, CpGi33167 / 70819methylation state information, genomic DNA-binding protein binding site information, and genetic variation information.

[0014] In some embodiments, the genomic DNA-protein contact information comprises one or more or all of: Fiber-seq inferred Regulatory Element (FiRE) information, CpG methylation information, CTCF footprinting score, and CTCF direction information.

[0015] In some embodiments, the haplotype-specific 3D chromatin interactions information or the plurality of chromatin features comprises 3D chromatin conformation information, and one or more or all of: chromatin accessibility information, CpG methylation state information, genomic DNA-binding protein binding site information, and genetic valuation information.

[0016] In some embodiments, the 3D chromatin conformation information comprises Hi-C contact matrices. In some embodiments, the Hi-C contact matrices comprise a map at a resolution of approximately 10 kilobases across a genomic window of approximately 2 million base pairs.

[0017] In some embodiments, the DNA-protein binding site information comprises binding site locations for DNA binding-proteins.

[0018] In some embodiments, the DNA-binding proteins are selected from the group consisting of transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins.

[0019] In some embodiments, the DNA-binding protein is CTCF.

[0020] In some embodiments, the CpG methylation state information comprises methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC) information. In some embodiments, the genetic variation information comprises regulatory DNA variation.

[0021] In some embodiments, the genomic DNA-protein contact information further comprises epigenomic DNA-protein contact information.

[0022] In some embodiments, the haplotype-specific 3D genomic structure information further comprises cell-specific 3D genomic structure information. In some embodiments, the plurality of chromatin features comprise cell-type specific and haplotype-specific chromatin features.

[0023] In some embodiments, the methods described herein may further comprise processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to identify one or more patterns or markers indicative of a rare disease; and generating, via one or more processors, a report indicating the likelihood of the rare disease in a subject based on the processing.

[0024] Also described herein is a computer-implemented method of training an artificial neural network model to predict (a) haplotype-specific 3D chromatin interactions and / or (b) a plurality of chromatin features. In some embodiments, the computer-implemented method may comprise receiving, via one or more processors, training genomic DNA-protein contact information via a single-molecule, long-read33167 / 70819sequencing assay, and processing the training genomic DNA-protein contact information using an artificial neural network model to learn one or more training parameters, until the artificial neural network model learns to accurately predict a data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0025] Also described herein is a computing system for predicting and visualizing a data set of haplotypespecific 3D chromatin interactions and / or a plurality of chromatin features from a single-molecule, long-read sequencing assay.

[0026] In some embodiments, the computing system may comprise one or more processors, and one or more memories, having stored thereon computer-executable instructions that, when executed, cause the computing system to receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay; and generate, via the one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0027] In some embodiments, the computing system may generate the visual representation by processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features, and processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0028] Also described herein is a non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed, cause a computer to receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generate, via the one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.In some embodiments, the non-transitory computer-readable medium may generate the visual representation by processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features, and processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.INCORPORATION BY REFERENCE

[0029] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference for all purposes and to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.33167 / 70819BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0031] FIG. 1A. depicts representative images of a plurality of graphs (e.g., top to bottom: CTCF-ChIP, CTCF foot printing, ATAC, FiRE, and CpG) generated using published data sets using a GM12878 cell line for chromosome 12.

[0032] FIG. IB depicts a schematic of a computer-implemented method for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions from a single-molecule, long-read sequencing assay.

[0033] FIG. 1C depicts Hi-C maps (experimental and prediction) generated using the method shown in the schematic of FIG. IB.

[0034] FIG. ID depicts representative images (top, middle, bottom) of a ATAC graph, a FiRE graph, and a m6a graph.

[0035] FIG. IE depicts representative images (top, middle, bottom) of a CTCF-ChIP graph, a CTCF foot printing graph, and a m6a graph.

[0036] FIG. IF depicts a computing environment 104 for predicting and visualizing haplotype-specific 3D chromatin interactions and a plurality of chromatin features from a single-molecule, long-read sequencing assay, according to some aspects.

[0037] FIG. 1G depicts a computer-implemented method 180 for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions from a single-molecule, long-read sequencing assay, according to some aspects.

[0038] FIG. 2A (left, middle, right) depicts Hi-C graphs and corresponding CTCF graphs, FiRE graphs, and mCpG graphs from data generated in K562 cells using the neural network model trained on the data acquired from GM 12878 cells.

[0039] FIG. 2B depicts a graph of the test set accuracy for Insulation Spearman, Insulation Pearson, Spearman R, and Pearson R.

[0040] FIG. 2C depicts a graph of the predicted contacts versus the experimental contacts.

[0041] FIG. 2D depicts a graph of the standardized feature stratification by correlation using the CTCF footprinting score, the FiRE data, and CpG methylation.

[0042] FIG. 3A depicts Hi-C maps (experimental-top left and top middle, predictive-middle left and middle middle, and overlaid-top right and middle right) and corresponding CTCF graphs, FiRE graphs, and mCpG graphs (bottom left, middle, right) for chromosome 13 in a held-out K562 cell line or a GM 12878 cell line to identify differential 3D chromatin organization patterns between the two cell lines.

[0043] FIG. 3B depicts Hi-C maps (experimental-top left and top middle, predictive-middle left and middle middle, and overlaid-top right and middle right) and corresponding CTCF graphs, FiRE graphs, and mCpG graphs (bottom left, middle, right) for chromosome lin a held-out K562 cell line or a GM12878 cell line to identify differential 3D chromatin organization patterns between the two cell lines.33167 / 70819

[0044] FIG. 3C depicts a graph of the test set accuracy for Insulation Spearman, Insulation Pearson, Spearman R, and Pearson R for a held-out cell line (K562).

[0045] FIG. 4A depicts a graph of the calculated mean absolute error between two predicted maps of the active and inactive X chromosomes, and all autosomes.

[0046] FIG. 4B (top, bottom) depicts chrX correlation or autosome correlation for averaged haplotypespecific predicted maps vs experimental maps.

[0047] FIG. 4C depicts Hi-C graphs of the haplotypes along the active and inactive X chromosome including an overlay of the Hi-C maps.

[0048] FIG. 4D depicts a subset of autosomal contact maps (e.g., chromosome 12) where notable differences in 3D chromatin architecture were apparent, including an overlay of the Hi-C maps.

[0049] FIG. 5A depicts haplotype maps (top, bottom) of the FiberFold results across 32 Mb window of chromosome X.

[0050] FIG. 5B depicts haplotype maps (top, bottom) of the FiberFold results including the breakpoint in chromosome X across a 12Mb window.

[0051] FIG. 6 depicts haplotype maps (top, bottom) of the FiberFold results across a 12Mb window of chromosome 13 .

[0052] FIG. 7 depicts differences between the haplotype maps of FIG. 5B (top) and FIG. 6 (bottom) of the FiberFold results.

[0053] FIGS. 8A-8D depict haplotype maps of chromosome 20 used as controls in Example 5.DETAILED DESCRIPTION OF THE INVENTION

[0054] As described herein, FiberFold is a valuable tool for advancing the study of genome regulation by enabling the rapid analysis of 3D chromatin architecture in conjunction with, in various embodiments, chromatin accessibility, CTCF binding, CpG methylation, and underlying genetic architecture. This disclosure is the first demonstration of the measurement of these 5 features of individual genomes in one assay, and thus will increase the utility of single-molecule sequencing.

[0055] It is to be understood that this application is not limited to particular formulations or process parameters, as these may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Further, it is understood that a number of methods and materials similar or equivalent to those described herein can be used in the practice of the present disclosure.

[0056] In accordance with the present application, there may be employed conventional molecular biology, microbiology, and recombinant DNA techniques as explained fully in the art. The definitions contained herein supplement those in the art and are directed to the cunent application and are not to be imputed to any related or unrelated case, e.g., to any commonly owned patent or application. Accordingly, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.i33167 / 70819

[0057] In this application, the use of the singular includes the plural unless specifically stated otherwise. It must be noted that, as used in the specification, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Furthermore, use of the term “including” as well as other forms, such as “include”, “includes,” and “included,” is not limiting.

[0058] The terms “and / or” and “any combination thereof’ and their grammatical equivalents as used herein, can be used interchangeably. These terms can convey that any combination is specifically contemplated. Solely for illustrative purposes, the following phrases “A, B, and / or C” or “A, B, C, or any combination thereof’ can mean “A individually; B individually; C individually; A and B; B and C; A and C; and A, B, and C.”

[0059] The term “or” can be used conjunctively or disjunctively, unless the context specifically refers to a disjunctive use.

[0060] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in pail on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0061] As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of the present disclosure, and vice versa. Furthermore, compositions of the disclosure can be used to achieve methods of the disclosure.

[0062] As used herein the term “consisting essentially of’ refers to those elements required for a given embodiment. The term permits the presence of elements that do not materially affect the basic and novel or functional characteristic(s) of that embodiment of the disclosure.

[0063] As used herein the term “consisting of’ refers to compositions, methods, and respective components thereof as described herein, which are exclusive of any element not recited in that description of the embodiment.

[0064] Reference in the specification to “some embodiments,” “an embodiment,” “one embodiment” or “other embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiments is included in at least some embodiments, but not necessarily all embodiments, of the present disclosure.

[0065] To gain a comprehensive understanding of the genomic regulatory landscape, it is important to investigate various features underlying genomic regulation, such as CpG methylation, protein-DNAi33167 / 70819interactions, chromatin accessibility, and 3D genome structure. Respective short-read ensemble methods, such as Bisulfite sequencing, ChlP-seq, ATAC-seq, and Hi-C, and other assays have been developed to quantify these features individually and have yielded many discoveries, but they can be costly, time consuming, and have technical short-comings intrinsic to short-read sequencing approaches.

[0066] Notably, these methods frequently fail to capture haplotype-specific information. Additionally, one skilled in the art would not typically think to use data from a single-molecule, long-read sequencing assay to predict 3D structure. Instead, those skilled in the art would solely rely upon data from a short read assay to predict 3D structure.

[0067] Building on recent advancements in genomic deep learning and single-molecule sequencing, disclosed herein is FiberFold, a computational method for predicting Hi-C maps while simultaneously assaying genetic variation, chromatin accessibility and CpG methylation state. There can be significant 3D folding differences that can be seen between haplotypes.

[0068] Applying data from a single-molecule, long-read sequencing assay, Fiber-seq, FiberFold can generate haplotype-resolved Hi-C maps in a both enecked GM 12878 lineage and can predict 3D genomic contacts in a de novo, cell-line-specific manner. Applying FiberFold to Fiber-seq data from a bottlenecked GM12878 cell line, all haplotype-specific, topologically complex 3D genomic structures are demonstrated to predominantly occur on an active, paternal X chromosome and that the inactive X chromosome largely lacks complex 3D chromatin patterning. FiberFold showcases the power of integrating single-molecule, long-read sequencing with deep learning tools, enhancing its utility both at the bench and potentially the bedside.

[0069] Fiber-seq, as described in W02021203047, the contents of which are incorporated by reference in their entirety herein, with its ability to capture chromatin accessibility and CTCF binding in a single assay, presents an opportunity to streamline the generation of 3D genomic contact information. We further applied the information obtained from Fiber-seq to a trained neural network model, C. Origami, as described in WO2023168079, the contents of which are incorporated by reference in their entirety herein, and generated cell-type specific Hi-C maps. FiberFold demonstrates a method that integrates chromatin accessibility, CTCF binding, and CpG methylation data from Fiber-seq alone to predict cell-type and haplotype-specific 3D chromatin interactions from a single assay.

[0070] In line with recent findings on the haplotype-specificity of chromatin accessibility (Vollger et al.2024), it was observed that a majority of 3D contact maps across the autosomes are highly concordant between haplotypes.Computer-implemented method for predicting and visualizing a data set of haplotype-specific three dimensional (“3D”) chromatin interactions or a plurality of chromatin features

[0071] Disclosed is a computer-implemented method for predicting and visualizing a data set of haplotypespecific 3D chromatin interactions from a single molecule, long-read sequencing assay. In somei33167 / 70819embodiments, the single-molecule, long-read sequencing assay may comprise the assay, “Fiber-seq” as will be described in more detail herein.

[0072] In some embodiments, the method may comprise the steps of receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generating, via one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin structure information. In some embodiments, generating the visual representation may comprise processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions, and processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions.

[0073] In some embodiments, the computer-implemented method for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions from a single molecule, long-read sequencing assay to determine genomic DNA-protein contact information. Advantageously, the computer-implemented method for predicting and visualizing the data set of haplotype-specific chromatin interactions can predict and visualize the data set of haplotype-specific chromatin interactions without sequence information, as the inputs for the computer-implemented method are epigenomic in nature rather than genomic in nature.

[0074] In some embodiments, the visual depicture of the data may include a Hi-C map. The Hi-C map may measure or depict an aspect (e.g., locus-specific interaction patterns or genome-level patterns) of the 3D structure of the genome. It can be appreciated that the methods described may generate other types of outputs. For example, in some embodiments, the output may comprise any one or more of Micro-C / Region Capture Micro-C I 4C. Micro-C is a chromatin conformation capture method useful for visualizing 3D contacts of specific regions of the genome. Micro-C can generate this visualization reproducible and at a single base pair resolution. Region Capture Micro-C combines MNase-based 3C with a tiling regioncapture approach to generate a deep 3D genome map. Capture Micro-C can reveal microcompartments that connect enhancers and promoters.

[0075] The outputs of the methods described herein may be used to predict any type of 3D chromatin interaction track.

[0076] In some embodiments, a 3D chromatin interaction may comprise a pairwise map of contact frequencies across bins in the genome. It can be appreciated that the pairwise map may not be necessarily sufficient to fully resolve 3D genomic structure in all senses (e.g. resolving the contact frequencies within a genomic region would not necessarily tell you where that region is located in the nucleus with respect to e.g. the nuclear lamina). In some embodiments, the 3D chromatin interaction may also include 3D genomic structure.

[0077] Put simply, the methods described here may take one or more outputs (e.g., FiRE, CpG methylation, CTCF footprinting score, or CTCF direction) of a single-molecule, long-read sequencing assay in a genetic window and apply consecutive 1 -dimensional residual -convolutions to compress the information in a latent33167 / 70819space that is fed into a transformer of an artificial neural network, wherein the artificial neural network learns relational long-range interactions. The output of the transformer may be fed into consecutive 2-dimensional dilated residual convolutions to output a Hi-C map at a 10 kb resolution.

[0078] Also disclosed herein is a computer-implemented method for concurrently predicting and visualizing a data set of a plurality of chromatin features from a single-molecule, long-read sequencing assay. In some embodiments, the method comprises receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generating, via one or more processors, a visual representation using an artificial neural network model trained to predict a plurality of chromatin features. In some embodiments, generating the visual representation using the artificial neural network model comprises processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of a plurality of chromatin feature, and processing, via one or more processors, the data set of a plurality of chromatin features to generate the visual representation of the data set of a plurality of chromatin features.

[0079] In some embodiments, concurrently predicting and visualizing may include wherein the predicting and visualizing of the data happens within approximately the same period of time. In some embodiments, concurrently predicting and visualizing may include wherein predicting and visualizing may occur in parallel.

[0080] In some embodiments, the computer-implemented method may further comprise a step of conducting the single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information. In some embodiments, the single-molecule, long-read sequencing assay may comprise Fiber-seq.

[0081] In some embodiments, the genomic DNA-protein contact information comprises one or more of: chromatin accessibility information, CpG methylation state information, genomic DNA-binding protein binding site information, and genetic variation information.

[0082] In some embodiments, chromatin accessibility comprises genomic regions devoid of nucleosomes, often containing important functional regulatory elements. For example, in some embodiments, these genomic regions may include promoters and / or enhancers.

[0083] It can be appreciated that the CpG dinucleotide motif exists in 2 primary states: methylated or unmethylated at the C position, in the genome. In some embodiments, the CpG methylation state comprises the quantity of the relative proportion of 5mCG / CG at all CpG sites.

[0084] In some embodiments, genomic DNA-binding protein binding site information may comprise single-molecule m6A marks that delineate patches of inaccessible chromatin that correspond to DNA-binding motifs of particular proteins. For example, genomic DNA-binding protein binding site information may include the CTCF footprinting information that will be described in more detail herein.33167 / 70819

[0085] In some embodiments, the identification of single-nucleotide polymorphisms (SNPs), and large structural variants can be identified using single-molecule long-read sequencing. This genetic variation information can then be mapped onto these identified genetic structures to provide more accurate 3D contact predictions in a genetically-aware fashion.

[0086] In some embodiments, the genomic DNA-protein contact information comprises one or more of: Fiber-seq inferred Regulatory Element (FiRE) information, CpG methylation information, CTCF foot printing score, and CTCF direction information.

[0087] In some embodiments, FIREs are MTase sensitive patches (MSPs) that are inferred to be regulatory elements on single chromatin fibers. In some embodiments, FiRE information may comprise information from the regulatory elements on single chromatin fibers. In some embodiments, the MSPs may be identified using semi-supervised machine learning to identify MSPs that are likely to be regulatory elements using the Mokapot framework and XGBoost. Every individual FIRE element is associated with a precision value, which indicates the probability that the FIRE element is a true regulatory element. Tire precision of FIREs elements are estimated using Mokapot and validation data not used in training. The model targeting FIRE elements is trained with at least 90% precision, MSPs with less than 90% precision are considered to have average level of accessibility expected between two nucleosomes, and are referred to as linker regions. Semi-supervised machine learning with Mokapot requires a mixed-positive training set and a clean negative training set. To create mixed positive training data MSPs are selected that overlapped DNase hypersensitive sites (DHSs) and CTCF ChlP-seq peaks, and to create a clean negative training set MSPs are selected that did not overlap DHSs or CTCF ChlP-seq peaks.

[0088] In some embodiments, the CTCF foot printing score may comprise the following: data from the single-molecule long-read sequencing assay may be aligned and used to identify the subset of nucleotides in the motif that can be methylated. All the other sites may be checked to see if all sites that could be methylated are unmethylated and if all eligible sites flanking the footprint are methylated. This is described in more detail in Stergachis et al., 2020

[0089] In some embodiments, the CTCF direction information comprises the directionality of the CTCF motif within the genome. The direction of the CTCF motif may be either in the sense orientation or the antisense orientation.

[0090] Also described herein is a computer-implemented method for concurrently determining a plurality of chromatin features from a single-molecule, long-read sequencing assay. In some embodiments, the method may comprise the following steps: conducting a single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information; applying the information from conducting a singlemolecule, long-read sequencing assay to determine genomic DNA-protein contact information to a trained neural network model and determining a plurality of chromatin features.

[0091] In some embodiments, in the method for concurrently determining a plurality of chromatin features from a single-molecule, long-read sequencing assay, the step of conducting the single-molecule, long-read33167 / 70819sequencing assay may be previously completed. The method may comprise the steps of applying the genomic DNA-protein contact information to a trained neural network model and determining a plurality of chromatin features.

[0092] The disclosed methods represent an advancement in the analysis of 3D genomic structure, offering significant advantages over existing approaches. Unlike traditional methods that rely on separate assays for distinct-omic features, these methods enable the simultaneous prediction of 3D genomic contacts alongside genetic variation, CpG methylation states, and chromatin accessibility from a single Fiber-seq dataset. These methods eliminate additional hands-on time, providing seamless integration with existing Fiber-seq workflows, thereby enhancing the commercial utility for companies and researchers already utilizing this assay.

[0093] In some embodiments, these methods demonstrate the first application of long-read sequencing and / or single-molecule foot printing and / or enrichment / amplification-free sequencing to generate a predictive model of 3D genomic contacts, expanding the utility of these technologies. U.S. Patent No.11,475,275 generally describes a computer-implemented method for inferring a 3D structure of a genome. The method is configured to use as its input genome interaction data including from a Hi-C experiment, different from the methods described herein.

[0094] These methods also demonstrate the first approach to measure and / or predict 5 multi-modal features from a single assay, wherein the 5 multi-modal features may include genetic variation, CpG methylation, chromatin accessibility, transcription factor binding occupancy, and 3D genomic contacts. This reduces costs, hands-on time, and expertise needed to measure epigenetic information.

[0095] Disclosed herein is the first trained artificial neural network model capable of haplotype-specific 3D genomic contact prediction, addressing a crucial gap in the field.

[0096] In some embodiments, the trained artificial neural network model may be the first predictive 3D genomic contact model that does not rely heavily on genome sequence information, making it more universal and applicable to different cell types.

[0097] Although both Fiber-seq and 3D contact predictive models have been published since 2020 but no other groups have thought to combine these two methods because they involve disparate expertise. Methods for directly measuring 3D genomic contacts require very high-coverage sequencing datasets and enrichment / amplification steps, and it would not be obvious to a skilled artisan that a much lower coverage, enrichment-free dataset like Fiber-seq would be sufficiently information rich to predict 3D contacts.

[0098] The particular and essential ways in which features are extracted from single-molecule sequencing data would not be obvious even to a skilled artisan possessing expertise in both areas. No other model has attempted haplotype-specific 3D genomic contact prediction, as it is a unique advantage of long-read sequencing and required dedicated model-building in order to implement.33167 / 70819

[0099] All previous models based on bulk sequencing information rely on genome sequence information, and a skilled artisan would predict that model performance would deteriorate when removing a feature as rich as the human genome sequence.

[0100] In some embodiments, the methods for predicting haplotype-specific 3D genomic structure information, described in more detail herein, allow for multi-omic feature extraction. For example, the multi-omic feature extraction includes extraction of CpG methylation state, wherein single-molecule CpG data may be collapsed at a per-reference base level. The multi-omic feature extraction may include chromatin accessibility, wherein the chromatin accessibility is inferred using specific patterns within the single-molecule, long read sequencing assay data set. The multi-omic feature extraction may include transcription factor binding state and orientation. In some embodiments, the transcription factor binding state and orientation may include CTCF binding state and orientation, wherein the CTCF motifs in the genome are identified and their orientation is encoded. The CTCF binding is evaluated for every motif on every molecule. For example, a CTCF motif located within an accessible site but itself not accessible is labeled as a footprinting element and the footprinting elements may be counted at each motif to generate a 1 -dimensional CTCF binding track analogous to a ChlP-seq profile. In some embodiments, the CTCF binding state and orientation may include a CTCF footprinting score.

[0101] In some embodiments, the methods for predicting haplotype-specific 3D genomic structure information may include regulatory element inference. For example, the trained neural network may include an XGBoost model that is employed to convert m6A calls into inferred regulatory elements. In some embodiments, these infened regulatory elements may comprise promoters and / or enhancers. The inferred regulatory elements are piled up into a 1 -dimensional track for downstream analysis.

[0102] In some embodiments, the methods for predicting haplotype-specific 3D genomic structure may comprise a deep learning model for Hi-C contact map prediction. In some embodiments, the deep learning model may comprise an artificial neural network model. In some embodiments, the deep learning model may include input data that comprises single-molecule features such as methylation (e.g., CpG methylation), chromatin accessibility, transcription factor binding (e.g., CTCF binding). These singlemolecule features may be combined into 1 -dimensional tracks. In some embodiments, the neural network model may comprise an encoding layer, a transformer layer, and a decoding layer, as will be described in more detail herein. In some embodiments, the encoding layer may comprise a 1 -dimensional neural network with 12 layers of variable size processes the input feature. It can be appreciated that other numbers of layers of variable size are contemplated. In some variations, the encoding layer having less than 12 layers may allow for the elimination of computational complexity by leveraging other inputs that can be derived from long-read sequencing data, including but not limited to: Co-variance matrices of CpG methylation state, co-variance matrices of FiRE peaks, or co-variance matrices of CTCF binding.

[0103] In some embodiments, the transformer layer may comprise 8 transformer heads configured to capture long-range dependencies. It can be appreciated that other numbers of transformer heads are contemplated. In some embodiments, the decoding layer may comprise a pseudo-decoding 2-dimensional neural network with 6 dilated residual layers that may have gradually increasing dilations. In someIB33167 / 70819embodiments, the 2-dimensional neural network may output a predicted 3D contact matrix. In some embodiments, the neural network may predict a Hi-C contact map at a 10 kb resolution across a genomic window of 2 million base pairs, although other resolutions and size of genomic windows are contemplated, as will be described in more detail herein.Genomic DNA-Protein Contact Information

[0104] In some embodiments, the computer-implemented methods of predicting haplotype-specific 3D genomic structure information from a single-molecule, long-read sequencing assay or determining a plurality of chromatin features from a single-molecule, long-read sequencing assay may comprise a step of determining genomic DNA-Protein contact information. In some embodiments, the genomic DNA-Protein contact information may comprise any one or more of: chromatin accessibility information, CpG methylation state information, genomic DNA-binding protein binding site information, and genetic variation information.

[0105] For example, in some embodiments, the genomic DNA-Protein contact information may comprise chromatin accessibility information. In some embodiments, the genomic DNA-Protein contact information may comprise CpG methylation state information.

[0106] In some embodiments, the genomic DNA-Protein contact information may comprise genomic DNA-binding protein binding site information. In some embodiments, the genomic DNA-binding protein binding site information may comprise genetic variation information.

[0107] In some embodiments, the genomic DNA-Protein contact information may comprise any one or more of: Fiber-seq inferred Regulatory Element (FiRE) information, CpG methylation information, CCCTC-binding factor (CTCF) foot printing score, and CTCF direction information.

[0108] In some embodiments, the genomic DNA-Protein contact information may comprise FiRE information. FIREs are MTase sensitive patches (MSPs) that are inferred to be regulatory elements on single chromatin fibers. In some embodiments, the FiRE information may be identified using machine learning that identify MSPs that are likely to be regulatory elements. The FiRE information may enable accurate quantification of chromatin accessibility across the genome with single molecule and / or single nucleotide precision.

[0109] In some embodiments, the genomic DNA-Protein contact information may comprise CpG methylation information. In some embodiments, the genomic DNA-Protein contact information may comprise CTCF foot printing score. In some embodiments, the genomic DNA-Protein contact information may comprise CTCF direction information.

[0110] In some embodiments, the haplotype-specific 3D genomic structure information may comprise 3D chromatin conformation information and one or more of (i) chromatin accessibility information; (ii) CpG methylation state information; (iii) genomic DNA-binding protein binding site information; and (iv) genetic variation information.il33167 / 70819

[0111] In some embodiments, the chromatin accessibility information may include any one or more of: Assay for Transposase- Accessible Chromatin (ATAC) sequencing (ATAC-seq) data, DNase-seq data, or MNase-seq data. DNase-seq data, or MNase-seq data. In some preferred embodiments, chromatin accessibility information can include Assay for Transposase-Accessible Chromatin (ATAC) sequencing (ATAC-seq) data. In some embodiments, chromatin accessibility data can include one or more of acetylated H3K4 (H3K4ac), acetylated H3K9 (H3k9ac), acetylated H3K27 (H3K27ac), H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3, and data describing the same. In some embodiments, the chromatin accessibility data is ChlP-seq data.

[0112] In some embodiments, the CpG methylation state information may comprise methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC) information.

[0113] In some embodiments, the genetic variation information may comprise regulatory DNA variation. In some embodiments, regulatory DNA variation may include a portion of a DNA sequence involved in regulating the expression of a gene.

[0114] In some embodiments, the plurality of chromatin features may comprise 3D chromatin conformation information and one or more or all of (i) chromatin accessibility information, as described above; (ii) CpG methylation state information, as described above; (iii) genomic DNA-binding protein binding site information, as described above; and (iv) genetic variation information, as described above.

[0115] In some embodiments, the 3D chromatin conformation information comprises Hi-C contact matrices. In some embodiments, the Hi-C contact matrices may comprise one or more maps, wherein each map may be at a resolution of approximately 10 kilobases (kb) across a genomic window of approximately 2 million base pairs (bps). It can be appreciated that the map may comprise different resolutions across different genomic windows. For example, in some embodiments, the map may be at a resolution of approximately 1 base pair, 10 base pairs, 50 base pairs, 100 base pairs, 200 base pairs, 300 base pairs, 400 base pairs, 500 base pairs, 600 base pairs, 700 base pairs, 800 base pairs, 900 base pairs, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 11 kb, 12 kb, 13 kb, 14 kb, 15 kb, 16 kb, 17 kb, 18 kb, 19 kb, 20 kb, or more including up to a million base pairs. In some embodiments, the map may have a resolution across a genomic window of approximately 5,000 bps, 10,000 bps, 15,000 bps, 20,000 bps, 50,000 bps, 75,000 bps, 90,00bps, 100,000 bps, 200,000 bps, 300,000 bps, 400,000 bps, 500,000 bps, 600,000 bps, 700,000 bps, 800,000 bps, 900,000 bps, 1 million bps, 1.2 million bps, 1.4 million bps, 1.5 million bps, 1.6 million bps, 1.8 million bps, 2.0 million bps, 2.2 million bps, 2.4 million bps, 2.5 million bps, 2.6 million bps, 2.8 million bps, 3.0 million bps, 3.2 million bps, 3.4 million bps, 3.5 million bps, 3.6 million bps, 3.8 million bps, 4.0 million bps, 4.2 million bps, 4.4 million bps, 4.5 million bps, 4.6 million bps, 4.8 million bps, 5.0 million bps, 5.2 million bps, 5.4 million bps, 5.5 million bps, 5.6 million bps, 5.8 million bps, 6.0 million bps, 6.2 million bps, 6.4 million bps, 6.5 million bps, 6.6 million bps, 6.8 million bps, 7.0 million bps. 7.2 million bps, 7.4 million bps, 7.5 million bps, 7.6 million bps, 7.8 million bps, 8.0 million bps, 8.2 million bps, 8.4 million bps, 8.5 million bps, 8.6 million bps, 8.8 million bps, 9.0 million bps, 9.2 million bps, 9.4 million bps, 9.5 million bps, 9.6 million bps, 9.8 million bps, 10 million bps, or more including up33167 / 70819to 256 million base pairs. It can be appreciated that the map may be at a resolution of about 100 base pairs to about 20 kb base pairs across a genomic window of about 100,000 bps to about 10 million bps.

[0116] In some embodiments, the genomic DNA-binding protein binding site information may comprise binding site locations for DNA binding proteins. In some embodiments, the DNA-binding proteins may include transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins. In some embodiments, the DNA-binding protein may include one or more of: CTCF, CTCFL, RAD21, STAG1, STAG2, SMC1, SMC3, ZNF143, YY1, NIPBL, WAPL, TRIM22, and BATF. In some aspects, the DNA-binding protein may comprise any member of the RNA polymerase II complex including but not limited to TFIIA, TFIIB, TFIID, IFIIE, TFIIF, TFIIH, and TFIIJ. In some aspects, the DNA-binding protein may comprise any member of the RNA polymerase I complex including but not limited to POLR1A, POLR1B, POLR1C, POLR1D, POLR1E, POLR2F, TAF1, TBP, and UTF. In some aspects, the DNA-binding protein may comprise any member of die RNA polymerase III complex including but not limited to RPC3, RPC6, RPC7, RPC4, RPC5, RPC7, RPC9, BDP1, BRF1, BRF2, TBP, TFIIIA, TFIIIB, and TFIIIC.

[0117] In some embodiments, the genomic DNA-protein contact information may further comprise epigenomic DNA-protein contact information. In some embodiments, epigenomic DNA-protein contact information may include all genome-wide DNA-protein contact information.

[0118] In some embodiments, the haplotype-specific 3D genomic structure information may further comprise cell-specific 3D genomic structure information. In some embodiments, the plurality of chromatin features comprise cell-type specific and haplotype-specific chromatin features.

[0119] Also disclosed herein is a method of simultaneously predicting 3D genomic contacts, genetic variation information, chromatin accessibility information, and CpG methylation state information from a single long-read sequencing assay dataset. In some embodiments, the method comprises conducting a single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information, applying the information from conducting a single-molecule, long-read sequencing assay to a trained neural network model architecture; and predicting haplotype-specific 3D chromatin interactions. In some embodiments, conducting the single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information may have previously occurred. The method of simultaneously predicting 3D genomic contacts, genetic variation information, chromatin accessibility information, and CpG methylation state information from a single long-read sequencing assay dataset may comprise the steps of applying the information from conducting a single-molecule, long-read sequencing assay to a trained neural network model architecture; and predicting haplotype-specific 3D chromatin interactions.

[0120] Also disclosed herein is a method of determining haplotype specific 3D genomic contact information as there can be significant 3D folding difference between haplotypes. In some embodiments, the method comprises conducting a single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information, applying the information from conducting a single-molecule, long-read33167 / 70819sequencing assay to a trained neural network model architecture; and predicting haplotype-specific 3D chromatin interactions.

[0121] Also disclosed herein is a method of identifying or predicting a rare disease. In some embodiments, the method comprises conducting a single-molecule, long-read sequencing assay to determine genomic DNA-protein contact information, applying the information from conducting a single-molecule, long-read sequencing assay to a trained neural network model architecture; and predicting haplotype-specific 3D chromatin interactions.Single-Molecule, Long-Read Sequencing Assay (Fiber-seq)

[0122] As described above, the methods described herein may comprise a step of conducting a singlemolecule, long-read sequencing assay to determine genomic DNA-protein contact information. In some embodiments, the single-molecule, long-read sequencing assay may include Fiber-seq as described in W02021203047, the contents of which are incorporated by reference in their entirety herein.

[0123] In some embodiments, the single-molecule, long-read sequencing assay may advantageously include multi-omic feature extractions. For example, the single-molecule CpG data is collapsed at a perreference base level for the CpG methylation state. Also, in the single-molecule, long-read sequencing assay, the chromatin accessibility is inferred using specific patterns within the single-molecule, long-read sequencing assay dataset.

[0124] In some embodiments, the single-molecule, long-read sequencing assay may identify CTCF motifs in the genome including their orientation. Their orientation may be encoded to determine CTCF binding state and orientation. In some embodiments, the single-molecule, long-read sequencing assay may evaluate CTCF binding for every motif on every molecule. If a CTCF motif is located within an accessible site but it is not accessible, the CTCF motif may be labeled as a footprinted element. In some embodiments, the footprinted elements may be counted at each motif to generate a 1 -dimensional CTCF binding hack that is analogous to a ChlP-seq profile.

[0125] Fiber-seq may comprise a method for identifying regions of genomic DNA bound to a protein. The protein may include any protein that limits the access of an adenine methyltransferase (A-MTase) to an adenine base present in the genomic sequence bound by the protein. The protein may be one of more of: nucleosomes, transcription factors, transcriptional repressors, and the like. Various steps and aspects of the methods will now be described in greater detail below.

[0126] In some embodiments, the method include contacting genomic DNA with an adenine methyltransferase (A-MTase), where the A-MTase causes methylation of adenine residues in regions of the genomic DNA not bound to a protein, and conducting single-molecule long-read sequencing of the contacted genomic DNA to detect locations in the genomic DNA lacking methylated adenine residues to identify regions of genomic DNA bound to a protein. This method may be described in US20230134592, the contents of which are incorporated by reference in their entirety herein.33167 / 70819

[0127] In certain aspects, the A-MTase is a N6-adenine methyltransferase (m6 A-MTase). In certain aspects, the m6 A-MTase is selected from the group consisting of: Hia5 or a biologically active fragment thereof, EcoGII or a biologically active fragment thereof, Btrl92IV or a biologically active fragment thereof, EcoGI or a biologically active fragment thereof, M.CviPI or a biologically active fragment thereof, and M.SssI or a biologically active fragment thereof.

[0128] In some embodiments, contacting genomic DNA with an A-MTase may involve introducing, into the cell, A-MTase or a nucleic acid encoding the A-MTase.

[0129] In some embodiments, the genomic DNA contacted with A-MTase may be from a single cell, a plurality of cells (e.g., cultured cells), tissue, an organ, or another organism (e.g., bacteria, yeast, or the like). In some embodiments, the genomic DNA is from a cell(s), tissue, organ, and / or the like of an animal.

[0130] In some embodiments, the method includes contacting genomic DNA with a non-adenine methylfransferase where the non-adenine methyltransferase causes methylation of non-adenine residues in regions of the genomic DNA not bound to a protein, and conducting single-molecule long-read sequencing of the contacted genomic DNA to detect locations in the genomic DNA lacking methylated non-adenine residues to identify regions of genomic DNA bound to a protein.

[0131] In some embodiments, the method includes contacting genomic DNA with an enzyme that modifies DNA in a covalent manner where the enzyme that modifies DNA in a covalent manner causes a covalent change in regions of the genomic DNA not bound to a protein, and conducting single-molecule long-read sequencing of the contacted genomic DNA to detect locations in the genomic DNA that lack the covalent change to identify regions of genomic DNA bound to a protein. In some embodiments, the enzyme that modifies DNA in a covalent manner includes any enzyme that modifies DNA in a covalent manner. In some embodiments, the enzyme that modifies DNA in a covalent manner may include enzymes that impart a covalent addition to the nucleotides of DNA including covalent additions to adenine, cytosine, guanine, or thymine, including for example, addition of methyl groups or other amino groups to the nucleotides. It can be appreciated that other covalent modifications to the DNA may be included within the scope including for example, cytosine deamination or ketoxal mediated guanine deamination (non-enzymatic).

[0132] In some embodiments, the animal is a mammal. For example, a mammal may include a mammal from the genus Homo, a rodent (e.g., a mouse or rat), a dog, a cat, a horse, a cow, or any other mammal of interest. In some embodiments, the genomic DNA may be from one or more cells, tissue, organ, and / or the like of a human. In other aspects, the genomic DNA is from a source other than a mammal, such as bacteria, yeast, insects, amphibians, viruses, plants, or any other non-mammalian source. In certain aspects, the genomic DNA is from a cancer cell.

[0133] In some embodiments, the genomic DNA may be cell-free. Such cell-free genomic DNA may be present in, or obtained from, any suitable source. In some embodiments, the cell-free genomic DNA is present in or obtained from a body fluid sample selected from the group consisting of: whole blood, blood33167 / 70819plasma, blood serum, amniotic fluid, saliva, urine, pleural effusion, bronchial lavage, bronchial aspirates, breast milk, colostrum, tears, seminal fluid, peritoneal fluid, pleural effusion, and stool.

[0134] In some embodiments, the genomic DNA is cell-free fetal DNA. In certain aspects, the genomic DNA is circulating tumor DNA. In some embodiments, the genomic DNA comprises infectious agent DNA. In some embodiments, the genomic DNA comprises DNA from a transplant. The term "cell-free genomic DNA" as used herein can refer to genomic DNA composition having no cells or substantially no cells. Genomic DNA does not necessarily imply that all of the genetic material of a cell is present, rather, genomic DNA can include a fraction of the genomic material of a cell. For example, genomic DNA may encompass isolated chromatin fragments, which may be any segment of genomic DNA isolated from a cell that is in association with a nuclear protein. Exemplary chromatin fragments may be oligonucleosomes, mononucleosomes, centromeres, telomeres or genomic DNA bound by a transcription factor or chromatin remodeling factor.

[0135] In some embodiments, the cells may be peripheral blood mononuclear cells (PBMCs), leukocytes, or may be isolated from bone marrow, thymus, tissue biopsy, tumor, lymphoma, lymph node, gut associated lymphoid tissue, mucosa associated lymphoid tissue, spleen, other lymphoid tissues, liver, lung, stomach, intestine, colon, kidney, pancreas, breast, bone, prostate, cervix, testes, ovaries, tonsil, or other organ, and / or cells derived therefrom. In some embodiments, the nucleic acid (e.g. genomic DNA, chromosomal DNA) to be assessed is from blood cells including blood cells from a sample of whole blood or a sub- population of cells in whole blood. Subpopulations of cells in whole blood may include platelets, red blood cells (erythrocytes), platelets and white blood cells (i.e., peripheral blood leukocytes, which are made up of neutrophils, lymphocytes, eosinophils, basophils and monocytes). White blood cells can be further divided into two groups, granulocytes (which are also known as polymorphonuclear leukocytes and include neutrophils, eosinophils and basophils) and mononuclear leukocytes (which include monocytes and lymphocytes). Lymphocytes can be further divided into T cells, B cells and NK cells. Peripheral blood cells are found in the circulating pool of blood and not sequestered within the lymphatic system, spleen, liver, or bone marrow.

[0136] In some embodiments, the method may involve analyzing genomic DNA obtained pre-treatment and genomic DNA obtained post-treatment. For example, genomic DNA may be analyzed after 1 day, 1 week, 10 days, 15 days, 1 month, 3 months, 6 months or more post-treatment to compare the regions of the DNA not bound by protein(s) and hence susceptible to adenine methylation. Comparison of adenine methylation pattern may be used to assess change in transcriptional profile of the genome.

[0137] In some embodiments, the method may be used to generate a reference chromatin structure and regulatory regions for a type of cell including for multiple types of human cells. The chromatin structure and regulatory regions in a cell from a subject having a disorder may be compared to the reference chromatin structure and regulatory regions for that cell type to determine any differences. Such differences may reveal previously unknown changes in chromatin structure and regulatory regions that may be used for diagnosis, prognosis, or treating the subject.33167 / 70819

[0138] In some embodiments, the population of cells used for the methods may be composed of any number of cells, including for example about 500 to about 106or more cells.

[0139] In some embodiments, the genomic DNA is present in its native environment during exposure to the methyl transferase. For example, the genomic DNA may be present in a cell (e.g., an intact cell or permeabilized cell) during exposure to the methyl transferase. In some embodiments, a cell-permeable methyl transferase that crosses an intact or permeabilized cell membrane may be employed. In some embodiments, a methyltransferase may be introduced into the cell using standard techniques. In certain aspects, the genomic DNA is present in a cell lysate during exposure to the methyltransferase.

[0140] In some embodiments, the genomic DNA is part of a nucleic acid sample isolated from a cell(s), tissue, organ, and / or the like of an organism, e.g., an animal, such as a human. Approaches, reagents and kits for isolating, purifying and / or concentrating nucleic acid molecules from sources of interest are known in the art and commercially available. For example, kits for isolating DNA from a source of interest include the DNeasy®, QIAamp®, QIAprep® and QIAquick® nucleic acid isolation / purification kits by Qiagen, Inc. (Germantown, Md); the DNAzol®, Charge Switch®, Purelink®, GeneCatcher® nucleic acid isolation / purification kits by Life Technologies, Inc. (Carlsbad, CA); the NucleoMag®, NucleoSpin®, and NucleoBond® nucleic acid isolation / purification kits by Clontech Laboratories, Inc. (Mountain View, CA).

[0141] In some embodiments, the nucleic acid may be isolated from a fixed biological sample, e.g., formalin-fixed, paraffin-embedded (FFPE) tissue. Genomic DNA from FFPE tissue may be isolated using commercially available kits - such as the AllPrep® DNA / RNA FFPE kit by Qiagen, Inc. (Germantown, Md), the RecoverAll® Total Nucleic Acid Isolation kit for FFPE by Life Technologies, Inc. (Carlsbad, CA), and the NucleoSpin® FFPE kits by Clontech Laboratories, Inc. (Mountain View, CA).

[0142] In some embodiments, subsequent to contacting the genomic DNA with a methyltransferase and prior to the sequencing, the genomic DNA may be processed for sequencing. For example, the methods may include treating the ends of the genomic DNA to produce blunt ends. Blunting is a process by which a single-stranded overhang is either “filled in”, by the addition of nucleotides on the complementary strand using the overhang as a template for polymerization, or by “chewing back” the overhang, using an exonuclease activity. DNA polymerases, such as the Klenow fragment of DNA Polymerase I and T4 DNA Polymerase may be used to fill in (5' 3’) and chew back (3’ 5’). Removal of a 5’ overhang can be accomplished with a nuclease, such as Mung Bean Nuclease. In some embodiments, genomic DNA may be sheared or enzymatically digested after treatment with the methyltransferase.

[0143] It can be appreciated that single molecule real time sequencing systems may be applied to the detection of methylated adenine through analysis of the sequence and / or kinetic data derived from such systems. In particular, methylated adenine may alter the enzymatic activity of a nucleic acid polymerase in various ways, e.g., by increasing the time for a bound nucleobase to be incorporated and / or increasing the time between incorporation events. In certain embodiments, polymerase activity is detected using a single molecule nucleic acid sequencing technology. In certain embodiments, polymerase activity is detected1033167 / 70819using a nucleic acid sequencing technology that detects incorporation of nucleotides into a nascent strand in real time.

[0144] In some embodiments, a single molecule nucleic acid sequencing technology is capable of realtime detection of nucleotide incorporation events. Such sequencing technologies are known in the art and include, e.g., the SMRT® sequencing and nanopore sequencing technologies. For more information on nanopore sequencing, see, e.g., U.S. Pat. No. 9,175,348; U.S. Pat. No. 5,795,782; Kasianowicz, et al. (1996) ProcNatl Acad Sci USA 93(24): 13770-3; Ashkenas, et al. (2005) Angew Chem Int Ed Engl 44(9): 1401-4; Howorka, et al. (2001) Nat Biotechnology 19(7): 636-9; and Astier, et al. (2006) J Am Chem Soc 128(5): 1705-10, all of which are incorporated herein by reference in their entireties for all purposes. With regards to nucleic acid sequencing, the term “template” refers to a nucleic acid molecule subjected to template-directed synthesis of a nascent strand. A template may comprise, e.g., DNA or analogs, mimetics, derivatives, or combinations thereof, as described elsewhere herein. Further, a template may be singlestranded, double- stranded, or may comprise both single- and double-stranded regions. A modification in a double-stranded template may be in the strand complementary to the newly synthesized nascent strand, or may be in the strand identical to the newly synthesized strand, i.e., the strand that is displaced by the polymerase.

[0145] In some embodiments, the direct methylation sequencing described may generally be carried out using single molecule real time sequencing systems, i.e., that illuminate and observe individual reaction complexes continuously over time, such as those developed for SMRT® DNA sequencing (see, e.g., P. M. Lundquist, et al., Optics Letters 2008, 33, 1026, which is incorporated herein by reference in its entirety for all purposes). The foregoing SMRT® sequencing instrument generally detects fluorescence signals from an array of thousands of zero mode waveguides (ZMWs) simultaneously, resulting in highly parallel operation. Each ZMW, separated from others by distances of a few micrometers, represents an isolated sequencing chamber.

[0146] Detection of single molecules or molecular complexes in real time, e.g., during the course of an analytical reaction, generally involves direct or indirect disposal of the analytical reaction such that each molecule or molecular complex to be detected is individually resolvable. In this way, each analytical reaction can be monitored individually, even where multiple such reactions are immobilized on a single substrate. Individually resolvable configurations of analytical reactions can be accomplished through a number of mechanisms, and typically involve immobilization of at least one component of a reaction at a reaction site. Various methods of providing such individually resolvable configurations are known in the art, e.g., see European Patent No. 1105529 to Balasubramanian, et al.; and Published International Patent Application No. WO 2007 / 041394, the full disclosures of which are incorporated herein by reference in their entireties for all purposes.

[0147] A reaction site on a substrate is generally a location on the substrate at which a single analytical reaction is performed and monitored, preferably in real time. A reaction site may be on a planar surface of the substrate or may be in an aperture in the surface of the substrate, e.g., a well, nanohole, or other aperture.1133167 / 70819

[0148] In some embodiments, such apertures are “nanoholes,” which are nanometer-scale holes or wells that provide structural confinement of analytic materials of interest within a nanometer-scale diameter, e.g., 1-300 nm. In some embodiments, such apertures comprise optical confinement characteristics, such as zeromode waveguides, which are also nanometer-scale apertures and are further described elsewhere herein. Typically, the observation volume (i.e., the volume within which detection of the reaction takes place) of such an aperture is at the attoliter (KG 18 L) to zeptoliter (KG21 L) scale, a volume suitable for detection and analysis of single molecules and single molecular complexes.

[0149] In some embodiments, the immobilization of a component of an analytical reaction can be engineered in various ways. For example, an enzyme (e.g., polymerase, reverse transcriptase, kinase, etc.) may be attached to the substrate at a reaction site, e.g., within an optical confinement or other nanometerscale aperture. In other embodiments, a substrate in an analytical reaction (for example, a nucleic acid template, e.g., DNA, derivatives, and mimetics thereof, or a target molecule for a kinase) may be attached to the substrate at a reaction site. Certain embodiments of template immobilization are provided, e.g., in U.S. patent application Ser. No. 12 / 562,690, filed Sep. 18, 2009, now U.S. Pat. No. 8,481,264, and incorporated herein by reference in its entirety for all purposes. One skilled in the art will appreciate that there are many ways of immobilizing nucleic acids and proteins into an optical confinement, whether covalently or non- covalently, via a linker moiety, or tethering them to an immobilized moiety. These methods are well known in the field of solid phase synthesis and micro-arrays (Beier et ah, Nucleic Acids Res. 27:1970-1-977 (1999)). Non-limiting exemplary binding moieties for attaching either nucleic acids or polymerases to a solid support include streptavidin or avidin / biotin linkages, carbamate linkages, ester linkages, amide, thiolester, (N)-functionalized thiourea, functionalized maleimide, amino, disulfide, amide, hydrazone linkages, among others. Antibodies that specifically bind to one or more reaction components can also be employed as the binding moieties. In addition, a silyl moiety can be attached to a nucleic acid directly to a substrate such as glass using methods known in the art.

[0150] In some embodiments, other processing steps useful for detecting the locations of the methylated adenine in the genomic DNA using a nanopore may be employed. For example, the methods may include adding one or more nanopore sequencing adapters or subregions thereof to one or more ends of the genomic DNA. By “nanopore sequencing adapter” is meant one or more nucleic acid domains that include at least a portion of a nucleic acid sequence (or complement thereof) utilized by a nanopore sequencing platform of interest, such as a nanopore sequencing platform provided by Oxford Nanopore Technologies, e.g., a Mini ON™, GridIONx5™, PromethlON™, or SmidglON™ nanopore-based sequencing system. Nanopore sequencing adapters of interest may be added via chemical or enzymatic ligation, or any other available approaches for joining one or more nucleic acid molecules to one or more ends of the double- stranded nucleic acid molecule. Suitable reagents (e.g., ligases) and kits for performing ligation reactions are known and available, e.g., the Instant Sticky-end Ligase Master Mix available from New England Biolabs (Ipswich, MA). Ligases that may be employed include, e.g., T4 DNA ligase (e.g., at low or high concentration), T4 DNA ligase, T7 DNA Ligase, E. coli DNA Ligase, Electro Ligase®, or the like. Conditions suitable for performing the ligation reaction will vary depending upon the type of ligase used.33167 / 70819

[0151] In some embodiments, single-molecule, circular consensus sequencing (CCS) may be used to generate accurate long read sequences. In some embodiments, CCS may involve rendering the DNA topologically circular and sequencing the DNA multiple times in order to create a consensus sequence. In certain aspects, the circular DNA may be sequenced up to 20 times, e.g., 5-20 times, 5-15 times, 10-20, or 10-15 times.

[0152] In some embodiments, prior to the sequencing, the genomic DNA may be processed to generate long fragments, including for example from about Ikb long to about 100 kb long. In some embodiments, the locations of methylated adenines (or other methylated or covalently modified nucleotides) are detected in a contiguous stretch of a strand of the double-stranded nucleic acid molecule of about 500 bases or greater.

[0153] Computational approaches (e.g., in the form of software) may be employed to detect the locations of the methylated adenines (or other methylated or covalently modified nucleotides) in single and / or double-stranded nucleic acid molecule, determine protein bound regions in the nucleic acid molecule based on the detected locations of methylated adenines, sequence the nucleic acid molecule, e.g., using single molecule real time sequencing, and optionally, CCS, and any combinations thereof.

[0154] Also disclosed are methods for determining nucleosome positions in genomic DNA. Such methods exploit the protected / inaccessible nature of nucleosome- associated genomic DNA from the methylase (e.g., aN6-adenine DNA methyltransferase, a non-adenine methylfransferase) employed, such that methylation does not occur in nucleosome-associated genomic DNA. The methods include detecting location of methylated adenine in genomic DNA that mark the locations of tinker genomic DNA in the genomic DNA. The nucleosome positions in the genomic DNA are determined based on the absence of methylated adenines. Such methods may also reveal the presence or absence of certain transcription factors bound to genomic DNA.

[0155] In some embodiments, the method disclosed may be conducted on one or a plurality of normal cells to generate a chromatin accessibility map for the region of genomic DNA sequenced, where the map indicates regions of chromatin not bound to protein(s) and hence accessible to the A-MTase (or non-adenine methylfransferase or enzyme that covalently modifies nucleotides) and regions of the chromatin bound to protein(s) and hence inaccessible to the A-MTase (or non-adenine methylfransferase or enzyme that covalently modifies nucleotides).

[0156] In some embodiments, the method disclosed may be conducted on one or a plurality of test cells to generate a chromatin accessibility map for genomic DNA of the test cell(s). The test cells may be from a subject, such as a mammal, e.g., a human patient. The subject, in some cases, may have or may be suspected of having a disease. The disease may be cancer.

[0157] In some embodiments, tlie method disclosed may further include comparing the chromatin accessibility map for the test cell to that of the normal cell, wherein the test cell and the normal cell are of the same cell type and comparing the genomic DNA sequences of the test and normal cells, wherein33167 / 70819presence of a difference in chromatin accessibility maps indicates a change in chromatin architecture in the test cell, wherein presence of a difference in genomic DNA sequence in absence of a difference in chromatin accessibility maps indicates that the sequence difference is not associated with a change in chromatin structure, and wherein presence of a difference in genomic DNA sequence and of a difference in chromatin accessibility maps indicates the sequence difference is associated with a change in chromatin structure.

[0158] In some embodiments, the method further comprises generating a database comprising information regarding chromatin accessibility map, the underlying genomic DNA sequence, and correlation, if any, to a condition or disease. In certain aspects, the normal cell and test cell may be epithelial cells, white blood cells, glial cells, osteoblasts, or chondrocytes. In certain aspects, the normal cell and the test cell may comprise a plurality of cells. In certain aspects, the plurality of cells comprises at least 10 cells, at least 30 cells, at least 100 cells, at least 300 cells, or at least 10,000 cells.

[0159] In some embodiments, the chromatin accessibility map encompasses at least 10% of a chromatin, e.g., at least 30%, at least 50%, or at least 80% of a chromatin. In some embodiments, the chromatin accessibility map encompasses at least 10% of the genome of the cell, including at least about 20%, at least 30%, at least 50%, or at least 80% of the genome of the cell. In certain aspects, the protein(s) bound to the genomic DNA includes nucleosomes, transcriptional regulator such as, transcriptional repressors and transcriptional activators, or both.

[0160] In some embodiments, detecting presence of methylated adenine (mA, e.g., m6A) or other modified nucleotides (including covalently modified nucleotides or methylated nucleotides that are not adenine) in the cell may involve visualization of an antibody bound to the mA. The visualization may be epifluorescence imaging when using a fluorescence label bound directly or indirectly to the antibody. A super-resolution microscopy method may be utilized for visualization of an antibody bound to the mA. The super-resolution microscopy method may be a deterministic super-resolution microscopy method, which utilizes a fluorophore's nonlinear response to excitation to enhance resolution. Exemplary deterministic super- resolution methods may include stimulated emission depletion (STED), ground state depletion (GSD), reversible saturable optical linear fluorescence transitions (RESOLFT), and / or saturated structured illumination microscopy (SSIM). A super resolution microscopy method may also include a stochastic super-resolution microscopy method, which utilizes a complex temporal behavior of a fluorophore, to enhance resolution. Exemplary stochastic super-resolution method may include super-resolution optical fluctuation imaging (SOFI), all single-molecular localization method (SMLM) such as spectral precision determination microscopy (SPDM), SPDMphymod, photo-activated localization microscopy (PALM), fluorescence photo-activated localization microscopy (FPALM), stochastic optical reconstruction microscopy (STORM), and dSTORM.

[0161] In some embodiments, detecting may include generating a map of spatial location of the methylated adenines in the genome of the cell. The detecting may include generating a map of spatial and temporal location of the methylated adenines in the genome of the cell.1133167 / 70819

[0162] In some embodiments, the method may include contacting a plurality of cells of the same type with the A-MTase (or non-adenine methyltransferase or enzyme that covalently modifies nucleotides) and generating a map of spatial location of the methylated adenines (or methylated nucleotides) in the genome of the cells.

[0163] In some embodiments, the method may include contacting a plurality of cells of the same type at at least two different time points with the A-MTase (or non-adenine methyltransferase or enzyme that covalently modifies nucleotides ) and generating a map of spatial and temporal location of the methylated adenines (or methylated nucleotides) in the genome of the cells. The two different time points may include a first time point and a second time point, wherein the first and second time points are separated by a time point at which a therapy is administered to the cells. The cells may be obtained from a subject and wherein the subject is administered the therapy. In some embodiments, the cells visualized by the disclosed methods may be live cells or fixed and permeabilized cells.Generating a Visual Representation using an Artificial Neural Network Model (C. Origami)

[0164] As described above, methods disclosed herein may comprise generating, via one or more processors, a visual representation using an artificial neural network model trained to predict haplotypespecific 3D chromatin structure information or a plurality of chromatin features. In some embodiments, the artificial neural network model may comprise C. Origami as described in WO2023168079, the contents of which are incorporated by reference in their entirety herein.

[0165] C. Origami is a deep neural network that synergistically integrates DNA sequence features and two essential cell type- specific genomic features, DNA-binding protein profile (e.g., CTCF binding profile (CTCF ChlP- seq signal)) and chromatin accessibility information (e.g., ATAC-seq signal) to typically generate one or more Hi-C maps. C. Origami achieved accurate prediction of cell type-specific chromatin architecture in both normal and rearranged genomes. Additionally, the high-performance of C. Origami enables in silico genetic perturbation experiments that interrogate the impact on chromatin interactions and moreover, allows the identification of cell type-specific regulators of genomic folding through in silico genetic screening. Taken together, it is believed that the underlying deep learning architecture, Origami, to be generalizable for predicting genomic features and discovering novel genomic regulations.

[0166] C. Origami, a neural network that accurately predicts cell type-specific genome folding, and enables in silico genetic studies of its regulation. C. Origami achieves cell type specificity by synergistically encoding both DNA sequence and minimum cell type-specific features. C. Origami is demonstrated to be able to de novo predict the genome folding of new cell types with high accuracy. Additionally, C. Origami enables in silico genetic perturbation studies for discovering new cell type-specific regulators of genomic folding. Collectively, it is believed that tire Origami architecture for integrating both DNA sequence information and cell type-specific features to be generalizable for future genomics studies and is capable of discovering novel regulatory mechanisms.33167 / 70819

[0167] It can be appreciated that other types of neural network architecture may also be used that utilize other features that can be derived from single-molecule sequencing data including but not limited to: CpG co-variances, FiRE-co-variances, CTCF binding co-variances, linker length distributions, high-order nucleosome phasing or the like, wherein the inclusion of these features may allow for higher-capacity neural network models with simpler architectures to be more computational ly efficient.

[0168] In some embodiments, the trained neural network model works by providing a method of predicting 3D genomic features in a target cell, the method comprising: training a neural network model architecture integrating (1) nucleotide-level DNA sequences, and (2) cell type specific genomic features, wherein the cell type-specific genomic features comprise (i) genomic DNA-binding protein binding profile information, and (ii) chromatin accessibility information, thereby generating a trained neural network model architecture; applying the trained neural network model architecture to a genomic window of a target cell; and identifying genomic features within die genomic window of die target cell. In some embodiments, die nucleotide-level DNA sequences comprise a naturally occurring wild type sequence, a mutated DNA sequence, or a synthetic DNA sequence.

[0169] In some embodiments, the cell type-specific genomic features comprise DNA binding profile information obtained for (1) transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins, or from (2) chromatin feature distribution profiles. In some embodiments, the chromatin feature distribution profiles comprise histone modifications, DNA modifications, chromatin accessibility information. In some embodiments, the genomic DNA-binding protein is selected from the group consisting of CTCF, CTCFL, RAD21, STAG1, STAG2, SMC1, SMC3, ZNF143, YY1, NIPBL, WAPL, TRIM22, and BATF. In some embodiments, the genomic DNA-binding protein is CTCF. In some embodiments, the genomic DNA-binding protein binding profile information comprises ChlP-seq data, CUT&RUN data, CUT&TAG data, or DamID data in the genomic window of the target cell. In some embodiments, the cell type-specific genomic features comprise chromatin feature distribution profiles. In some embodiments, the chromatin feature distribution profiles comprise histone modification data, DNA modification data. In some embodiments, the chromatin accessibility information comprises one or more of H3K4ac, H3K9ac, H3K27ac, H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3. In some embodiments, the chromatin accessibility information is selected from the group consisting of ATAC-seq data, DNase-seq data, or MNase-seq data. In some embodiments, the cell type-specific genomic information comprises a DNA modification profile. In some embodiments, the DNA modification profile comprises DNA methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC), or DNA formylated cytosine (5hmC), or carboxylated cytosine (5caC). In some embodiments, the chromatin accessibility information comprises ATAC- seq data in the genomic window of the target cell. In some embodiments, genomic features comprise identification of a topologically associating domain (TAD). In some embodiments, the genomic window comprises a contiguous genomic region of 2 million bases. In some embodiments, the model architecture comprises two encoders, a transformer module, and a decoder. In some embodiments, the decoder is a decoder associated with Hi-C contact matrices for predicting complex chromatin architecture.IB33167 / 70819

[0170] In some embodiments, the present disclosure provides a computer-implemented machine for predicting 3D genomic features in a target cell using the trained neural network model, comprising: a processor; a neural network comprising a first encoder, a second encoder, a transformer module, and a decoder; and a tangible computer-readable medium operatively connected to the processor and including computer code configured to: train a neural network model architecture integrating (1) nucleotide-level DNA sequences, and (2) cell type-specific genomic features, wherein the cell type-specific genomic features comprise (i) genomic DNA-binding protein binding profile information, and (ii) chromatin accessibility information, thereby generating a trained neural network model architecture; apply the trained neural network model architecture to a genomic window of a target cell; and identify genomic features within the genomic window of the target cell. In some embodiments, the nucleotide-level DNA sequences comprise a naturally occurring wild type sequence, a mutated DNA sequence, or a synthetic DNA sequence.

[0171] In some embodiments, the trained artificial neural network model may comprise C. Origami. In some embodiments, the cell type-specific genomic features comprise DNA binding profile information obtained for (1) transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins, or from (2) chromatin feature distribution profiles. In some embodiments, the chromatin feature distribution profiles comprise histone modifications, DNA modifications, chromatin accessibility information. In some embodiments, the genomic DNA-binding protein is selected from the group consisting of CTCF, CTCFL, RAD21, STAG1, STAG2, SMC1, SMC3, ZNF143, YY1, NIPBL, WAPL, TRIM22,B ATF, any members of RNAPII-complex, any members of RNAPI-complex, and any members of RNAPIII-complex. In some embodiments, the genomic DNA-binding protein is CTCF. In some embodiments, the genomic DNA-binding protein binding profile information comprises ChlP-seq data, CUT&RUN data, CUT&TAG data, or DamID data in the genomic window of the target cell. In some embodiments, the cell type-specific genomic features comprise chromatin feature distribution profiles. In some embodiments, the chromatin feature distribution profiles comprise histone modification data, DNA modification data.

[0172] In some embodiments, the chromatin accessibility information comprises one or more of H3K4ac, H3K9ac, H3K27ac, H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3. In some embodiments, the chromatin accessibility information is selected from the group consisting of ATAC-seq data, DNase-seq data, or MNase-seq data. In some embodiments, the cell type-specific genomic information comprises a DNA modification profile. In some embodiments, the DNA modification profile comprises DNA methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC), or DNA formylated cytosine (5hmC), or carboxylated cytosine (5caC). In some embodiments, the genomic DNA- binding protein binding profile information comprises ChlP-seq data for the genomic DNA-binding protein in the genomic window of the target cell. In some embodiments, the chromatin accessibility information comprises ATAC-seq data in the genomic window of the target cell. In some embodiments, genomic features comprise identification of a topologically associating domain (TAD). In some embodiments, the genomic window comprises a contiguous genomic region of 2 million bases. In some embodiments, the model architecture comprises two encoders, a transformer module, and a decoder. In some embodiments, the decoder is a decoder associated with Hi-C contact matrices for predicting complex chromatin architecture.1133167 / 70819

[0173] In some embodiments, DNA modification data may include data that relates to any modification of DNA that specifically or preferentially targets DNA not bound by proteins (e.g., accessible DNA). For example, this data may be read out using single-molecule long-read sequencing. In some embodiments, the DNA modification data may include data related to methyltransferases (adenine MTases or cytosine Mtases), deaminases (ddA / ddB, or the like) or reactive small molecules that covalently attach to or alter the DNA. For example, DNA modification data may include acetylated H3K4 (H3K4ac), acetylated H3K9 (H3k9ac), acetylated H3K27 (H3K27ac), H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3, and data describing the same.

[0174] Also provided are methods for predicting chromatin structure including 3D genomic features in a target cell. In some embodiments, the provided methods enable the prediction of 3D chromatin architecture within a target cell.

[0175] In some embodiments, genomic features comprise genome organization, including 3D genome organization. In some embodiments, genomic features comprise genome folding. In some embodiments, the method comprises training a neural network model architecture integrating genomic structure data, epigenomic data, and / or genomic sequence data. Genomic structure data can include, for example, chromatin folding data, topological associating domain (TADs) and TAD boundary data, and other known metrics for assessing genome structure, including 3D chromatin structure. Genomic sequence data generally includes genomic DNA sequence data, such as continuous sequences of DNA within a chromosome. The genomic sequence data can be obtained by applying known genomic DNA sequencing methods to a target cell, or from previously-generated genomic DA sequence data, such as from a genomic DNA sequence database. Epigenomic data can include, for example, transcriptional regulatory data, such as genomic DNA-binding protein data (e.g., CTCF-binding data). In some embodiments, the epigenomic data is obtained for a genomic window of a target cell

[0176] In some embodiments, a trained artificial neural network model integrates nucleotide-level DNA sequences, and / or cell type-specific genomic features. DNA sequences can include a wild type sequence, a mutated DNA sequence, or a synthetic DNA sequence. Cell type-specific features can include one or more of: genomic DNA-binding protein binding profile information, and chromatin accessibility information. Cell type-specific features can include DNA binding profile information obtained for transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins, or from chromatin feature distribution profiles. Chromatin feature distribution profiles can include data describing histone modifications, DNA modifications, chromatin accessibility information.

[0177] Genomic DNA-binding protein binding profile information can include ChlP- sequencing (ChlP-seq) data, CUT&RUN data, CUT&TAG data, Hi-C / 4C, DamID, MadID, pA-DamlD DiMeLo-Seq data obtained for a genomic DNA-binding protein in a target cell. Non -limiting examples of genomic DNA-binding proteins include CCCTC-binding factor (CTCF), CTCFL, RAD21, STAG1, STAG2, SMC1, SMC3, ZNF143, YY1, NIPBL, WAPL, TRIM22, BATF, or the like.IB33167 / 70819

[0178] In some embodiments, DiMeLo-seq may comprise a method for determining the genomic location of at least one biomolecule-genomic DNA interaction and DiMeLo-Seq data may comprise data generated or retrieved using this method. DiMeLo-seq is described in WO2022256469, the contents of which are incorporated by reference in their entirety herein.

[0179] In one embodiment, the DiMeLo-seq method comprises one or more or all of the following steps, (a) incubating a biomolecule of interest under conditions that allow the biomolecule of interest to contact a genomic DNA sequence; (b) isolating and permeabilizing nuclei from the cells in (a) under conditions that allow isolation of genomic DNA bound by the biomolecule of interest; (c) contacting the biomolecule bound to genomic DNA with a first binding moiety capable of specifically binding to the biomolecule of interest; (d) contacting the first binding moiety with a second binding moiety capable of specifically binding to the first binding moiety, wherein said second binding moiety is conjugated to an enzyme capable of modifying genomic DNA; (e) incubating the first binding moiety and second binding moiety of (d) under conditions that allow modification of genomic DNA; (f) isolating and preparing the genomic DNA for sequencing, wherein said preparing does not require amplification of the DNA; and (g) sequencing the genomic DNA under conditions that allow determining the location of the biomolecule-DNA interaction.

[0180] In some embodiments, chromatin accessibility information can include Assay for Transposase-Accessible Chromatin (AT AC) sequencing (ATAC-seq) data, DNase-seq data, or MNase-seq data. DNase-seq data, or MNase-seq data. In some embodiments, chromatin accessibility information can include Assay for Transposase-Accessible Chromatin (ATAC) sequencing (ATAC-seq) data In some embodiments, chromatin accessibility data can include one or more of acetylated H3K4 (H3K4ac), acetylated H3K9 (H3k9ac), acetylated H3K27 (H3K27ac), H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3, and data describing the same. In some embodiments, the chromatin accessibility data is ChlP-seq data.

[0181] In some embodiments, cell type-specific genomic information comprises a DNA modification profile. In some embodiments, the DNA modification profile comprises DNA methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC), or DNA formylated cytosine (5hmC), or carboxylated cytosine (5caC).

[0182] In some embodiments, a trained artificial neural network is applied to a genomic window in a target cell. As used herein, the term “genomic window” refers to a contiguous segment of genomic DNA. A genomic window can contain at least 100 bases to at least about at least 10 million bases. In some embodiments, the genomic window may be as described above, containing anywhere from at least about 100 bases to 100,000 bases, from at least about 100,000 bases to about 10 million bases, or from at least about 5 kilobases to about 256 megabases.

[0183] Provided is a method of predicting 3D genomic features in a target cell, the method comprising training an artificial neural network model integrating (1) nucleotide-level DNA sequences, and (2) cell type-specific genomic features, wherein the cell type-specific genomic features comprise (i) genomic DNA-binding protein binding profile information, and (ii) chromatin accessibility information, thereby generating33167 / 70819a trained artificial neural network model; applying the trained neural network model architecture to a genomic window of a target cell; and identifying genomic features within the genomic window of the target cell. In some embodiments, predicting genomic features comprise identifying or characterizing a topologically associated domain (TAD).Computer Implemented Machines for Predicting 3D Chromatin Structure or 3D genomic features

[0184] Also provided herein are computer implemented machines for predicting chromatin structure in a target cell. In some embodiments, the provided machines enable the prediction of 3D cliromati n architecture within a target cell.

[0185] Also provided herein are computer-implemented machines for predicting 3D genomic features in a target cell. In some embodiments, a machine comprises at least one processor, a trained artificial neural network, and a tangible computer-readable medium operatively connected to the processor. In some embodiments, the trained artificial neural network comprises a first encoder, a second encoder, a transformer module, and a decoder. In some embodiments, the tangible computer-readable medium includes computer code. In some embodiments, the model architecture comprises two encoders, a transformer module, and a decoder. In some embodiments, the decoder is a decoder associated with Hi-C contact matrices for predicting complex chromatin architecture.

[0186] In some embodiments, a computer implemented machine described herein is configured to: train an artificial neural network model architecture integrating (1) nucleotide-level DNA sequences, and (2) cell type-specific genomic features, wherein the cell type-specific genomic features comprise (i) genomic DNA-binding protein binding profile information, and (ii) chromatin accessibility information, thereby generating a trained neural network model architecture; apply the trained neural network model architecture to a genomic window of a target cell; and identify genomic features within the genomic window of the target cell.

[0187] In some embodiments, a computer implemented machine is configured to train a neural network model architecture integrating nucleotide-level DNA sequences, and / or cell type-specific genomic features. In some embodiments, a neural network model architecture integrates nucleotide-level DNA sequences and cell type-specific genomic features. Cell type-specific features can include one or more of: genomic DNA-binding protein binding information, and chromatin accessibility information.

[0188] Genomic DNA-binding protein binding information can include ChlP-sequencing (ChlP-seq) data obtained for a genomic DNA-binding protein in a target cell. Nonlimiting examples of genomic DNA-binding proteins include CCCTC-binding factor (CTCF), CTCFL, RAD21, STAG1, STAG2, SMC1, SMC3, ZNF143, YY1, NIPBL, WAPL, TRIM22, BATF, or the like.

[0189] In some embodiments, chromatin accessibility information can include Assay for Transposase-Accessible Chromatin (ATAC) sequencing (ATAC-seq) data, DNase-seq data, or MNase-seq data. DNase-seq data, or MNase-seq data. In some embodiments, chromatin accessibility information can include Assay for Transposase-Accessible Chromatin (ATAC) sequencing (ATAC-seq) data In some embodiments,33167 / 70819chromatin accessibility data can include one or more of acetylated H3K4 (H3K4ac), acetylated H3K9 (H3k9ac), acetylated H3K27 (H3K27ac), H3K4mel, H3K4me2, H3K4me3, H3K9me3, H3K27me3, H3K36me3, and data describing the same. In some embodiments, the chromatin accessibility data is ChlP-seq data.

[0190] In some embodiments, cell type-specific genomic information comprises a DNA modification profile. In some embodiments, the DNA modification profile comprises DNA methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC), or DNA formylated cytosine (5hmC), or carboxylated cytosine (5caC).

[0191] In some embodiments of a computer implemented machine as described herein, a trained neural network is applied to a genomic window in a target cell. As used herein, the term “genomic window” refers to a contiguous segment of genomic DNA. A genomic window can contain at least 100 bases to at least about 5 million bases.Systems for Use in predicting and visualizing a data set of haplotype-specific 3D chromatin interactions

[0192] Also disclosed herein are systems which find use, e.g., in practicing the subject methods, including carrying out one or more of any of the steps of the methods described above.

[0193] Also described herein is a computing system for predicting and visualizing a data set of haplotypespecific 3D chromatin interactions and / or a plurality of chromatin features from a single-molecule, long-read sequencing assay.

[0194] In some embodiments, the computing system may comprise one or more processors, and one or more memories, having stored thereon computer-executable instructions that, when executed, cause the computing system to receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay; and generate, via the one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0195] In some embodiments, the computing system may generate the visual representation by processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features, and processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0196] Also described herein is a non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed, cause a computer to receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generate, via the one or more processors, a visual representation using an artificial neural network1133167 / 70819model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0197] In some embodiments, the non-transitory computer-readable medium may generate the visual representation by processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features, and processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

[0198] In some embodiments, encompassed by the systems that find use in determining bound regions in genomic DNA are systems that find use in determining nucleosome positions in genomic DNA. The instructions for such systems cause the system to sequence genomic DNA that has been treated with an adenine methylase, and record the locations of the methylated adenines in the genomic DNA. The instructions for such systems may further cause the system to assess transcriptional accessibility of certain regions of the genome based on the determined positions of methylated adenines in the genomic DNA. The instructions for such systems may further cause the system to assess differential nucleosome occupancy or phasing near the promoters of genes. In some embodiments, the instructions for such systems cause the system to sequence genomic DNA that has been treated with a non-adenine methylase or an enzyme that covalently modifies DNA, and record the locations of the methylated non-adenines or nucleotides that have been covalently modified in the genomic DNA.

[0199] In some embodiments, the systems may be adapted (e.g., include instructions) to sequence a contiguous stretch of the genomic DNA of 500 bases or greater to about 100 kb or greater, and record the locations of such methylated adenines, methylated non-adenines, or other covalently modified nucleotides.

[0200] In some embodiments, the system includes a sequencing device such as a commercially available sequencer, e.g., PacBio sequencer.

[0201] The present disclosure includes computer-readable medium, including non- transitory computer-readable medium, which stores instructions for methods, or portions thereof, described herein, and which may be part of the systems of the present disclosure. Aspects of the present disclosure include computer-readable medium storing instructions that, when executed, cause the system to perform one or more steps of a method as described herein.

[0202] In some embodiments, instructions in accordance with the methods and systems described herein can be coded onto a computer-readable medium in the form of “programming”, where the term "computer-readable medium" as used herein refers to any storage or transmission medium that participates in providing instructions and / or data to a computer for execution and / or processing. Examples of storage media include a floppy disk, hard disk, optical disk, magneto optical disk, CD-ROM, CD-R, magnetic tape, non-volatile memory card, ROM, DVD-ROM, Blue-ray disk, solid state disk, and network attached storage (NAS), whether or not such devices are internal or external to the computer. A file containing information can be33167 / 70819“stored” on computer-readable medium, where “storing” means recording information such that it is accessible and retrievable at a later date by a computer.

[0203] Any steps of the methods or those carried out by the systems of the present disclosure can be executed using programming that can be written in one or more of any number of computer programming languages. Such languages include, for example, Java (Sun Microsystems, Inc., Santa Clara, CA), Visual Basic (Microsoft Corp., Redmond, WA), and C++ (AT&T Corp., Bedminster, NJ), as well as any many others.Applications

[0204] The disclosed methods may be useful for a number of different applications. In some embodiments, the disclosed methods may be useful for accurately predicting the effects of genetic and epigenetic variation on 3D chromatin interactions in subject samples acquired from a subject having a disease. In some embodiments, these predictions of the effects of genetic and epigenetic variation on 3D chromatin interaction may be useful for resolving the potential etiology of a disease including for example a genetic disease. In some embodiments, the subject may include a mammal including a human. For example, if a subject has a chromosomal translocation, a single-molecule, long-read sequence assay may identify the chromosomal translocation, and then with the same data, the trained artificial neural network model may be able to predict whether elements on the translocated genomic element interact with the new elements it is now proximal to on another chromosome. Previously, predicting the effects on generic and epigenetic variation on 3D chromatin interactions in a subject sample using other methods would be cost prohibitive. Advantageously, predicting the effects on generic and epigenetic variation on 3D chromatin interactions in a subject sample using the methods describe herein is less cost prohibitive.

[0205] For example, in some embodiments, the disclosed methods may be useful for predicting the effects on 3D chromatin architecture of a translocation breakpoint from a rare disease in a patient. In this embodiment, sequencing data may be used to perform a de novo assembly of the patient’s genome. That genome assembly, which includes the chromosome rearrangements, may be used as the reference for aligning Fiber-seq reads and generating input features for Fiber-fold. Advantageously, this is one useful way of detecting the functional impact of genetic variation using long-read sequencing data. Because the rearrangements occur on only one haplotype each, this also underscores the importance of having haplotype-specific predictions, which would not be possible with short-read methods.

[0206] In some embodiments, the disease may comprise any one or more of: cancer, cystic fibrosis, sickle cell anemia, Huntington’s disease, Marfan syndrome, Ehlers-Danlos syndrome, Hemophilia, Hirschsprungs Disease, Tay-Sachs disease, Polycystic kidney disease, Duchenne muscular dystrophy, Breast cancer, Parkinson’s disease, heart disease, Down syndrome, Angelman syndrome, Arrhythmogenic Right Ventricular Dysplasia / Cardiomyopathy, Ankylosing spondylitis, Apert syndrome, Brugada Syndrome, Charcot-Marie-Tooth disease, Cleidocranial Dysplasia, congenital adrenal hyperplasia, Fabry disease, Fragile X syndrome, hemochromatosis, neurofibromatosis, Noonan syndrome, Prader-Willi syndrome, Rett syndrome, Turner syndrome, Von Willebrand disease, Von Hippel-Lindau Disease (VHL), Williams33167 / 70819syndrome. In some embodiments, the disease may comprise any rare disease that may have a genetic origin, including those diagnosed or undiagnosed.EXAMPLES

[0207] Leveraging advances in deep learning and single-molecule sequencing techniques, FiberFold enables researchers to capture chromatin accessibility, CTCF binding, CpG methylation state, genetic variants, and 3D genome structure all from a single, long-read assay that can accurately map across the genome in a haplotype-specific manner. FiberFold distinctly does not utilize sequence in the underlying model in an aim to increase the generalizability of de-novo Hi-C predictions and we observe accuracies similar to state-of-the-art sequence based models, indicating that chromatin accessibility, CTCF binding, and CpG methylation alone are sufficient to predict 3D chromatin structure.

[0208] FiberFold demonstrated the ability to accurately capture unique 3D chromatin organization in a novel cell type. It is shown that along autosomes, the 3D chromatin organization, at 2MB scale, is nearly identical between haplotypes. In contrast, in an activated X chromosome-haplotype there is a much more topologically intricate pattern of chromatin organization compared to the inactive chromosome.

[0209] Here, it is demonstrated that a single assay is sufficient to generate genetic information, CpG methylation state, accessibility, TF binding, and 3D genomic contacts, all in a long-read format. For example, with disparate experimental approaches, generating the data used in this study would have required 5 experiments / cell line, increasing the number of sequencing experiments in our study from 3 to 15. At scale, the time and cost savings provided by integration of deep learning with single-molecule sequencing may be massive. The rapid throughput of multi-omic studies enabled by the combination of Fiber-seq with FiberFold will prove particularly valuable to groups aiming to sequence hundreds or thousands of samples.Example 1 -FiberFold Architecture, Training and Evaluation

[0210] The FiberFold model builds upon the foundational framework of C. Origami combined with the inputs from Fiber-seq, with several important modifications to enhance performance and generalizability. Notably, FiberFold was designed to exclude the underlying DNA sequence from its training inputs, a deliberate choice aimed at improving the model's ability to generalize across different cell types and genomic contexts. Instead of relying on sequence information, FiberFold integrates three distinct feature tracks derived entirely from Fiber-seq data: Fiber-seq Inferred Regulatory Elements (FIREs, a proxy for chromatin accessibility), CpG methylation states, and CTCF footprinting scores. The directionality of CTCF motifs was also included, given that this information cannot be accurately learned from our inputs, yet is functionally relevant (Fig IB) (Rowley and Corces 2018; Guo et al. 2015).33167 / 70819

[0211] FiberFold takes in 4 features (FiRE, CpG methylation, CTCF footprinting score, CTCF direction) in a 2,097,152 bp window and applies a consecutive ID residual-convolutions to compress the information in a latent space that is then fed into a transformer, which learns relational long-range interactions. The output of the transformer is then fed into consecutive 2-D dilated residual convolutions to output a Hi-C map at lOkb resolution.

[0212] To train our FiberFold model, previously published datasets were utilized from an extensively profiled, tier-1 ENCODE cell line, GM12878 (Vollger et al. 2023). FiREs have been extensively profiled and have been shown to recapitulate signals captured by DNasel-footprinting and / or ATAC-seq (FIG.1A and FIG. ID) (Vollger et al. 2024; Vollger et al. 2023; Stergachis et al. 2020). Additionally, this work has established the robustness of Fiber-seq for CTCF footprinting, which is corroborated by our data (FIG. 1 A-1E). A monotonic increase was observed in our Fiber-seq derived CTCF footprinting signal relative to CTCF ChlP-seq data (ENCFF485TGR, pearson R on quantile normalized data = 0.98, p = 1.31E-7), albeit theirs is limited resolution at very weakly bound footprints due to the depth-constrained dynamic range of single-molecule data compared to ensemble methods.

[0213] FiberFold is trained to predict Hi-C contact maps by minimizing the mean squared error between experimentally derived and model-predicted contact maps. The genome was split into 2MB windows, sliding across the genome with a 36kb window with a randomized shift parameter applied to these windows to increase training data diversity. Regions that overlapped with the hg38 Encode Blacklist or where we have significant gaps in coverage were not used. For model evaluation, chromosomes were split into training, validation, and test sets, with chromosome 10 used for validation and chromosome 15 reserved for testing.

[0214] FiberFold’ s performance on the holdout chromosome exhibited strong concordance with experimental Hi-C maps, yielding a holdout mean Pearson correlation coefficient of 0.938 and a test-set mean Spearman correlation of 0.906 (FIG. 2A-2C). It was reasoned that the signal across and closely proximal to the diagonal along the Hi-C map is a function of read depth. In order to calculate a correlation of meaningful contacts driven be active processes, a pixel-wise comparison was calculated across all testset maps at contacts => lOOkb, revealing a mean Pearson correlation of 0.88 (FIG. 2C). Moreover, insulation score analysis at 500kb windows revealed a high correlation between predicted and experimental insulation tracks (test-set pearson R = 0.836), further supporting FiberFold’s accuracy in delineating topologically associated domains (TADs) at this scale.Example 2 - Rad21-associated CTCF binding drive insulation

[0215] To better understand the role of individual features in driving chromatin compartmentalization, the relative contribution of each input track was assessed as a function of the insulation score across predicted Hi-C maps. Our analysis revealed a positive correlation between CTCF footprinting scores and insulation strength (FIG. 2D), consistent with prior evidence that CTCF-mediated loop-extrusion is one of the major primary contributor to chromatin compartmentalization at this scale (33 and 46). In contrast, there does not33167 / 70819appeal' to be a monotonic relationship between chromatin accessibility or CpG methylation, and insulation strength (FIG. 2D).

[0216] To understand the nature of the CTCF sites driving 3D chromatin structure predictions, the pixelwise relationship between CTCF footprint events was analyzed, predicted insulation value, and Rad21 ChlP-seq data. It was observed that only in the quartile with the strongest Rad21 signal was there a monotonically increasing relationship between CTCF occupancy and insulation score (data not shown). This finding is in line with the current mechanistic understanding that CTCF’s involvement in loopextrusion is cohesin-dependent and suggests that the model is identifying biologically plausible CTCF signals (Hansen et al. 2017; Dekker and Mirny 2024; Pugacheva et al. 2020).Example 3: FiberFold generalizes to a separate lymphoblastoid cell line

[0217] Broadly, 3D genomic contact maps are highly conserved across diverse cell types (Rao et al. 2014). However, even closely related cell lines can exhibit unique 3D genomic interactions that have physiological or pathologic relevance. For instance, different subtypes of acute myeloid leukemia (AML) show distinct 3D chromatin contacts, including neoTADs associated with structural variants that are functionally relevant (Xu et al. 2022). Alterations in 3D genome organization are observed in many other cancers (Sarni et al.2020; Dubois et al. 2022; Hnisz et al. 2016; Franke et al. 2016). Other indications where aberrant 3D genomic topologies have been observed include aging related neurodegeneration, asthma, and polydactyly (Dileep et al. 2023; Schmiedel et al. 2016; Lupianez et al. 2015).

[0218] Given that orthogonal models (Tan et al. 2023; Gao et al. 2024) are capable of predicting cell-type specific 3D genomic contacts when provided cell-type specific information on chromatin structure, it was reasoned that FiberFold would accurately predict de-novo specific 3D topologies. To assess the de novo predictive capacity of FiberFold, our GM12878-trained model was applied to predict Hi-C maps in K562 cells. Our model maintained high predictive accuracy in K562, achieving a pixel-wise Pearson correlation of 0.93 between experimental and predicted Hi-C maps (FIG. 3C). Additionally, FiberFold successfully identified differential 3D chromatin organization patterns between GM 12878 and K562 (FIG. 3 A), highlighting its capability to capture cell type-specific chromatin architecture.Example 4: Dissection of haplotype specific autosomal 3D genomic contacts with FiberFold

[0219] Human somatic cells contain two complete copies of the genome (haplotypes), totaling approximately 6 billion base pairs. Understanding the differential regulation of each haplotype is important for understanding the effects of human genetic and structural variation. For example, alterations in 3D genome architecture specific to one haplotype have been linked to the autosomal dominant inheritance of central iris hypoplasia (Sun et al. 2024). Studies in mice with a highly polymorphic X chromosome demonstrate a role for 3D genome organization in X chromosome inactivation in neuronal progenitor cells (Giorgetti et al. 2016).

[0220] However, current technologies limit our ability to study these haplotype-specific 3D genomic features in humans. In Rao et al., for instance, only 9.8% (478 million of 4.9 billion sequenced contacts)33167 / 70819could be assigned to a specific haplotype (Rao et al. 2014). Dip-C, an imputation based single-cell Hi-C protocol has shed light on the haplotype-resolved 3D structure of our genomes but is extremely expensive, impacted by sparse SNP coverage, and resolution limited in resolving haplotype-specific structures (Tan et al. 2018). Additionally, short-read chromatin profiling assays struggle to accurately map to regions with zero or two or more variants, making it challenging for predictive models to predict haplotype-specific 3D structures. To date, there are no predictive models that can generate haplotype-resolved 3D chromatin information.

[0221] Recently, deep sequencing using Fiber-seq has enabled the generation of haplotype-specific regulatory maps across the human genome. The long, single-molecule nature of Fiber-seq data permits phasing of 88% of GM12878 reads, resulting in the discovery of over 1,000 haplotype-specific regulatory elements, 64% of which could not be resolved using short-read methods (Vollger et al. 2024).

[0222] It was hypothesized that applying FiberFold to this high-resolution, haplotype-resolved Fiber-seq dataset would allow us to predict haplotype-specific 3D chromatin structure within a botdenecked GM 12878 cell line population. Given that 99% of accessible chromatin peaks are shared between haplotypes, it was expected that the predicted Hi-C maps for each haplotype would be largely similar across autosomes. Previous studies have indicated that the inactive X chromosome likely exhibits a more compact morphology, while the active X chromosome adopts a more diffuse morphology (Tan et al. 2018; Giorgetti et al. 2016). So, it was also hypothesized that there would be a significant difference between the 3D chromatin state between the active and inactive X chromosomes. To test this, FiberFold was applied to phased input features for each haplotype.

[0223] To quantify the level of topological divergence between haplotype-phased 3D contact maps, the mean absolute error (MAE) was calculated between the two predicted maps. As anticipated, analysis of the resulting haplotype-specific Hi-C maps revealed that most autosomal contact maps were nearly identical between haplotypes along the autosomes, particularly when compared to the difference between chromosome X haplotypes which are expected to have appreciable divergence (FIG. 4A and FIG. 4C). However, a subset of autosomal contact maps was identified where notable differences in 3D chromatin architecture were apparent (FIG. 4D). Importantly, when the Hi-C maps were averaged between haplotypes, the result closely matched experimentally derived Hi-C maps (FIG. 4B, autosome pearson R = 0.938), achieving accuracy comparable to that of predictions generated from bulk data. This agreement persisted even in regions with substantial differences between haplotypes, such as chrX (FIG. 4B, pearson R 0.91), highlighting FiberFold’s potential to capture haplotype-specific 3D chromatin structures with high fidelity.

[0224] In contrast, a marked difference was observed between haplotypes along the active (paternal) and inactive (maternal) X chromosome (FIGS. 4A, 4C, and 4D).

[0225] Given that the experimental Hi-C map should be a bulked average of the 3D contacts predicted across the 2 haplotypes, it was reasoned that stratification of the predicted, haplotype-specific contact contribution by the experimental pixel value should provide insight into which haplotype is driving contact33167 / 70819in our bulked maps. Indeed, it was observed that at the strongest experimentally derived contacts, the active paternal haplotype contributes a majority of highest predicted contact values (FIG. 4C). Additionally, at experimentally derived sites with the weakest contacts the predicted contact map from the active haplotype disproportionally contributed to the weakest predicted contact values (FIG. 4C).

[0226] Example 5: Using FiberFold to predict 3D interactions around a translocation breakpoint from a rare disease patient.

[0227] To test if FiberFold can predict 3D interactions, FiberFold was performed on data acquired from a patient with a rare germline chromosomal rearrangement (47) between chromosome X and chromosome 13, leading to a breakpoint that disrupts the NBEA gene. The patient has a global developmental delay, bilateral retinoblastomas, bilateral SNHL, lactic acidosis, hypotonia, dysmorphic facial features, and polymicrogyria.

[0228] First, Fiber-seq was performed, as described above, using data acquired from the patient sample (47). Next, phased de novo assembly of the genome was performed using the same Fiber-seq reads, and the assembly recapitulated the padent’s chromosome rearrangements (47). This assembly was used as the reference onto which the Fiber-seq reads were mapped for generating input features for FiberFold. Lastly, FiberFold was applied to the de novo assembly to predict 3D interactions as described above. The data is shown in FIGS. 5 A-8D and demonstrates that there are specific 3D chromatin arrangements that are specific to the genetic translocation affecting this patient.

[0229] FIG. 5 A demonstrates that the chrX haplotypes in this patient exhibit distinct, large-scale changes across sites of chrX homology, as seen in the haplotype maps (top and bottom).

[0230] For example, FIG. 5B depicts a zoomed-in haplotype map of the breakpoint. Homologous sequences along chrX have distinct 3D architecture.

[0231] In comparison, chromosome 13 also contains a translocation, however, the distinct 3D architecture is not so obvious across chromosome 13 haplotype comparisons as seen in FIG. 6 (top and bottom).

[0232] However, when plotting the differences between the haplotype maps (chrX vs chrl3) relative to the breakpoint, the distinct 3D architecture becomes clearer, as seen in FIG. 7 (top and bottom).

[0233] This is further demonstrated when the haplotype maps are compared to haplotype maps of chromosome 20 arms, used as controls. Chromosome 20 has no structural variants, and the haplotype maps are nearly identical.

[0234] Exemplary Computing Environment and Exemplary Computer-Implemented Methods

[0235] As discussed, FIG. IB depicts a block-flow diagram 100 of an artificial neural network architecture for training / operating a computational model used for genomics analysis. This model may be configured to process single-cell assay data, with specific inputs and processes that generate output relating to the 3D organization of the genome.IB33167 / 70819

[0236] The diagram 100 may include a "Fiber-seq (single-assay)" (block 102A) which refers to a type of genomic assay used as the input data. Below this, there are three interconnected blocks (blocks 102B) representing different types of genomic features or annotations that are derived from the Fiber-seq data: "mCpG," "CTCF Binding," and "FiRE (Access.)". These features / annotations may differ, in some aspects.

[0237] These genome annotations are then fed into a deep learning architecture, consisting of the following components, arranged vertically and connected by arrows pointing downwards:

[0238] 1. "mCpG | CTCF Orientation | CTCF Footprints | FiREs" (block 102C) - This block represents the combination of input features to be processed.

[0239] 2. "ConvlD Res-Block (13x)" (block 102D) - Indicates that there are 13 consecutive onedimensional convolutional residual blocks used for feature transformation.

[0240] 3. "Transformer (8x)" (block 102E) - Indicates that the model uses a transformer architecture, for capturing long-range dependencies, repeated 8 times.

[0241] 4. "Conv2D Dilated Res-Block (5x)" - Indicates the use of two-dimensional convolutional dilated residual blocks, repeated 5 times, (block 102F)

[0242] On the right side, there's a block that pertains to "Genetic and Structural Variants", (block 102G) which is used alongside "Hi-C maps / 2Mb Win I lOkb Res" as additional input data, for training the model to learn the 3D genomic structure (block 102H), in some aspects.

[0243] At the very right, the block flow diagram 100 provides details about the dataset split for machine learning purposes (block 1021): "Training" data comes from chromosome sets 1-9, 11-14, 16-22, and X; "Validation" uses chromosome 10; and "Test" uses chromosome 15. These splits in machine learning data enable the process and structure depicted in the block flow diagram 100 to train models and evaluate their performance on unseen data.

[0244] FIG. IF depicts a computing environment 104 for predicting and visualizing haplotype-specific 3D chromatin interactions and a plurality of chromatin features from a single-molecule, long-read sequencing assay. The computing environment 104 includes a computing system 106 designed to receive genomic DNA-protein contact information via a single-molecule, long-read sequencing assay, and generate visual representations using an artificial neural network model trained to predict haplotype-specific 3D chromatin structure information.

[0245] The computing system 106 comprises a processor 108, a memory 110, and a network interface controller (NIC) 112. The processor 108 may include any number of processors and / or processor types, such as central processing units (CPUs), graphics processing units (GPUs), and the like, configured to execute software instructions stored in the memory 110. The memory 110 may include volatile and / or nonvolatile fixed and / or removable memory, such as read-only memory (ROM), random access memory (RAM), and others. The memory 110 has stored thereon one or more sets of computer-executable33167 / 70819instructions, including a data collection module 114A, an Al analysis module 114B, a dependency identification module 114C, and an actionable task generation module 114D. In some aspects, the memory 110 may include more or fewer sets of instructions / modules.

[0246] The NIC 112 facilitates networking over the network between the computing system 100 and external data sources, user interfaces, and other systems used in the operation of the computing environment. For example, the NIC 112 may enable the computing system 100 to communicate bidirectionally via a network 120 The network may be a single communication network or may include multiple communication networks of one or more types, such as wired and / or wireless local area networks (LANs) and / or wide area networks (WANs) such as the Internet. For example, the computing system 106 may be communicatively coupled via the network 120 to one or more client computing devices 130. For example, the client computing devices 130 may display one or more visualizations based on output of the computing system 106 (e.g., one or more HiC maps).

[0247] In operation, the data collection module 114A collects genomic DNA-protein contact information via a single-molecule, long-read sequencing assay. The Al analysis module 114B applies artificial intelligence to analyze the collected data to generate relationships and predict stability across the layers. The dependency identification module 114C identifies dependencies within the ecosystem. The actionable task generation module 114D generates actionable tasks to optimize processes based on the analysis. These modules support operations across the computing environment 104, enabling the prediction and visualization of haplotype-specific 3D chromatin interactions and a plurality of chromatin features from a single-molecule, long-read sequencing assay.

[0248] The computing environment 104 may perform the steps of the block-flow diagram 100 of FIG. IB, in some aspects. In particular, the relational long-range interactions captured by the transformer layer at block 102E of FIG. IB enables the system 106 to train one or more models to understand 3D genomic structure, as these interactions reflect the complex spatial relationships between different genomic regions. These interactions enable accurate prediction of the 3D chromatin architecture, including the formation of topologically associating domains (TADs) and chromatin loops, which are components of the Hi-C map.

[0249] When the transformer layer captures these long-range interactions, it processes and integrates the information derived from the input features, such as chromatin accessibility, CTCF binding, CpG methylation, and the directionality of CTCF motifs, across extended genomic distances. This allows the model to understand how different regions of the genome interact with each other in three-dimensional space, even if they are far apart linearly along the DNA sequence.

[0250] The output from the transformer, which encapsulates these long-range relational interactions, is then fed into the 2D dilated residual convolutions in the decoding layer. These convolutions are designed to output a predicted Hi-C contact map at a resolution of 10 kilobases (for example) across a genomic window of approximately 2 million base pairs. The 2D dilated residual convolutions use the processed information from the transformer to construct a detailed map that represents the frequency of physical1033167 / 70819contacts between genomic regions, effectively translating the abstracted long-range interactions into a concrete, visual representation of the 3D chromatin structure.

[0251] Thus, the relational long-range interactions captured by the transformer layer may be directly related to the Hi-C map generated by the 2D dilated residual convolutions, as they provide foundational understanding of genomic organization used for accurately predicting the spatial arrangement of chromatin within the nucleus. This process enables the FiberFold model to reconstruct 3D genome architecture from single-molecule sequencing data, offering insights into the complex regulatory landscape of the genome.

[0252] To convert predicted long-range relational interactions into a visual format like a Hi-C contact map, several data structures and computational steps may be involved. This process may include generating input data representation, wherein data is encoded as multidimensional arrays or tensors. The input features (chromatin accessibility, CTCF binding, CpG methylation, and CTCF directionality) are represented as numerical values in a structured format, typically multidimensional arrays or tensors, where each dimension corresponds to a specific feature or genomic position. This process may further include transformer layer processing, wherein tensors with relational embeddings are generated. Specifically, the transformer layer may process the input data to capture long-range relational interactions. It may output tensors where the values represent learned embeddings that encapsulate the relationships between different genomic regions. These embeddings may be high-dimensional vectors that abstractly represent how regions interact over long distances. This process may include processing the output from the transformer using a decoding layer (e.g., 2D Dilated Residual Convolutions, or a 2D matrix or tensor for Hi-C contact map). The output from the transformer may be fed into the decoding layer, which uses 2D dilated residual convolutions to construct the Hi-C contact map. This process may include applying filters over the embeddings to predict the frequency of physical contacts between genomic regions. The output may be a 2D matrix (or tensor) where each cell represents the predicted contact frequency between pairs of genomic loci. The process may include a further step of Hi-C Contact Map Visualization, wherein a 2D heatmap or graphical representation is generated. Specifically, the 2D matrix representing the Hi-C contact map is visualized as a heatmap or another graphical format. In a heatmap, the x and y axes represent genomic positions, and the color intensity of each cell indicates the frequency of contact between the corresponding loci. This visual representation makes it easier to identify patterns such as TADs and chromatin loops. Further analysis might involve overlaying additional information or annotations on the Hi-C map, such as gene locations, TAD boundaries, or regions of interest. This step might use data structures like lists or dictionaries to map genomic coordinates to relevant biological features, facilitating the interpretation of the 3D chromatin architecture in a biological context. Throughout this workflow, the transition from raw input data to a visually interpretable Hi-C map involves complex transformations and representations. Each step abstracts and refines the information, ultimately leading to a detailed and informative depiction of the 3D genome organization. The computing system 106 may include one or more additional modules in the memory 110 for carrying out the aforementioned computational steps.

[0253] FIG. 1G depicts a computer-implemented method 180 for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions from a single-molecule, long-read sequencing assay. The1133167 / 70819method 180 includes receiving genomic DNA-protein contact information via a single-molecule, long-read sequencing assay (block 182A). This step involves capturing comprehensive genomic information, including chromatin accessibility, CpG methylation state, genomic DNA-binding protein binding site information, and genetic variation information, through a single assay, thereby streamlining the data acquisition process and reducing the complexity and cost associated with multi-assay approaches.

[0254] The method 180 further includes generating a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin structure information (block 182B). This involves processing the genomic DNA-protein contact information to generate the data set of haplotypespecific 3D chromatin interactions (block 182C) and processing the data set of haplotype-specific 3D chromatin interactions to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions (block 182D). This step leverages the power of deep learning to integrate and analyze complex genomic data, enabling the prediction of 3D chromatin structures and their visualization in a manner that is both informative and accessible to researchers.

[0255] The method 180 may include conducting the single-molecule, long -read sequencing assay to determine genomic DNA-protein contact information (block 182E). This step underscores the importance of using advanced sequencing technologies to capture high-resolution genomic data, which serves as a foundation for accurate prediction and visualization of 3D chromatin interactions, in some aspects.

[0256] The genomic DNA-protein contact information may include one or more or all of chromatin accessibility information, CpG methylation state information, genomic DNA-binding protein binding site information, and genetic variation information. This approach to data collection ensures that the model has access to a wide range of genomic features, enhancing its ability to accurately predict 3D chromatin structures.

[0257] The haplotype-specific 3D chromatin interactions information or the plurality of chromatin features may include 3D chromatin conformation information, and one or more or all of chromatin accessibility information, CpG methylation state information, genomic DNA-binding protein binding site information, and genetic variation information, in some aspects. This highlights the multifaceted nature of chromatin interactions, which are influenced by a variety of genomic and epigenomic factors.

[0258] In some aspects, the method 180 further includes processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to identify one or more patterns or markers indicative of a rare disease; and generating, via one or more processors, a report indicating the likelihood of the rare disease in a subject based on the processing. This application of the method demonstrates its potential for clinical relevance, offering a novel approach to diagnosing and understanding rare diseases through the lens of 3D chromatin architecture.

[0259] The method 180 may be performed, for example, by the computing environment 100 of FIG. IF, in some aspects. The method 180 may include steps for training one or more machine learning models, and / or for operating already trained models.IB33167 / 70819Additional Considerations

[0260] The various embodiments described above can be combined to provide further embodiments. All U.S. patents, U.S. patent application publications, U.S. patent application, foreign patents, foreign patent application and non-patent publications referred to in this specification and / or listed in the Application Data Sheet are incorporated herein by reference, in their entirety. Aspects of the embodiments can be modified if necessary to employ concepts of the various patents, applications, and publications to provide yet further embodiments.

[0261] These and other changes can be made to the embodiments in light of the above-detailed description. In general, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and the claims but should be construed to include all possible embodiments along with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the disclosure.References:1. Fudenberg, G., Kelley, D. R., & Pollard, K. S. (2020). Predicting 3D genome folding from DNA sequence with Akita. Nature Methods, 17(11), 1111-1117. https: / / doi.org / 10.1038 / s41592-020- 0958-x2. Gao, V. R., Yang, R., Das, A., Luo, R., Luo, H., McNally, D. R., Karagiannidis, I., Rivas, M. A., Wang, Z., Barisic, D., Karbalayghareh, A., Wong, W., Zhan, Y. A., Chin, C. R., Noble, W. S., Bilmes, J. A., Apostolou, E., Kharas, M. G., Beguelin, W., . . . Leslie, C. S. (2024). ChromaFold predicts the 3D contact map from single-cell chromatin accessibility. Nature Communications, 15(1). https: / / doi.org / 10.1038 / S41467-024-53628-03. Huang, C., Shuai, R. W., Baokar, P., Chung, R., Rastogi, R., Kathail, P., & loannidis, N. M.(2023). Personal transcriptome variation is poorly explained by current genomic deep learning models. Nature Genetics, 55(12), 2056-2059. https: / / doi.org / 10.1038 / s41588-023-01574-w 4. Jerkovic, I., & Cavalli, G. (2021). Understanding 3D genome organization by multidisciplinary methods. Nature Reviews Molecular Cell Biology, 22(8), 511-528.https : / / doi.org / l 0.1038 / s41580-021 -00362-w5. Kathail, P., Shuai, R. W., Chung, R., Ye, C. J., Loeb, G. B., & loannidis, N. M. (2024). Current genomic deep learning models display decreased performance in cell type-specific accessible regions. Genome Biology, 25(1). https: / / doi.org / 10.1186 / sl3059-024-03335-26. Schwessinger, R., Gosden, M., Downes, D., Brown, R. C., Oudelaar, A. M., Telenius, J., Teh, Y.W., Lunter, G., & Hughes, J. R. (2020). DeepC: predicting 3D genome folding using megabasescale transfer learning. Nature Methods, 17(11), 1118-1124. https: / / doi.org / 10.1038 / s41592-020- 0960-37. Stergachis, A. B., Debo, B. M., Haugen, E., Churchman, L. S., & Stamatoyannopoulos, J. A.(2020). Single-molecule regulatory architectures captured by chromatin fiber sequencing.Science, 565(6498), 1449-1454. https: / / doi.org / 10.1126 / science.aazl6468. Tan, J., Shenker-Tauris, N., Rodriguez-Hernaez, J., Wang, E., Sakellaropoulos, T., Boccalatte, F., Thandapani, P., Skok, J., Aifantis, I., Fenyb, D., Xia, B., & Tsirigos, A. (2023). Cell-type-specific prediction of 3D chromatin organization enables high-throughput in silico genetic screening. Nature Biotechnology, 41(8), 1140-1150. https: / / doi.org / 10.1038 / s41587-022-01612-89. Tan, L., Xing, D., Chang, C., Li, H., & Xie, X. S. (2018). Three-dimensional genome structures of single diploid human cells. Science, 361(6405), 924—928.https: / / doi.org / 10.1126 / science.aat564133167 / 70819Tan, L., Xing, D., Chang, C., Li, H., & Xie, X. S. (2018b). Three-dimensional genome structures of single diploid human cells. Science, 567(6405), 924—928.https: / / doi.org / 10.1126 / science.aat5641Tang, Z., Toneyan, S., & Koo, P. K. (2023). Current approaches to genomic deep learning struggle to fully capture human genetic variation. Nature Genetics, 55(12), 2021-2022. https: / / doi.org / 10.1038 / s41588-023-01517-5Vollger, M. R., Korlach, J., Eldred, K. C., Swanson, E., Underwood, J. G., Cheng, Y. H., Ranchalis, J., Mao, Y., Blue, E. E., Schwarze, U., Munson, K. M., Saunders, C. T., Wenger, A. M., Allworth, A., Chanprasert, S., Duerden, B. L.. Glass, I., Horike-Pyne, M., Kim, M., . . . Stergachis, A. B. (2023). Synchronized long-read genome, methylome, epigenome, and transcriptome for resolving a Mendelian condition. bioRxiv ( Cold Spring Harbor Laboratory). https: / / doi.org / 10.1101 / 2023.09.26.559521Vollger, M. R., Swanson, E. G., Neph, S. J., Ranchalis, J., Munson, K. M., Ho, C., Sedeno-Cortes, A. E., Fondrie, W. E., Bohaczuk, S. C., Mao, Y., Parmalee, N. L., Mallory, B. J., Harvey, W. T., Kwon, Y., Garcia, G. H., Hoekzema, K., Meyer, J. G., Cicek, M., Eichler, E. E., . . .Stergachis, A. B. (2024). A haplotype-resolved view of human gene regulation. bioRxiv ( Cold Spring Harbor Laboratory), https: / / doi.org / 10.1101 / 2024.06.14.599122Zhou, J. (2022). Sequence-based modeling of three-dimensional genome architecture from kilobase to chromosome scale. Nature Genetics, 54(5), 725-734. https: / / doi.org / 10.1038 / s41588-022-01065-4Dekker, J., & Mirny, L. A. (2024). The chromosome folding problem and how cells solve it. Cell, 187(23), 6424-6450. https: / / doi.Org / 10.1016 / j.cell.2024.10.026Krumm, A., & Duan, Z. (2018). Understanding the 3D genome: Emerging impacts on human disease. Seminars in Cell and Developmental Biology, 90, 62-77. https: / / doi.Org / 10.1016 / j.semcdb.2018.07.004Giorgio, E., Robyr, D., Spielmann, M., Ferrero, E., Di Gregorio, E., Imperiale, D., Vaula, G., Stamoulis, G., Santoni, F., Atzori, C., Gasparini, L., Ferrera, D., Canale, C., Guipponi, M., Pennacchio, L. A., Antonarakis, S. E., Brussino, A., & Brusco, A. (2015). A large genomic deletion leads to enhancer adoption by the lamin B 1 gene: a second path to autosomal dominant adult-onset demyelinating leukodystrophy (ADLD). Human Molecular Genetics, 24( 11), 3143-3154. https: / / doi.org / 10.1093 / hmg / ddv065Sarni, D., Sasaki, T., Tur-Sinai, M. I., Miron, K., Rivera-Mulia, J. C., Magnuson, B., Ljungman, M., Gilbert, D. M., & Kerem, B. (2020). 3D genome organization contributes to genome instability at fragile sites. Nature Communications, 77(1). https: / / doi.org / 10.1038 / s41467-020-17448-2Dubois, F., Sidiropoulos, N., Weischenfeldt, J., & Beroukhim, R. (2022). Structural variations in cancer and the 3D genome. Nature Reviews. Cancer, 22(9), 533-546. https: / / doi.org / 10.1038 / s41568-022-00488-9Dileep, V., Boix, C. A., Mathys, H., Marco, A., Welch, G. M., Meharena, H. S., Loon, A., Jeloka, R., Peng, Z., Bennett, D. A., Kellis, M., & Tsai, L. (2023). Neuronal DNA double-strand breaks lead to genome structural variations and 3D genome disruption in neurodegeneration. Cell, 186(20), 4404-442 l.e20. https: / / doi.Org / 10.1016 / j.cell.2023.08.038Xu, J., Song, F., Lyu, H., Kobayashi, M., Zhang, B., Zhao, Z., Hou, Y., Wang, X., Luan, Y., Jia, B., Stasiak, L., Wong, J. H., Wang, Q., Jin, Q., Jin, Q., Fu, Y., Yang, H., Hardison, R. C., Dovat, S., . . . Yue, F. (2022). Subtype-specific 3D genome alteration in acute myeloid leukaemia.Nature, 67 / (7935), 387-398. https: / / doi.org / 10.1038 / s41586-022-05365-xZheng, H., & Xie, W. (2019). The role of 3D genome organization in development and cell differentiation. Nature Reviews Molecular Cell Biology, 20(9), 535-550.https : / / doi .org / 10.1038 / s41580-019 -0132-4Shipony, Z., Marinov, G. K., Swaffer, M. P., Sinnott- Armstrong, N. A., Skotheim, J. M., Kundaje, A., & Greenleaf, W. J. (2020). Long-range single-molecule mapping of chromatinBO33167 / 70819accessibility in eukaryotes. Nature Methods, 17(3), 319-327. https: / / doi.org / 10.1038 / s41592-019-0730-2Altemose, N., Maslan, A., Smith, O. K., Sundararajan, K., Brown, R. R., Mishra, R., Detweiler, A. M., Neff, N., Miga, K. H., Straight, A. F., & Streets, A. (2022). DiMeLo-seq: a long-read, single-molecule method for mapping protein-DNA interactions genome wide. Nature Methods, 19(6), 711-723. https: / / doi.org / 10.1038 / s41592-022-01475-6Dubocanin, D., Hartley, G. A., Cortes, A. E. S., Mao, Y., Hedouin, S., Ranchalis, J., Agarwal, A., Logsdon, G. A., Munson, K. M., Real, T., Mallory, B. J., Eichler, E. E., Biggins, S., O’Neill, R. J., & Stergachis, A. B. (2023). Conservation of dichromatin organization along regional centromeres. bioRxiv ( Cold Spring Harbor Laboratory). https: / / doi.org / 10.1101 / 2023.04.20.537689Dubocanin, D., Cortes, A. E. S., Ranchalis, J., Real, T., Mallory, B., & Stergachis, A. B. (2022). Single-molecule architecture and heterogeneity of human telomeric DNA and chromatin. bioRxiv ( Cold Spring Harbor Laboratory), https: / / doi.org / 10.1101 / 2022.05.09.491186Grasberger, H., Dumitrescu, A. M., Liao, X., Swanson, E. G., Weiss, R. E., Srichomkwun, P., Pappa, T., Chen, J., Yoshimura, T., Hoffmann, P., Franca. M. M., Tagett, R., Onigata, K., Costagliola, S., Ranchalis, J., Vollger, M. R., Stergachis, A. B., Chong, J. X., Bamshad, M. J., . . . Refetoff, S. (2024). STR mutations on chromosome 15q cause thyrotropin resistance by activating a primate-specific enhancer of MIR7-2 / MIR1179. Nature Genetics, 56(5), 877-888. https: / / doi.org / 10.1038 / s41588-024-01717-7Vollger, M. R., Swanson, E. G., Neph, S. J., Ranchalis, J., Munson, K. M., Ho, C., Sedeno-Cortes, A. E., Fondrie, W. E., Bohaczuk, S. C., Mao, Y., Parmalee, N. L., Mallory, B. J., Harvey, W. T., Kwon, Y., Garcia, G. H., Hoekzema, K., Meyer, J. G., Cicek, M., Eichler, E. E., . . .Stergachis, A. B. (2024b). A haplotype-resolved view of human gene regulation. bioRxiv ( Cold Spring Harbor Laboratory), https: / / doi.org / 10.1101 / 2024.06.14.599122Vollger, M. R., Dishuck, P. C., Harvey, W. T., DeWitt, W. S., Guitart, X., Goldberg, M. E., Rozanski, A. N., Lucas, J., Asri, M., Abel, H. J., Antonacci-Fulton, L. L., Baid, G., Baker, C. A., Belyaeva, A., Billis, K., Bourque, G., Buonaiuto, S., Carroll, A., Chaisson, M. J. P., . . . Eichler, E. E. (2023). Increased mutation and gene conversion within human segmental duplications. Nature, 617(1960), 325-334. https: / / doi.org / 10.1038 / s41586-023-05895-yTan, J., Shenker-Tauris, N., Rodriguez-Hernaez, J., Wang, E., Sakellaropoulos, T., Boccalatte, F., Thandapani, P., Skok, J., Aifantis, I., Fenyb, D., Xia, B., & Tsirigos, A. (2023b). Cell-type-specific prediction of 3D chromatin organization enables high-throughput in silico genetic screening. Nature Biotechnology, 41(8), 1140-1150. https: / / doi.org / 10.1038 / s41587-022-01612-8 Kathail, P., Shuai, R. W., Chung, R., Ye, C. J., Loeb, G. B., & loannidis, N. M. (2024b). Current genomic deep learning models display decreased performance in cell type-specific accessible regions. Genome Biology, 25(1). https: / / doi.org / 10.1186 / sl3059-024-03335-2Sasse, A., Ng, B., Spiro, A. E., Tasaki, S., Bennett, D. A., Gaiteri, C., De Jager, P. L., Chikina, M., & Mostafavi, S. (2023). Benchmarking of deep neural networks for predicting personal gene expression from DNA sequence highlights shortcomings. Nature Genetics, 55(12), 2060-2064. https: / / doi.org / 10.1038 / s41588-023-01524-6Rowley, M. J., & Corces, V. G. (2018). Organizational principles of 3D genome architecture. Nature Reviews Genetics, 19(12), 789-800. https: / / doi.org / 10.1038 / s41576-018-0060-8 Guo, Y., Xu, Q., Canzio, D., Shou, J., Li, J., Gorkin, D. U., Jung, I., Wu, H., Zhai, Y., Tang, Y., Lu, Y., Wu, Y., Jia, Z., Li, W., Zhang, M. Q., Ren, B., Krainer, A. R., Maniatis, T., & Wu, Q. (2015). CRISPR inversion of CTCF sites alters genome topology and Enhancer / Promoter function. Cell, 162(4), 900-910. https: / / doi.Org / 10.1016 / j.cell.2015.07.038Hansen, A. S., Pustova, I., Cattoglio, C., Tjian, R., & Darzacq, X. (2017). CTCF and cohesin regulate chromatin loop stability with distinct dynamics. eLife, 6.https: / / doi.org / 10.7554 / elife.2577633167 / 70819Pugacheva, E. M., Kubo, N., Loukinov, D., Tajmul, M., Kang, S., Kovalchuk, A. L., Strunnikov, A. V., Zentner, G. E., Ren, B., & Lobanenkov, V. V. (2020). CTCF mediates chromatin looping via N-terminal domain-dependent cohesin retention. Proceedings of die National Academy of Sciences, 117(4), 2020-2031. https: / / doi.org / 10.1073 / pnas.1911708117Rao, S. S., Huntley, M. H., Durand, N. C., Stamenova, E. K., Bochkov, I. D., Robinson, J. T., Sanborn, A. L., Machol, I., Omer, A. D., Lander, E. S., & Aiden, E. L. (2014). A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell, 159(7), 1665— 1680. https: / / doi.Org / 10.1016 / j.cell.2014.ll.021Hnisz, D., Weintraub, A. S., Day, D. S., Valton, A., Bak, R. O., Li, C. H., Goldmann, J., Lajoie, B. R., Fan, Z. P., Sigova, A. A., Reddy, J., Borges-Rivera, D., Lee, T. I., Jaenisch, R., Porteus, M. H., Dekker, J., & Young, R. A. (2016). Activation of proto-oncogenes by disruption of chromosome neighborhoods. Science, 357(6280), 1454-1458.https: / / doi.org / 10.1126 / science.aad9024Franke, M., Ibrahim, D. M., Andrey, G., Schwarzer, W., Heinrich, V., Schbpflin, R., Kraft, K., Kempfer, R., Jerkovic, I., Chan, W., Spielmann, M., Timmermann, B., Wittier, L., Kurth, L, Cambiaso, P., Zuffardi, O., Houge, G., Lambie, L., Brancati, F., . . . Mundlos, S. (2016).Formation of new chromatin domains determines pathogenicity of genomic duplications. Nature, 538(7624), 265-269. https: / / doi.org / 10.1038 / naturel9800Schmiedel, B. J., Seumois, G., Samaniego-Castruita, D., Cayford, J., Schulten, V., Chavez, L., Ay, F., Sette, A., Peters, B., & Vijayanand, P. (2016). 17q21 asthma-risk variants switch CTCF binding and regulate IL-2 production by T cells. Nature Communications, 7(1).https : / / doi.org / l 0.1038 / ncomms 13426Lupianez, D. G., Kraft, K., Heinrich, V., Krawitz, P., Brancati, F., Klopocki, E., Horn, D., Kayserili, H., Opitz, J. M., Laxova, R., Santos-Simarro, F., Gilbert-Dussardier, B., Wittier, L., Borschiwer, M., Haas, S. A., Osterwalder, M., Franke, M., Timmermann, B., Hecht, J., . . .Mundlos, S. (2015). Disruptions of topological chromatin domains cause pathogenic rewiring of Gene-Enhancer interactions. Cell, 161(5), 1012-1025. https: / / doi.Org / 10.1016 / j.cell.2015.04.004 Tan, J., Shenker-Tauris, N., Rodriguez-Hernaez, J., Wang, E., Sakellaropoulos, T., Boccalatte, F., Thandapani, P., Skok, J., Aifantis, I., Fenyb, D., Xia, B., & Tsirigos, A. (2023c). Cell-type-specific prediction of 3D chromatin organization enables high-throughput in silico genetic screening. Nature Biotechnology, 41(8), 1140-1150. https: / / doi.org / 10.1038 / s41587-022-01612-8 Sun, W., Xiong, D., Ouyang, J., Xiao, X., Jiang, Y., Wang, Y., Li, S., Xie, Z., Wang, J., Tang, Z., & Zhang, Q. (2024). Altered chromatin topologies caused by balanced chromosomal translocation lead to central iris hypoplasia. Nature Communications, 75(1). https: / / doi.org / 10.1038 / s41467-024-49376-wGiorgetti, L., Lajoie, B. R., Carter, A. C., Attia, M., Zhan, Y., Xu, J., Chen, C. J., Kaplan, N., Chang, H. Y., Heard, E., & Dekker, J. (2016). Structural organization of the inactive X chromosome in the mouse. Nature, 535(7613), 575-579. https: / / doi.org / 10.1038 / naturel8589 Monteagudo-Sanchez, A., Albert, J. R., Scarpa, M., Noordermeer, D., & Greenberg, M. V. C. (2024). The impact of the embryonic DNA methylation program on CTCF-mediated genome regulation. Nucleic Acids Research, 52(18), 10934-10950. https: / / doi.org / 10.1093 / nar / gkae724 Nichols, M. H., & Corces, V. G. (2015). A CTCF code for 3D genome architecture. Cell, 162(4), 703-705. https: / / doi.org / 10.1016 / j-cell.2015.07.053Vollger, M.R., Korlach, J., Eldred, K.C. et al. Synchronized long-read genome, methylome, epigenome and transcriptome profiling resolve a Mendelian condition. Nat Genet (2025). https: / / doi.Org / 10.1038 / s41588-024-02067-0IB

Claims

33167 / 70819CLAIMS1. A computer-implemented method for predicting and visualizing a data set of haplotypespecific 3D chromatin interactions from a single-molecule, long-read sequencing assay, said method comprising the steps of:(a) receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay;(b) generating, via one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin structure information by:i. processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions; and ii. processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to generate the visual representation of the data set of haplotypespecific 3D chromatin interactions.

2. The computer-implemented method of claim 1, further comprising conducting the singlemolecule, long-read sequencing assay to determine genomic DNA-protein contact information.

3. A computer-implemented method for concurrently predicting and visualizing a data set of a plurality of chromatin features from a single-molecule, long-read sequencing assay, said method comprising the steps of:(a) receiving, via one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay;(b) generating, via one or more processors, a visual representation using an artificial neural network model trained to predict a plurality of chromatin features by:i. processing, via one or more processors, the genomic DNA-protein contact information to generate the data set of a plurality of chromatin features; and ii. processing, via one or more processors, the data set of a plurality of chromatin features to generate the visual representation of the data set of a plurality of chromatin features.

4. The computer-implemented method of claim 3, further comprising conducting the singlemolecule, long-read sequencing assay to determine genomic DNA-protein contact information.33167 / 708195. The method of any one of claims 1-4, wherein the genomic DNA-protein contact information comprises one or more or all of:i. chromatin accessibility information;ii. CpG methylation state information;iii. genomic DNA-binding protein binding site information; andiv. genetic variation information.

6. The method of any one of claims 1-4, wherein the genomic DNA-protein contact information comprises one or more or all of:i. Fiber-inferred Regulatory Element (FiRE) information;ii. CpG methylation information;iii. CTCF footprinting score; andiv. CTCF direction information.

7. The method of any one of claims 1-6, wherein the haplotype-specific 3D chromatin interactions information or the plurality of chromatin features comprises:(a) 3D chromatin conformation information, and(b) one or more or all of:(i) chromatin accessibility information;(ii) CpG methylation state information;(iii) genomic DNA-binding protein binding site information; and(iv) genetic variation information.

8. The method of claim 7. wherein the 3D chromatin conformation information comprises Hi-C contact matrices.

9. The method of claim 8, wherein the Hi-C contact matrices comprise a map at a resolution of approximately 10 kilobases across a genomic window of approximately 2 million base pairs.

10. The method of claim 7, wherein the genomic DNA-protein binding site information comprises binding site locations for DNA binding-proteins.

11. The method of claim 10, wherein the DNA-binding proteins are selected from the group consisting of transcription factor proteins, chromatin binding proteins, and chromatin-associated proteins.ii33167 / 7081912. The method of claim 11, wherein the DNA-binding protein is CTCF.

13. The method of claim 7, wherein the chromatin accessibility information comprises DNA modification data.

14. The method of claim 7, wherein the CpG methylation state information comprises methylated cytosine (5mC), DNA hydroxylmethylaed cytosine (5hmC) information.

15. The method of claim 7, wherein the genetic variation information comprises regulatory DNA variation.

16. The method of any one of claims 1-15, wherein the genomic DNA-protein contact information further comprises epigenomic DNA-protein contact information.

17. The method of any one of claims 1-2, wherein the haplotype-specific 3D genomic structure information further comprises cell-specific 3D genomic structure information.

18. The method of any one of claims 3-4, wherein the plurality of chromatin features comprises cell-type specific and haplotype-specific chromatin features.

19. The method of any one of claims 1-18, further comprising:processing, via one or more processors, the data set of haplotype-specific 3D chromatin interactions to identify one or more patterns or markers indicative of a rare disease; and generating, via one or more processors, a report indicating the likelihood of the rare disease in a subject based on the processing.

20. A computer-implemented method of training an artificial neural network model to predict (a) haplotype-specific 3D chromatin interactions and / or (b) a plurality of chromatin features, comprising:receiving, via one or more processors, training genomic DNA-protein contact information via a single-molecule, long-read sequencing assay; andprocessing the training genomic DNA-protein contact information using an artificial neural network model to learn one or more training parameters, until the artificial neural network model learns to accurately predict a data set of haplotype-specific 3D chromatin i nteractions and / or a plurality of chromatin features.33167 / 7081921. A computing system for predicting and visualizing a data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features from a single-molecule, long-read sequencing assay, comprising:one or more processors, andone or more memories, having stored thereon computer-executable instructions that, when executed, cause the computing system to:(a) receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay;(b) generate, via the one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features by:i. processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features; andii. processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.

22. A non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed, cause a computer to:(a) receive, via the one or more processors, genomic DNA-protein contact information via a single-molecule, long-read sequencing assay;(b) generate, via the one or more processors, a visual representation using an artificial neural network model trained to predict haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features by:i. processing, via the one or more processors, the genomic DNA-protein contact information to generate the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features; andii. processing, via the one or more processors, the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features to generate the visual representation of the data set of haplotype-specific 3D chromatin interactions and / or a plurality of chromatin features.