Methods and devices for analyzing chromatin interaction data
By segmenting genomic elements into boxes of different sizes and generating variable sized box-pair matrices, the problems of poor interactions in the existing technology for detecting distal enhancer and high-energy consumption are solved, and high-precision genomic element contact mapping and long-range interaction detection are achieved.
Patent Information
- Application Number
- CN201980034320.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-22
- Filing Date
- 2019-03-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-03-21
AI Technical Summary
The prior art does not perform well in detecting distal enhancer interactions of genomic elements, and requires a large amount of computing resources and memory at high resolution, making it impossible to effectively detect long-range cis and trans interactions.
By segmenting genomic elements into boxes of different sizes and generating variable sized boxes pair matrix, identifying paired contact subsets in box pairs, calculating interaction frequencies, and normalizing according to density functions, an enriched or depleted set of contacts is generated.
High-precision mapping of genomic element contacts is achieved, reducing memory requirements and computing resources, and able to detect long-range cis and trans interactions mediated by functional elements, especially significantly improving detection capabilities between TADs.
Smart Images

Figure CN112272849B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority and the filing date of U.S. Provisional Application Serial No. 62 / 646,433, filed on March 22, 2018, entitled "Method and Apparatus for Analysis of Chromatin Interaction Data", the entire disclosure of which is hereby expressly incorporated herein by reference. Field of the Invention
[0003] This application relates to chromatin interaction analysis, and more particularly, to a method and system for efficiently identifying contacts of genomic elements using variable-sized bins with statistical techniques. Background of the Invention
[0004] Today, genomic element contacts are mapped using Hi-C sequencing or other similar methods, such as genome architecture mapping, ChIA-PET, 4C, 5C, Combi-C, Micro-C, etc. In such methods, paired-end sequencing reads represent pairs of genomic positions that have spatial contacts in a biological cell sample that is processed to generate a Hi-C sequencing library. Multiple such paired-end reads are compiled into a graph or frequency matrix representing the frequency of spatial interaction of genomic position pairs.
[0005] To perform the mapping, the data set is compiled into fixed-sized bins, which are uniformly sized portions of the genome that are adjacent to each other. However, this method requires the selection of a fixed resolution, which makes it inherently limited. At low resolution, the loci of interest are combined with unrelated loci, while other loci are split in half. Genes are typically regulated by enhancer elements that are far away in sequence space from the gene, called distal cis, or located on different chromosomes, called trans. However, due to data sparsity, these methods perform poorly in detecting distal enhancer interactions. Trans and distal cis interactions exhibit severe data sparsity because read pairs in a linear genome are mapped into a square genome with an area of over nine million megabase pairs (Mb). At high resolution, this method is very memory-intensive and requires a large amount of computational resources.
[0006] In addition, the read - segment density varies over five orders of magnitude with genomic distance, and most of the measured interactions are concentrated on the axis. Thus, for a fixed bin, fine resolution will result in more than 99.9% of the entries in the genome - wide matrix being empty, while coarse resolution will not benefit at all from the mediation of long - range contacts of functional elements, cutting them into pieces and combining them with adjacent sequence regions, thus dissipating the signals that researchers hope to detect.
[0007] Topologically associating domains (TADs) have been identified as effective spatial and functional genomic units. Approximately 80% of the sequence length of the human genome is partitioned into approximately 2,500 TADs, which are highly robust and conserved across human cell types, different humans, and disease states. TADs also function as replication domains. In addition, TADs mediate long - range spatial interactions: the contact frequency in any given part of the genomic square will be more closely related to more distant sequence parts within the same TAD pair than to proximal sequence parts across TAD boundaries.
[0008] Recent work has begun to address the drawbacks of fixed bins. The SHAMAN package does away with fixed bins and matrix compilation and employs a different approach to detect contacts. It uses a base - pair - resolution sparse matrix and then generates a random matrix that meets distance - frequency and edge - coverage criteria sampled from the real matrix. It uses this random matrix to compare with the real matrix, generating p - values, and then compares the p - values with FDR statistics to address random errors in the Hi - C matrix. However, the p - values are generated based on the Kolmogorov - Smirnov D statistics of the K - nearest - neighbor clustering density around each individual read - pair in the database. Pairs with significantly dense K - nearest neighbors can be considered enriched. Thus, choosing the K value for a particular experiment represents an important trade - off between resolution and statistical power, much like bin - size selection in traditional Hi - C compilation.
[0009] For distal contacts, the SHAMAN package is affected because it does not account for the mediation of contacts by large sequence elements. The K - nearest neighbors of a particular read - pair may not be significantly enriched, while the entire TAD pair in which the read - pair lies may be enriched. For a suitable K value, these would be approximately consistent, but SHAMAN does not provide a way to choose this K, which would change the genome - wide situation in any case. In addition, read - pairs adjacent to TAD pairs with strong clustering can "stowaway" on densely sequenced reads that are close in sequence, thus producing near - neighbor spillover contact detection in the way of fixed bins.
[0010] Accordingly, compared to existing systems, a system for precisely mapping genomic element contacts is needed to maintain high precision and reduce memory requirements and computational resources. There is also a need for a system that partitions related loci in the same bin and does not split a locus in half to detect long-range cis- and trans-interactions mediated by functional elements. Summary of the Invention
[0011] To map genomic element contacts, a chromatin interaction system obtains a set of genomic elements (e.g., loci) and partitions the set of elements into bins of different sizes. The bins can be selected to include related genomic elements in the same bin and prevent splitting a genomic element in half. For example, each bin can correspond to a contiguous segment of a deoxyribonucleic acid (DNA) sequence and can represent a cutting site increment or a functional element such as a gene, a chromatin state segment, a loop domain, a chromatin domain, a topologically associating domain (TAD), etc. Then two sets of bins (e.g., a first set of bins corresponding to chromosome 1 and a second set of bins corresponding to chromosome 8) are selected and placed in n × m a matrix (a square genomic region) to generate a set of bin pairs. Thus, the square genomic region can have variable size and shape. In some embodiments, the two sets of bins are the same (e.g., each corresponding to chromosome 1). In any case, the chromatin interaction system uses, for example, a binary search tree to identify position pairs corresponding to paired-end reads or other spatially interacting positions that may contain the bin pairs that contain the position pairs (i.e., where one of the bins in the pair contains one of the loci and the other bin contains the other locus) (e.g., Chr1:950000 and Chr8:15000).
[0012] Then, based on the genomic element contacts within the corresponding bin pairs, the interaction frequency of each bin pair is generated. Additionally, the interaction frequency is normalized according to the variation of the density of paired contacts within each bin pair with genomic distance. More specifically, the variation of the density of paired contacts with genomic distance can be determined to generate a density function. Such a function can be corrected for the GC sequence percentage in a specific bin sequence, the sequence coverage of a specific bin sequence in a Hi-C sequencing dataset, or other suitable factors for Hi-C normalization. Then, for a specific bin pair, the density function is integrated over the square genomic region of the bin pair to determine the expected density of the bin pair. Then, the expected density of the bin pair can be compared with the actual density (i.e., the number of paired contacts within the square genomic region of the bin pair) using, for example, a statistical test such as a Poisson distribution p-value (e.g., to which the Benjamini false discovery rate can be applied) to generate a set of enriched and depleted chromatin contacts in an adjusted manner for distance (and other suitable features) on a local or genome-wide basis. The chromatin interaction system can then provide an indication of bin pairs with enriched or depleted contacts for display on a user interface.
[0013] In this way, the enriched or depleted contacts can be used to predict the molecular phenotype of a subject based on the spatial interactions of loci within the corresponding genome. The enriched or depleted contacts can also be used to model the 3D and 4D structures of chromosomes and identify altered TAD boundaries and spatial interactions in tissue samples to determine genetic diseases or oncology. Additionally, the enriched or depleted contacts can be used to determine whether a pair of loci in a specific tissue or cell line interact. Moreover, the enriched or depleted contacts can be used to locate trans and distal cis-binding partners of functional TADs and construct a spatial contact network. This embodiment advantageously detects long-range contacts not found in existing systems using traditional methods in the same dataset of comparable bins with fixed size and spacing. In experiments, compared with traditional methods, the embodiments of the present invention detected 2.5 times more significant long-range cis interactions between TADs.
[0014] In addition, compared with traditional methods, by using variable bin sizes, the present embodiment advantageously reduces the memory requirements and computational resources for mapping spatial interactions. As with traditional methods, when using fixed-sized bins to map spatial interactions, a resolution high enough to ensure that the boundaries of each bin are within the selected range must be chosen. For example, when using fixed-sized bins to map spatial interactions between TADs, the resolution must be chosen such that the DNA sequence fragments corresponding to each bin are shorter than the shortest TAD. In other words, if the smallest TAD is 100 kilobases (kB), then the resolution of the fixed-sized bins must be at most 100 kB. To improve accuracy, the resolution is typically much smaller than the shortest TAD (e.g., 1 kB or 10 kB), and several bins are aggregated. On the other hand, using variable bin sizes, the present embodiment selects bins such that each bin represents a different TAD (or other functional elements, such as genes, chromatin state segments, loop domains, chromatin domains, etc.), regardless of its length. For example, if the average TAD is 1 megabase (MB) long, the present embodiment can effectively use a 1 MB resolution to map the spatial interactions of the same functional element (TAD), whereas the resolution for mapping spatial interactions between TADs using traditional methods is 1 kB or 10 kB. Thus, compared with traditional methods, the present embodiment has lower memory intensity and computational complexity. Where n is the number of read pairs, k is the number of bins in the square matrix, the complexity of each step is approximately O(n) for alignment and quality control, approximately O(n*log(k)) for compilation, approximately O(k) for integration, and approximately O(k^2) for statistical control and data output. The following references Figure 8 will be further described in detail for each of these steps.
[0015] In one embodiment, a computer-implemented method for analyzing the spatial and temporal organization of chromatin is provided. The method includes obtaining a set of paired contacts of genomic elements, partitioning the genomic elements into a plurality of bins, wherein the bin sizes of the plurality of bins are inconsistent, identifying a first set of the plurality of bins and a second set of the plurality of bins, and generating n × m a matrix of bin pairs, wherein n corresponds to the first set of the plurality of bins, and m corresponds to the second set of the plurality of bins. The method further includes: identifying a subset of the paired contacts within each bin pair of the bin pairs, determining the interaction frequency of each bin pair of the bin pairs, normalizing each interaction frequency of the interaction frequencies to generate a normalized interaction frequency for each bin pair, and providing a map of chromatin interactions for display on a user interface, including an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies.
[0016] In another embodiment, a computing device for analyzing the spatial and temporal organization of chromatin is provided. The computing device includes a communication network, one or more processors, and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon. When executed by the one or more processors, the instructions cause the system to obtain a set of paired contacts of genomic elements, divide the genomic elements into a plurality of bins, where the bin sizes of the plurality of bins are inconsistent, identify a first set of the plurality of bins and a second set of the plurality of bins, and generate n × m a matrix of bin pairs, where n corresponds to the first set of the plurality of bins, and m corresponds to the second set of the plurality of bins. The instructions further cause the system to identify a subset of paired contacts within each bin pair of the bin pairs, determine the interaction frequency of each bin pair of the bin pairs, normalize each interaction frequency of the interaction frequencies to generate a normalized interaction frequency for each bin pair, and provide, via the communication network, a map of chromatin interactions for display on a user interface, including an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A shows a block diagram of a computer network and system on which an exemplary chromatin interaction system according to the presently described embodiments can operate;
[0018] Figure 1B is a block diagram of an exemplary chromatin interaction server that can operate in a Figure 1A system according to the presently described embodiments;
[0019] Figure 1C is a block diagram of an exemplary client device that can operate in a Figure 1A system according to the presently described embodiments;
[0020] Figure 2 depicts an example set of bins according to the presently described embodiments, each bin corresponding to a contiguous segment of a locus in a chromosome;
[0021] Figure 3 depicts an example square genomic matrix of bin pairs according to the presently described embodiments;
[0022] Figure 4 depicts an example spatial interaction map of bin pairs and corresponding normalized interaction frequencies according to the presently described embodiments;
[0023] Figure 5depicts example density functions according to the presently described embodiments, each density function representing the density of paired contacts as a function of genomic distance;
[0024] Figure 6 is a block diagram representing an exemplary process for generating chromatin interaction data from a biological sample of a subject;
[0025] Figure 7A describes an example display of enriched contacts and corresponding loci and / or single nucleotide polymorphisms (SNPs) associated with molecular phenotypes according to the presently described embodiments;
[0026] Figure 7B depicts an example chromatin interaction network generated from spatial interaction data according to the presently described embodiments;
[0027] Figure 8 shows a flowchart representing an exemplary method for analyzing chromatin spatial organization according to the presently described embodiments; and
[0028] Figure 9 shows a Venn diagram comparing the number of contacts identified by a chromatin interaction system and the number of contacts identified by an alternative system. DETAILED DESCRIPTION
[0029] Although the following text sets forth a detailed description of many different embodiments, it should be understood that the legal scope of this description is defined by the words of the claims set forth at the end of this disclosure. The detailed description should be interpreted as merely exemplary and does not describe every possible embodiment, as it would be impractical, if not impossible, to describe every possible embodiment. Many alternative embodiments can be implemented using current technology or technology developed after the filing date of this patent application, and such embodiments will still fall within the scope of the claims.
[0030] It should also be understood that unless a term is explicitly defined in this patent by a sentence such as "As used herein, the term '______' is defined herein to mean... " or a similar sentence, there is no intent to limit the meaning of the term, whether explicitly or by implication, beyond its ordinary or common meaning, and the term should not be construed as being limited in scope by any statement made in any section of this patent (except the language of the claims). For any term that is referred to in this patent in a manner consistent with a single meaning in the claims recited at the end of this patent, this is done solely for clarity so as not to confuse the reader and is not intended to limit the claim term by implication or otherwise to that single meaning. Finally, unless a claim element is limited by reference to the word "means" and a recited function without any structure, no claim element is intended to be construed under the provisions of 35 U.S.C. § 112, paragraph 6.
[0031] Accordingly, as used herein, the term "read pair" or "paired genomic element contact" may refer to a pair of loci within a corresponding part of the genome. For example, a read pair can be Chr1:950000, Chr8:15000.
[0032] In addition, as used herein, the term "genomic element" may refer to a specific unit of a deoxyribonucleic acid (DNA) sequence. Genomic elements can be reads, loci within a chromosome, base pairs, and the like.
[0033] As used herein, the term "bin" may refer to a continuous segment of DNA sequence within the genome of a human or other organism that is considered a unit for chromatin contact analysis. Such bins can be selected for various purposes according to a particular analysis and can, for example, contain sequence regions corresponding to TADs, inter-TAD fragments, genes, superTADs, subTADs, loop domains, common domains, enhancer bodies, and / or promoter bodies, exons, introns, chromatin state fragments, restriction enzyme fragments, genetically engineered insertion constructs, translocation elements, LADs, NORs, SARs, MARs, or combinations of these elements.
[0034] Furthermore, as used herein, the term "bin pair" may refer to a rectangular region in a square genomic space corresponding to the Cartesian product of two bins represented in a linear genomic space. This term may also refer to two bins considered as a pair, rather than the region of the square genomic space represented by the pair.
[0035] As used herein, the term "subject" may refer to any human or other organism or combination thereof whose health, lifespan, or other biological outcome is the object of clinical or research interest, study, or endeavor.
[0036] As used herein, the term "pharmacological phenotype" can refer to any distinguishable phenotype that may affect drug therapy, subject lifespan and outcomes, quality of life, etc. in clinical care, the management and finance of clinical care, and pharmaceutical and other medical and biomedical research on humans and other organisms. Such phenotypes can include pharmacokinetic (PK) and pharmacodynamic (PD) phenotypes, all phenotypes that include the rates and characteristics of drug absorption, distribution, metabolism, and excretion (ADME), as well as drug responses related to drug efficacy, drug treatment dose, half-life, plasma level, clearance, etc., and adverse drug events, adverse drug reactions, and the corresponding severity of adverse drug events or adverse drug reactions, organ damage, drug abuse and dependence and their likelihood, and body weight and its changes, mood and behavior changes and disturbances. Such phenotypes can also include favorable and adverse responses to drug combinations, drug-gene interactions, social and environmental factors, dietary factors, etc. The phenotypes can also include compliance with pharmacological or non-pharmacological treatment regimens. The phenotypes may also include medical phenotypes, such as a subject's propensity to contract a certain disease or complication, the outcome and prognosis of the disease, whether a subject will exhibit specific disease symptoms, and the subject's outcomes (such as lifespan, clinical scores and parameters, test results, healthcare expenditures), and other phenotypes.
[0037] As used herein, the term "molecular phenotype" can refer to a pharmacological phenotype or any other phenotype of a human or other organism that can be measured or distinguished, either individually or collectively, at a specific point in time or retrospectively and that can be detected, evaluated, estimated, or modified, affected, or altered for any useful purpose.
[0038] Generally, techniques for mapping spatial interactions can be implemented in one or more client devices, one or more web servers, or a system that includes a combination of these devices. However, for clarity, the following examples mainly focus on one embodiment in which a chromatin interaction server obtains a set of read pairs or paired genomic element contacts, such as a pair of loci (e.g., Chr1:950000, Chr8:15000). The chromatin interaction server also obtains a set of bins, which are non-overlapping contiguous segments of DNA. In some embodiments, the set of read pairs and / or the set of bins can be obtained from a client device of a researcher or healthcare professional. For example, a researcher or healthcare professional can select a specific set of read pairs for analysis. Additionally, a researcher or healthcare professional can select a specific set of bins at a specific resolution. For example, a researcher or healthcare professional can select a set of bins in which each bin represents a different TAD. In another example, a researcher or healthcare professional can select a set of bins in which each bin represents a different gene.
[0039] In any case, the bins can have variable sizes or sequence lengths and can be used to divide one or more parts of a genome into segments, where each bin represents a cleavage site increment or a functional element such as a gene, a chromatin state segment, a loop domain, a chromatin domain, a TAD, etc. (e.g., Chr1:1000-2000). Then the chromatin interaction server selects two sets of bins (e.g., the first set of bins corresponds to chromosome 2, the second set of bins corresponds to chromosome 5, two identical sets of bins each correspond to chromosome 3, etc.). Each set can represent n × m the axes of a square genomic matrix, where the first set contains n bins and the second set contains m bins. Thus, the square genomic matrix can contain n × m bin pairs, where a bin pair is an entry or a rectangle in the square genomic matrix (e.g., Chr1:1000-2000*Chr8:10000-20000). The chromatin interaction server can also assign each read pair to a corresponding bin pair. For example, the read pair Chr1:1010, Chr8:15000 can be assigned to the bin pair Chr1:1000-2000*Chr8:10000-20000 because this read pair lies within the boundaries of the square genomic region corresponding to the bin pair. Specifically, the bins and multiple sets of bins can be constructed to analyze cis (i.e., the bins in a bin pair are on the same chromosome) or trans (i.e., the bins in a bin pair are on different chromosomes) contacts.
[0040] Furthermore, the chromatin interaction server can determine the interaction frequency of each bin pair based on the density of read pairs in the bin pair. Each interaction frequency can be normalized by calculating a density function of the entire set of read pairs as a function of genomic distance or the distance between the two loci (also referred to herein as "reads") in each read pair. This function can be corrected for the percentage of GC sequence in a particular bin sequence, the sequence coverage of a particular bin sequence in a Hi-C sequencing dataset, the density of cleavage sites in a particular sequence region, or other suitable factors for Hi-C normalization. For each bin pair, the density function can be integrated over the rectangular region of the bin pair to determine the expected density of the bin pair. Then, a statistical method (such as a Poisson distribution p-value) (e.g., to which a Benjamini false discovery rate can be applied) can be used to compare the expected density of the bin pair with the actual density of the bin pair (the density of read pairs in the bin pair). Based on the p-value and the false discovery rate, the chromatin interaction server can identify bin pairs with enriched contacts and bin pairs with depleted contacts. Then, the chromatin interaction server can provide an indication of the bin pair and an indication of its corresponding normalized interaction frequency (such as a p-value) to the client device for display. The client device can present a spatial interaction map (such as a heat map), where bin pairs with higher normalized interaction frequencies are represented by darker colors. In other embodiments, the client device can present a numerical indication of the normalized interaction frequency, such as the p-value for each bin pair. Thus, a healthcare professional or researcher can view the spatial interaction map or numerical indication on her client device to view bin pairs with enriched or depleted contacts.
[0041] In some embodiments, the actual read counts of multiple sets of contacts in bin pairs, such as corresponding to different biological cell systems or different physiological conditions, can be analyzed together. For example, such systems can consist of two different tissues in the human body, tissue samples from two different individuals, a cell line that has been subjected to a medical treatment compared to a control sample, or multiple cell cycle conditions (such as interphase or metaphase) or cell differentiation states of the same tissue, cell line, or organism. Such an analysis can determine, for example, a set of differential contacts between a pair of contact sets by, for example, separately comparing the enriched and depleted contacts from each dataset. Differential contacts can also be determined, for example, by using a multiple sampling distribution of a Poisson distribution or other statistical distribution to generate a p-value corresponding to the probability of an accidentally observed differential interaction frequency, which can then be corrected with a false discovery rate (FDR) or other method, as described herein.
[0042] Reference Figure 1A, the exemplary chromatin interaction system 100 identifies enriched or depleted contacts within a selected set of read pairs and bins of bins. Bin pairs with enriched or depleted contacts can be highlighted and displayed on a client device of a researcher or healthcare professional. For example, bin pairs with enriched or depleted contacts can be highlighted in a spatial interaction map. A healthcare professional or researcher can then use the enriched or depleted contacts to predict the molecular phenotype of a subject based on the spatial interactions of loci within the corresponding genome. Such prediction can be made using other methods described elsewhere using spatial interactions and other forms of clinical and / or panoramic information. Enriched or depleted contacts can also be used to model the 3D and 4D structures of chromosomes or genomes within the cell nucleus, and to identify altered TAD boundaries and spatial interactions in tissue samples to determine genetic diseases or oncology. In addition, enriched or depleted contacts can be used to determine whether a pair of loci in a particular tissue or cell line interact.
[0043] The chromatin interaction system 100 includes a chromatin interaction server 102 and a plurality of client devices 106 - 116 that can be communicatively connected via a network 130, as described below. In one embodiment, the chromatin interaction server 102 and the client devices 106 - 116 can communicate over the communication network 130 via wireless signals 120, which can be any suitable local area network or wide area network, including WiFi networks, Bluetooth networks, cellular networks (such as 3G, 4G, Long Term Evolution (LTE), 5G), the Internet, etc. In some cases, the client devices 106 - 116 can communicate with the communication network 130 via an intervening wireless or wired device 118, which can be a wireless router, a wireless repeater, a base transceiver station of a mobile phone provider, etc. By way of example, the client devices 106 - 116 can include a tablet computer 106, a sequencer 107, a network-enabled cellular phone 108, a sequence database 109 containing sequence data from published literature, clinical trials, consortia, academia, etc., a personal digital assistant (PDA) 110, a mobile device smart phone 112 (also referred to herein as a "mobile device"), a laptop computer 114, a desktop computer 116, a wearable biosensor, a portable media player (not shown), a tablet computer, any device configured for wired or wireless RF (radio frequency) communication, etc. In addition, any other suitable client device that records genomic data of a subject, receives multiple sets of read pairs / bins, or displays an indication of enriched contacts can also communicate with the chromatin interaction server 102.
[0044] Each of client devices 106-116 may interact with the chromatin interaction server 102 to provide a selected set of read pairs and / or selected sets of bins. For example, sequencer 107 may generate sequence data provided to chromatin interaction server 102. In yet another example, sequence database 109 may provide pre-existing sequence data to chromatin interaction server 102 generated from, for example, published literature, clinical trials, consortia, academia, etc. Chromatin interaction server 102 may then identify a set of read pairs and / or sets of bins from the sequence data. Each client device 106-116 may also interact with chromatin interaction server 102 to receive one or more indications of bin pairs and an indication of the normalized interaction frequency of the bin pairs. The indication may be a numerical indication, and the client device may present the numerical indication via a user interface for display to a healthcare professional or researcher. The client device may also present a graphical representation of the bin pairs and the normalized interaction frequency, such as a heat map, where square genomic regions corresponding to bin pairs with a higher normalized interaction frequency (e.g., enriched contacts) are highlighted in a darker color.
[0045] In an example implementation, chromatin interaction server 102 may be a cloud-based server, application server, web server, etc., and includes a memory 150, one or more processors (CPUs) 142 (such as a microprocessor coupled to memory 150), a network interface unit 144, and an I / O module 148, which may be, for example, a keyboard or a touch screen.
[0046] Chromatin interaction server 102 may also be communicatively connected to a database 154 of read pairs and bins. For example, database 154 may store a collection of bins across a genome or a portion of a genome, where each bin represents a set of loci corresponding to a TAD (e.g., Chr1:1280000-1840000). In some embodiments, chromatin interaction server 102 may retrieve a set of read pairs and / or sets of bins from database 154. In other embodiments, the set of read pairs and / or sets of bins are provided by client devices 106-116. In still other embodiments, chromatin interaction server 102 may retrieve bins from the database, and a healthcare professional or researcher may select sets of bins for each axis of a square genomic matrix (e.g., a first set of bins corresponding to chromosome 1 and a second set of bins corresponding to chromosome 4).
[0047] Memory 150 may be tangible non - transitory memory and may include any type of suitable memory module, including random access memory (RAM), read - only memory (ROM), flash memory, other types of persistent memory, etc. Memory 150 may store, for example, instructions for an operating system (OS) 152 that can be executed on processor 142, and the operating system may be any type of suitable operating system, such as a modern smartphone operating system. Memory 150 may also store, for example, instructions for a spatial organization module 160 that can be executed on processor 142. A more detailed description of the chromatin interaction server 102 will be given below with reference to Figure 1B In some embodiments, the spatial organization module 160 may be part of one or more of the client devices 106 - 116, the chromatin interaction server 102, or a combination of the chromatin interaction server 102 and the client devices 106 - 116.
[0048] In any case, the spatial organization module 160 may obtain a set of read - pair and multiple sets of bins from the database 154 and / or the client devices 106 - 116. The spatial organization module 160 may then use each set of bins as axes to generate n × m a square genomic matrix to identify n × m bin pairs. Additionally, for each bin pair, the spatial organization module 160 may identify a subset of the read - pairs corresponding to the bin pair. Then, the spatial organization module 160 may identify the normalized interaction frequency of each bin pair by comparing the actual density of the read - pairs in the bin pair with the expected density based on a density function that varies with genomic distance for all read - pairs. Such a function may be corrected for the percentage of GC sequence in a particular bin sequence, the sequence coverage of a particular bin sequence in the Hi - C sequencing dataset, or other suitable factors for Hi - C normalization. Various statistical methods may be used to perform the comparison to generate, for example, a p - value, which can be compared with a confidence threshold to determine whether a particular bin pair has enriched contacts. The spatial organization module 160 may provide an indication of the bin pair and an indication of the corresponding normalized interaction frequency for display on the client devices 106 - 116. These indications may be displayed in numerical form or graphical form (such as in the form of a spatial interaction map), as described in more detail below with reference to FIG. 7.
[0049] The chromatin interaction server 102 can communicate with client devices 106-116 via network 130. The digital network 130 can be a private network, a secure public Internet, a virtual private network, and / or some other type of network, such as a dedicated access line, an ordinary conventional telephone line, a satellite link, a combination of these, etc. In the case where the digital network 130 includes the Internet, data communication can be carried out on the digital network 130 via Internet communication protocols.
[0050] Now turning to Figure 1B , the chromatin interaction server 102 can include a controller 224. The controller 224 can include a program memory 226, a microcontroller or microprocessor (MP) 228, a random access memory (RAM) 230, and / or input / output (I / O) circuitry 234, all of which can be interconnected via an address / data bus 232. In some embodiments, the controller 224 can also include a database 239, or otherwise be communicatively connected to the database or other data storage mechanisms (e.g., one or more hard disk drives, optical storage drives, solid state storage devices, etc.). The database 239 can contain data such as subject information, read pair data, bin data, spatial interaction mapping templates, web page templates, and / or web pages, as well as other data required for interacting with users via the network 130. The database 239 can contain data similar to the database 154 described above with reference to Figure 1A description.
[0051] It should be understood that although Figure 1B only one microprocessor 228 is depicted, the controller 224 can include multiple microprocessors 228. Similarly, the memory of the controller 224 can include multiple RAMs 230 and / or multiple program memories 226. Although Figure 1B the I / O circuitry 234 is described as a single block, the I / O circuitry 234 can include many different types of I / O circuitry. The controller 224 can implement one or more RAMs 230 and / or program memories 226 as, for example, semiconductor memories, magnetically readable memories, and / or optically readable memories.
[0052] As Figure 1BAs shown, the program memory 226 and / or the RAM 230 can store various applications for execution by the microprocessor 228. For example, the user interface application 236 can provide a user interface to the chromatin interaction server 102, which can, for example, allow a system administrator to configure, troubleshoot, or test various aspects of the server operation. The server application 238 can operate to receive a set of read pairs and multiple sets of bins, generate a square genomic matrix of bin pairs, identify the normalized interaction frequency for each bin pair, and provide an indication of the bin pairs and an indication of the normalized interaction frequency to the client devices 106-116 of healthcare professionals or researchers. The server application 238 can be a single module 238, such as the spatial organization module 160 or multiple modules 238A, 238B.
[0053] Although the server application 238 is depicted in Figure 1B as including two modules 238A and 238B, the server application 238 can include any number of modules that complete tasks related to the implementation of the chromatin interaction server 102. It should be understood that although only one chromatin interaction server 102 is depicted in Figure 1B , multiple chromatin interaction servers 102 can be provided for distributing server loads, serving different web pages, etc. These multiple chromatin interaction servers 102 can include web servers, entity-specific servers (such as Apple® servers, etc.), servers located in retail or private networks, etc.
[0054] Now referring to Figure 1C, the laptop computer 114 (or any one of the client devices 106-116) may include a display 240, a communication unit 258, a user input device (not shown), and a controller 242 similar to the chromatin interaction server 102. Similar to the controller 224, the controller 242 may include a program memory 246, a microcontroller or microprocessor (MP) 248, a random access memory (RAM) 250, and / or input / output (I / O) circuitry 254, all of which may be interconnected via an address / data bus 252. The program memory 246 may include an operating system 260, a data storage device 262, a plurality of software applications 264, and / or a plurality of software routines 268. For example, the operating system 260 may include Microsoft Windows®, OS X®, Linux®, Unix®, etc. The data storage device 262 may include data such as subject information, application data of the plurality of applications 264, routine data of the plurality of routines 268, and / or other data necessary for interacting with the chromatin interaction server 102 via the digital network 130. In some embodiments, the controller 242 may also include other data storage mechanisms (e.g., one or more hard disk drives, optical storage drives, solid state storage devices, etc.) residing within the laptop computer 114, or otherwise communicatively connected to the other data storage mechanisms.
[0055] The communication unit 258 may communicate with the chromatin interaction server 102 via any suitable wireless communication protocol network such as a wireless telephone network (e.g., GSM, CDMA, LTE, etc.), a Wi-Fi network (802.11 standard), a WiMAX network, a Bluetooth network, etc. The user input device (not shown) may include a "soft" keyboard displayed on the display 240 of the laptop computer 114, an external hardware keyboard (e.g., a Bluetooth keyboard) communicating via a wired or wireless connection, an external mouse, a microphone for receiving voice input, or any other suitable user input device. As discussed with reference to the controller 224, it should be understood that although Figure 1C only one microprocessor 248 is depicted, the controller 242 may include multiple microprocessors 248. Similarly, the memory of the controller 242 may include multiple RAMs 250 and / or multiple program memories 246. Although Figure 1C the I / O circuitry 254 is described as a single block, the I / O circuitry 254 may include many different types of I / O circuitry. The controller 242 may implement one or more RAMs 250 and / or program memories 246 as, for example, semiconductor memories, magnetically readable memories, and / or optically readable memories.
[0056] In addition to other software applications, one or more processors 248 may be adapted and configured to execute any one or more of a plurality of software applications 264 and / or any one or more of a plurality of software routines 268 residing in program memory 246. One application of the plurality of applications 264 may be a client application 266, which may be implemented as a series of machine-readable instructions for performing various tasks associated with receiving information at the laptop computer 114, displaying information on the laptop computer, and / or sending information from the laptop computer.
[0057] One application of the plurality of applications 264 may be a native application and / or a web browser 270 (such as Apple's Safari®, Google Chrome™, Microsoft Internet Explorer®, and Mozilla Firefox®), which may be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the chromatin interaction server 102 while also receiving input from a user such as a healthcare professional or researcher. Another application of the plurality of applications may include an embedded web browser 276, which may be implemented as a series of machine-readable instructions for receiving, interpreting, and / or displaying web page information from the chromatin interaction server 102.
[0058] One of the plurality of routines may include a spatial organization display routine 272 that obtains an indication of bin pairs and an indication of the normalized interaction frequency and presents a spatial interaction map on the display 240. Another routine of the plurality of routines may include a data input routine 274 that obtains a set of read pairs, a selection of a set of bins or two sets of bins to include as axes in a square genomic matrix, and sends the set of read pairs, the selection of the set of bins or two sets of bins to the chromatin interaction server 102.
[0059] Preferably, a user may initiate the client application 266 from a client device (such as one of the client devices 106 - 116) to communicate with the chromatin interaction server 102, thereby implementing the chromatin interaction system 100. Additionally, the user may also initiate or instantiate any other suitable user interface application (e.g., a native application or web browser 270, or any other application of the plurality of software applications 264) to access the chromatin interaction server 102, thereby implementing the chromatin interaction system 100.
[0060] As described above, Figure 1AThe chromatin interaction estimation server 102 shown can include a memory 150 that can store instructions for the spatial organization module 160 executable on a processor 142.
[0061] Figure 2 A set of example bins 200 is shown, where each bin represents a contiguous segment of a chromosome illustratively labeled as chromosome A. In this example, each bin contains several loci within chromosome A. The first bin is from locus 1 to locus 100, the second bin is from locus 100 to locus 186, the third bin is from locus 192 to locus 304, the fourth bin is from locus 308 to locus 396, the fifth bin is from locus 396 to locus 472, the sixth bin is from locus 478 to locus 672, the seventh bin is from locus 672 to locus 716, and the eighth bin is from locus 716 to locus 904. Each bin can represent different cleavage site increments or functional elements within chromosome A, such as genes, TADs, chromatin state segments, loop domains, chromatin domains, etc. In the example set, the bins are non - overlapping and the size of each bin (or the length of the chromosomal segment of each bin) varies. As described above, the bins can be selected by a healthcare professional or researcher, can be determined from previous studies, can be pre - stored bins in the database 154, or can be selected in any suitable manner. Although the set of example bins 200 is within chromosome A, the bins can be generated across the entire genome or any suitable portion thereof and can be of any suitable size.
[0062] As described above, the chromatin interaction server 102, more specifically, the spatial organization module 160 can obtain two sets of bins similar to the set of bins 200 and can generate a square genomic matrix, where each set of bins is an axis of the matrix. In some embodiments, the two sets of bins are the same and correspond to the same chromosome or other genomic region. In other embodiments, the two sets of bins correspond to the same chromosome or genomic region, but the bins are different, i.e., the chromosome or genomic region of each axis is divided differently. In still other embodiments, the two sets of bins correspond to different chromosomes or other genomic regions. In any case, a healthcare professional or researcher can select the set of bins to be used as the axes in the matrix via the client devices 106 - 116, or the set of bins can be selected in any suitable manner.
[0063] Although the above has described a set of bins with reference to chromosomes (e.g., a set of bins corresponding to chromosome A), this is merely an example for ease of illustration. A set of bins can correspond to any suitable set of DNA sequence fragments in the genome of a human or other organism, such as a genome-wide collection of TADs, a genome-wide collection of genes, a genome-wide collection of chromatin state fragments, a collection of loci of interest in a particular biomedical context, etc. In addition to genome-wide collections, a set of bins can be allele-specific and can correspond to a particular haplotype and / or diplotype. Furthermore, multiple sets of bins can be generated depending on the ploidy level and / or copy number. More generally, a set of bins can contain any collection of bins, where each of the bins corresponds to the same type of functional element (e.g., a set of TADs, genes, chromatin state fragments, loci, loop domains, chromatin domains, etc.). For example, a set of bins can be selected for a genome-wide search for long-range interactions, a focused search for interaction partners of a particular locus or set of loci, a genome-wide mapping of regulatory circuits, a comprehensive assessment of cell-type variability in long-range interactions, Hi-C-based diagnostic and prognostic biomarkers, etc. However, a set of bins does not necessarily have to correspond to the same type of functional element and can contain any suitable collection of bins.
[0064] Figure 3 Shown is a bin pair 300 of an example square genomic matrix that can be generated by the chromatin interaction server 102, more specifically, the spatial organization module 160. In this example, a set of bins from chromosome A can correspond to one axis of the matrix, and a set of bins from chromosome B can correspond to the other axis. A rectangular region of the matrix corresponding to a bin from one axis and a bin from the other axis can be referred to as a bin pair. For example, bin pair 302 corresponds to ChrA: 478-672*ChrB: 320-488. As shown in the example square genomic matrix, bin pairs have different shapes and sizes. Some bin pairs are rectangular, while others are more square-like (e.g., bin pair 302). Additionally, in the matrix, the rectangular regions of the bin pairs are different.
[0065] In addition to generating the matrix, chromatin interaction server 102 also identifies read pairs within each bin pair. A set of read pairs can be obtained from database 154, provided by a researcher or healthcare professional via client device 106 - 116, or obtained in any suitable manner. In any case, when both reads are within the rectangular region occupied by the bin pair, the read pair can be identified as being within the bin pair. For example, the rectangular region occupied by the bin pair containing read pair 304 spans from ChrA: 478 - 672 * ChrB: 1 - 320. This means that any read pair having a locus on chromosome A between 478 and 672 and a locus on chromosome B between 1 and 320 is within the bin pair. Read pair 304 can contain the loci ChrA: 570, ChrB: 160, which is within the rectangular region of ChrA: 478 - 672 * ChrB: 1 - 320. In some embodiments, a binary search tree, another type of search tree such as a quadtree, k-d tree, or B-tree, or any other suitable data structure for efficient searching such as a hash table can be used to match read pairs with bin pairs.
[0066] Chromatin interaction server 102 can then identify a subset of read pairs corresponding to each bin pair. For each bin pair, the corresponding subset of read pairs can be used to determine the actual density or interaction frequency of the read pairs of the bin pair. In some embodiments, the actual density of the read pairs of a bin pair can be the number of read pairs within the bin pair or the number of read pairs divided by the rectangular region occupied by the bin pair. In any case, the interaction frequency of each bin pair can be normalized according to a density function.
[0067] In some embodiments, chromatin interaction server 102, and more specifically, spatial organization module 160, can provide an indication of the bin pair and an indication of the normalized interaction frequency to the client device 106 of a researcher or healthcare professional. Client device 106 can display a graphical representation of the bin pair and the normalized interaction frequency. Figure 4 An example spatial interaction map 400 or heat map of bin pairs and the corresponding normalized interaction frequencies that can be presented on client device 106 is shown. Figure 4 The set of bins shown represents the human genome (chromosomes 1 - 22, X). In example spatial interaction map 400, the normalized interaction frequencies of the bin pairs are represented by a color scale to produce a two-dimensional map of chromatin organization. More specifically, bin pairs having a larger normalized interaction frequency are highlighted in a darker color. Spatial interaction map 400 can be similar to the square genomic matrix described graphically above with reference to Figure 3 in a graphical form. As Figure 4As shown, the bin pairs along the diagonal from the upper left to the lower right on the spatial interaction map 400 are presented in a darker color than other bin pairs in the spatial interaction map 400. Thus, these bin pairs may have enriched contacts. Additionally, the contacts in the squares along the axis representing intra-chromosomal contacts are darker than the inter-chromosomal contacts in the off-axis regions of the square genomic space. Generally speaking, the "cis" regions within a chromosome have a higher contact frequency than the "trans" regions between chromosomes. In corresponding embodiments, the sets of regions with enriched and depleted contacts and the degree of contact can be evaluated for many useful purposes.
[0068] The spatial interaction map 400 and / or other representations of such contacts can be used to generate 3D and 4D chromatin structures, such as the 3D chromatin structure 410. The 3D chromatin structure 410 depicts the chromosomes located in the chromatin-binding regions of the nucleus as 4D nucleosomes. Euchromatin is characterized by DNase 1 hypersensitivity and a specific combination of histone marks that define active genomic regulatory elements. For example, promoters typically carry the marks H3K4me3 and H3K27ac, and enhancers typically carry the marks H3K4me1 and H3K27ac. Enhancers can increase or decrease transcription in their target genes, which can be sequence-proximal, and / or spatially localized (e.g., by the methods described above) and / or functionally linked, either alone or in combination (e.g., by molecular QTL linkage), to the enhancer. Heterochromatin is located in the interior of chromosomal regions and at the periphery of the nucleus, near the nuclear lamina and nucleolus, and is characterized by its own repressive chromatin marks and DNA-binding proteins, as well as spatial compaction and linker histones. Recent studies have shown that in the brain, the DNA sequence CAC is a common site of methylation, which is contrary to other tissues where CpG is most frequently methylated. Additionally, in the brain, a reactive species of a unique element carrying epigenetic information - 5-hydroxymethylcytosine (5hmC) - is relatively common. In contrast, in the periphery, methylcytosine (hmC) is common.
[0069] As described above, to determine the normalized interaction frequency, the spatial organization module 160 can apply a density function to the bin pairs to calculate the expected density of each bin pair. Figure 5Shows an example curve graph 500, which depicts three example density functions 510, 520, 530, each density function representing the change of the true or ideal density of read pairs across all read pairs in this group with genomic distance. As shown in the example curve graph 500, each of the density functions 510 - 530 changes in an overall downward-sloping manner with genomic distance because a large number of read pairs contain reads that are very close to each other. The density function 510 decreases monotonically, while the density functions 520 and 530 first increase with distance and then decrease, and have variable slopes and levels at each distance position. The density function 510 is a synthetic function based on a published power-law spline model, which is the type of synthetic distance density curve sometimes used to normalize Hi-C datasets, while the other two density functions 520, 530 are empirical density functions. The density function 520 is based on a dataset from SK-N-SH cells (neuronal cells), and the density function 530 is based on a dataset from fibroblasts (skin cells). However, these are just a few examples of the density functions of several groups of read pairs. The change of density with genomic distance can exhibit other patterns for other groups of read pairs (e.g., this function can decay at a faster or slower rate with the increase of genomic distance). Such functions for a specific database of sequence contacts can be generated empirically and appropriately adjusted.
[0070] In any case, the chromatin interaction server, more specifically, the spatial organization module 160 can apply one of the density functions 510 - 530 shown in the example curve graph 500 to calculate the expected density of each bin pair. In some embodiments, the spatial organization module 160 can select an empirical density function applicable to the selected bin pair. For example, when the bin group contains bins representing DNA sequence fragments expressed in skin cells, the spatial organization module 160 can select the density function 530 based on the dataset from fibroblasts. When the bin group contains bins representing DNA sequence fragments expressed in neurons, the spatial organization module 160 can select the density function 520 based on the dataset from SK-N-SH cells. In other embodiments, the spatial organization module 160 can select the synthetic density function 510.
[0071] For a particular pair of bins, the spatial organization module 160 can integrate a selected density function (e.g., density function 520) across the rectangular region occupied by the pair of bins to determine the expected density of the pair of bins. Then, for each pair of bins, the spatial organization module 160 can compare the expected density of the pair of bins with the actual density using various statistical methods to determine whether the expected density differs from the actual density by a statistically significant amount. For example, the null hypothesis can be that the actual density of the pair of bins is not greater than the expected density. The spatial organization module 160 can compare the expected density with the actual density according to a Poisson distribution or any other suitable distribution using a one-tailed test to generate a p-value. When the p-value is less than a threshold confidence level (e.g., a p-value of.05 corresponds to 95% confidence, a p-value of.01 corresponds to 99% confidence, etc.), the null hypothesis can be rejected, and the spatial organization module 160 can determine that the pair of bins contains enriched contacts. In some embodiments, the spatial organization module 160 can apply a false discovery rate to the p-value, such as the Benjamini false discovery rate, or other statistical methods for multiple comparison control.
[0072] In another instance, the null hypothesis can be that the actual density of the pair of bins is not less than the expected density. The spatial organization module 160 can compare the expected density with the actual density according to a Poisson distribution or any other suitable distribution using a one-tailed test to generate a p-value. When the p-value is less than a threshold confidence level (e.g., a p-value of.05 corresponds to 95% confidence, a p-value of.01 corresponds to 99% confidence, etc.), the null hypothesis can be rejected, and the spatial organization module 160 can determine that the pair of bins contains depleted contacts. In some embodiments, the spatial organization module 160 can apply a false discovery rate to the p-value, such as the Benjamini false discovery rate, or other statistical methods for multiple comparison control.
[0073] In yet another instance, the null hypothesis can be that the actual density of the pair of bins is the same as the expected density. The spatial organization module 160 can compare the expected density with the actual density according to a Poisson distribution or any other suitable distribution using a two-tailed test to generate a p-value. When the p-value is less than a threshold confidence level (e.g., a p-value of.05 corresponds to 95% confidence, a p-value of.01 corresponds to 99% confidence, etc.), the null hypothesis can be rejected, and the spatial organization module 160 can determine that the pair of bins contains differential or abnormal contacts (i.e., enriched or depleted contacts). In some embodiments, the spatial organization module 160 can apply a false discovery rate to the p-value, such as the Benjamini false discovery rate, or other statistical methods for multiple comparison control.
[0074] Although the statistical analysis is described herein with reference to the Poisson distribution, this is merely one type of statistical test that can be used to determine whether there is a statistically significant difference between the actual density and the expected density of the bins. Other statistical tests can include t-tests, chi-square tests, G-tests, regression tests, etc. In addition, in addition to statistical tests, machine learning methods can also be used, including but not limited to regression algorithms (e.g., ordinary least squares regression, linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines, locally estimated scatterplot smoothing, etc.), instance-based algorithms (e.g., k-nearest neighbors, learning vector quantization, self-organizing maps, locally weighted learning, etc.), regularization algorithms (e.g., ridge regression, least absolute shrinkage and selection operator, elastic net, least angle regression, etc.), decision tree algorithms (e.g., classification and regression trees, C4.5, C5, chi-square automatic interaction detection, decision stumps, M5, conditional decision trees, etc.), clustering algorithms (e.g., k-means, k-medians, expectation maximization, hierarchical clustering, spectral clustering, mean shift, density-based spatial clustering of applications with noise, ordering points to identify the clustering structure, etc.), association rule learning algorithms (e.g., Apriori algorithm, Eclat algorithm, etc.), Bayesian algorithms (e.g., naive Bayesian, Gaussian naive Bayesian, multinomial naive Bayesian, averaged one-dependence estimators, Bayesian belief networks, Bayesian networks, etc.), artificial neural networks (e.g., perceptrons, Hopfield networks, radial basis function networks, etc.), deep learning algorithms (e.g., multi-layer perceptrons, deep Boltzmann machines, deep belief networks, convolutional neural networks, stacked autoencoders, generative adversarial networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, principal component regression, partial least squares regression, Sammon mapping, multidimensional scaling, projection pursuit, linear discriminant analysis, mixture discriminant analysis, quadratic discriminant analysis, flexible discriminant analysis, factor analysis, independent component analysis, non-negative matrix factorization, t-distributed stochastic neighbor embedding, etc.), ensemble algorithms (e.g., boosting, bagging, AdaBoost, stacking generalization, gradient boosting machines, gradient boosting regression trees, random decision forests, etc.), reinforcement learning (e.g., temporal difference learning, Q-learning, learning automata, state-action-reward-state-action, etc.), support vector machines, mixture models, evolutionary algorithms, probabilistic graphical models, etc.
[0075] In addition, although the method described herein uses the false discovery rate for multiple comparison control, any suitable multiple comparison control method (such as false coverage rate, Bayesian methods, etc.) can be applied to the p-values.
[0076] Then, bins pairs with enriched or depleted contacts can be used to predict a subject's molecular phenotype based on the spatial interactions of loci within the corresponding genome. Enriched or depleted contacts can also be used to model the 3D and 4D structures of chromosomes and identify altered TAD boundaries and spatial interactions in tissue samples to determine genetic diseases or oncology. Additionally, enriched or depleted contacts can be used to determine whether and / or to what extent a pair of loci interact in a particular tissue or cell line. Further still, enriched or depleted contacts can be used to identify ploidy and translocations based on abnormal contacts and / or the total contact density in the square genomic space.
[0077] For example, a healthcare professional can obtain a biological sample (e.g., from a cheek swab, skin sample, biopsy section, blood sample, lymph fluid, bone marrow, cell line, tissue, model organism, etc.) for measuring chromatin interaction data of a subject and provide the laboratory results obtained by analyzing the biological sample to a chromatin interaction server.
[0078] In Figure 6 An example process 600 for generating chromatin interaction data from a subject's biological sample is shown. The process can be performed by an analytical laboratory or other suitable institution. At block 602, a healthcare professional obtains a biological sample of the subject and sends it to a laboratory for analysis. The biological sample can include the subject's skin, blood, lymph fluid, bone marrow, cheek cells, cell line, tissue, etc. Then at block 604, cells are extracted from the biological sample and reprogrammed into stem cells (such as induced pluripotent stem cells (iPSCs)) at block 606. Then at block 608, the iPSCs are differentiated into multiple tissues (such as neurons, cardiomyocytes, etc.) and analyzed at block 610 to obtain chromatin interaction data. The chromatin interaction data can include 5C data, Hi-C data, ChIA-PET data, Combi-C data, genome architecture mapping data, Micro-C data, etc.
[0079] In addition, a single locus of a DNA sequence can be identified as being associated with or causally related to a specific molecular phenotype. A set of bins can also be identified as containing a single locus. Then, when chromatin interaction data of a subject is analyzed with respect to a specific molecular phenotype or a set of molecular phenotypes (e.g., a molecular phenotype indicative of a response to valproic acid), iPSCs of loci associated with or causally related to a specific set of molecular phenotypes can be analyzed. The set of bins corresponding to the loci identified from the assay can be compared with contact data of such a set of bins in other biological cell systems. For example, such a system can consist of two different tissues in a human body, tissue samples from two different individuals, a cell line that has been subjected to a medical treatment compared to a control sample, or multiple cell cycle conditions or cell differentiation states of the same tissue, cell line, or organism. Then, the chromatin interaction server 102 can predict the molecular phenotype of the subject based on this comparison. For example, if the iPSCs of a subject contain read pairs in a set of bins having loci associated with or causally related to a specific response to valproic acid, the chromatin interaction server 102 can predict that the subject will have a specific response to valproic acid.
[0080] More generally, which chromatin interaction data to select for the assay can be based on chromatin interaction data identified as being associated with or causally related to the set of molecular phenotypes being assayed in the subject.
[0081] More specifically, cells are reprogrammed into iPSCs by introducing transcription factors or "reprogramming factors" or other reagents into a given cell type. For example, cells can be reprogrammed into iPSCs using Yamanaka factors (comprising the transcription factors Oct4 (POU5F1), Sox2 (SOX2), cMyc (MYC), and Klf4 (KLF4)). The iPSCs can then be differentiated into various tissues such as neurons, adipocytes, cardiomyocytes, pancreatic beta cells, etc. After differentiating the iPSCs, various detection techniques (such as DNA methylation analysis, DNase footprinting assays, filter binding assays, etc.) can be used to detect the differentiated iPSCs to identify epigenomic information. In fact, the system performs a virtual biopsy, and the differentiated iPSCs have at least to some extent the phenotypic and epigenomic characteristics of their corresponding tissues.
[0082] In the above embodiments, cells are extracted from a biological sample of a subject, reprogrammed into stem cells, differentiated into various tissues, and analyzed to obtain chromatin interaction data (differentiated, reprogrammed cell assays). Alternatively, in some embodiments, a biological sample of a patient is assayed without extracting cells (cell-free assays). In other embodiments, cells are extracted from a biological sample of a patient and assayed without reprogramming or differentiating the cells (primary cell assays). In other embodiments, the cells are reprogrammed into iPSCs and analyzed without differentiating the cells (reprogrammed stem cell assays). For example, iPSCs can be analyzed without differentiation to obtain stem cell omics. Although these are just some example processes for generating chromatin interaction data from a biological sample of a subject, the assay can be performed at any suitable stage in the process, and the chromatin interaction data can be generated in any suitable manner.
[0083] In some embodiments, the spatial organization module 160 can then provide an indication of the bin pairs and an indication of the normalized interaction frequencies to a client device 106 of a researcher or healthcare professional. Figure 7A An example display 700 of enriched contacts and corresponding loci and SNPs associated with a molecular phenotype that can be presented on the client device 106 is shown. The example display 700 includes TADs (each TAD consisting of a set of individual loci of a DNA sequence) that contain a set of regulatory SNPs that are considered to be significantly responsive to the phenotypes of two drugs (in particular valproic acid and ketamine) that induce neurogenesis in adults. As an exemplary instance of an embodiment of the present invention, scientists studying these drugs may wish to discern the mechanisms by which they act. These scientists may wish to identify the target genes of these regulatory SNPs and thus find the spatial contact partners of the TADs containing them in the system of interest. These scientists may also wish to find out which of these TADs contact each other in the system of interest. In other embodiments, the set of loci of interest can be identified as corresponding to a set of variants of a drug, a disease, or other molecular phenotype, or a set of loci of interest of a particular subject, etc.
[0084] The example display 700 includes a chromosome (e.g., chromosome 17), a locus (e.g., 33720000 - 35360000), a TAD (e.g., 1977), a candidate contact (e.g., 1), and a target gene associated with a bin pair having an enriched contact (e.g., CCL2). Each TAD in display 700 shows some distal contacts across the genome, with the number ranging from three (e.g., TAD 1977) to hundreds (e.g., TAD 2112). TADs containing pharmacokinetic loci, such as CYP genes that metabolize these drugs, seem to have the most contacts. Display 700 also includes SNPs associated with enriched contacts (e.g., rs2857654) and drugs whose response is associated with the SNP (e.g., valproic acid). In this way, healthcare professionals or researchers can view SNPs associated with enriched contacts and their corresponding molecular phenotypes. However, this is merely an example of a digital indication of bin pairs shown for illustrative purposes only. In other embodiments, client device 106 can display a digital indication of each bin pair (e.g., Chr11:8560000-10720000*Chr11:4580000-4780000), a digital indication of the interaction frequency of the bin pairs (such as a p-value), an indication of whether the bin pair has enriched contacts, etc.
[0085] In still other embodiments, client device 106 can display a chromatin interaction network generated from spatial contact data. Figure 7B Shows an example display of chromatin interaction network 750 generated from spatial contact data according to the reference Figure 7A description. In some embodiments, chromatin interaction network 750 can be presented on client device 106. In any case, as shown in display 700, TADs with at least one candidate contact are included in chromatin interaction network 750. For example, TAD 1977 is a candidate contact and is thus included in chromatin interaction network 750. Additionally, TAD1977 is associated with its candidate contact (TAD 2112). Further, TAD 2245 is connected to TADs 2258 and 2112, and TADs 1567, 1862, and 693 have zero candidate contacts and are thus not included in chromatin interaction network 750. Thus, in Figure 7AOf the 13 identified TADs shown in the display, 10 contain connections to other TADs in groups with dense interaction spatial networks of TADs related to formation function. In any case, TAD 1636 is connected to TADs 2258 and 2112, and TAD 1832 is connected to TADs 2258, 2112, and 1418. Additionally, TAD 2112 is connected to every other TAD in chromatin interaction network 750, and TAD 2258 is connected to every other TAD in chromatin interaction network 750 except TAD 1977. Additionally, TAD 1063 is connected to its four candidate contacts (TADs 2112, 2258, 1418, and 1343). TAD 1343 also has four candidate contacts and is connected to each of these contacts (TADs 2245, 1063, 2258, 1343, and 1418) in chromatin interaction network 750. TAD 1418 is connected to TADs 1832, 2112, 2258, 1063, and 1343 in chromatin interaction network 750, and TAD 1472 is connected to TADs 2258 and 2112. None of the TADs contained in chromatin interaction network 750 are recognized by an alternative system such as HOMER, described in more detail below. TADs 1064, 2258, 2112, and 1418 are pharmacokinetic TADs, while TADs 1472, 2245, 1636, 1977, 1343, and 1832 are pharmacodynamic TADs.
[0086] In this way, a healthcare professional or researcher viewing chromatin interaction network 750 on client device 106 can see the strength of the relationships within chromatin interaction network 750. For example, a healthcare professional or researcher might see that TAD 2112 has a relationship with all other TADs in chromatin interaction network 750, while TAD 1977 is part of chromatin interaction network 750 but has a relationship with only one TAD. Given the different groups of genes and variants present in each TAD, and their different biological functions and significance in various biomedical and research contacts, the accurate detection and display of chromatin contacts can serve many useful purposes in various embodiments.
[0087] Figure 9 Shows the comparison results between the chromatin interaction system and HOMER, a widely used Hi-C compiler. The chromatin interaction system and HOMER analyzed a dataset of human fibroblasts, where the chromatin interaction system produced a matrix of TAD bins on two axes. For HOMER, fixed bins with a resolution of 1Mb were used. As Figure 9As shown, HOMER detected 18,220 contacts, while the chromatin interaction system detected 17,720 contacts, and the two systems detected the same 5,648 contacts. The chromatin interaction system detected 10,193 contacts that were not detected by HOMER. Therefore, due to the improved discrimination ability of the chromatin interaction system, the chromatin interaction system identified contacts that were not detected by the previous system. Therefore, the chromatin interaction system was able to identify more long-range contacts than the previous system (e.g., cis interactions in the range greater than 10 Mb, which were enriched two-fold or more and passed the comparison control, square genome-wide).
[0088] HOMER detected 12,572 contacts that were not detected by the chromatin interaction system. However, among these contacts, 82% did not pass the fold change threshold in the chromatin interaction system, 90% did not pass the FDR threshold, and 72% did not pass both thresholds. Among these pairs, 92% of the neighboring TAD pairs did detect contacts in the chromatin interaction system. The non-neighboring discordant HOMER contacts included 1,054 contacts.
[0089] Figure 8 A flowchart depicting an exemplary method 800 for analyzing the spatial organization of chromatin is shown. Method 800 can be executed on the chromatin interaction server 102. In some embodiments, method 800 can be implemented in a set of instructions stored on a non-transitory computer-readable memory and executable on one or more processors on the chromatin interaction server 102. For example, method 800 can be executed by Figure 1A the spatial organization module 160.
[0090] The method may include steps of alignment, quality control, compilation, integration, statistical testing, and result output. More specifically, at block 802, the spatial organization module 160 can obtain a set of paired genomic element contacts or read pairs. The set of read pairs can be obtained from, for example, Figure 1AThe illustrated database 154 is obtained from the client device 106 of a researcher or professional, or in any other suitable manner. In some embodiments, a researcher or professional may select a particular set of read pairs for a particular analysis or study and provide the selected set to the chromatin interaction server 102 via the client device 106. In another example, a sequencer 107 may generate sequence data that is provided to the chromatin interaction server 102. The spatial organization module 160 may then identify a set of read pairs from the sequence data. In yet another example, a sequence database 109 may provide pre-existing sequence data generated from, for example, published literature, clinical trials, consortia, academia, etc. to the chromatin interaction server 102. The spatial organization module 160 may then identify a set of read pairs from the sequence data. In some embodiments, the spatial organization module 160 divides the read pairs into their single-end components and aligns the single-end components. A plurality of pairs are then selected, where there is a less than threshold probability (e.g., 0.05) of misalignment in either read for the two reads in the pair.
[0091] The spatial organization module 160 may also divide genomic element contacts or reads into bins (block 804). Each bin may represent a different cleavage site increment or functional element in the genome or a portion of the genome, such as a gene, TAD, chromatin state segment, loop domain, chromatin domain, etc. The bins are non-overlapping and the size of the bins is inconsistent, i.e., the size of each bin (or the length of the genomic segment for each bin) varies. In some embodiments, the bins may be selected by a healthcare professional or researcher via the client device 106, may be determined from a previous study such as the sequence database 109, may be pre-stored bins in the database 154, or may be selected in any suitable manner. For example, a researcher or healthcare professional may select a set of bins where each bin represents a different TAD. In another example, a researcher or healthcare professional may select a set of bins where each bin represents a different gene.
[0092] Then at block 806, a first set n of bins and a second set m of bins are selected, where each set corresponds to an axis of an n × m square genomic matrix. The axes may be selected by a healthcare professional or researcher via the client device 106 - 116, or may be selected in any suitable manner. In some embodiments, the two sets of bins are the same and correspond to the same chromosome. In other embodiments, the two sets of bins correspond to the same chromosome, but the bins are different, i.e., the chromosome for each axis is divided differently. In yet other embodiments, the two sets of bins correspond to different chromosomes. In any case, the spatial organization module 160 may generate a containment n × mA square genomic matrix of bin pairs (box 810), where a bin pair is an entry or rectangle in the square genomic matrix (e.g., Chr1:1000 - 2000 * Chr8:10000 - 20000).
[0093] Then, the spatial organization module 160 compiles read pairs into bin pairs. More specifically, the spatial organization module 160 can use, for example, a binary search tree to identify subsets of read pairs corresponding to each bin pair (box 810). A read pair can be identified as being within a bin pair when both reads are within the rectangular region occupied by the bin pair. For example, as Figure 3 shown, the rectangular region occupied by the bin pair containing read pair 304 spans from ChrA: 478 - 672 * ChrB: 1 - 320. This means that any read pair having a locus on chromosome A between 478 and 672 and a locus on chromosome B between 1 and 320 is within the bin pair. Read pair 304 can contain the loci ChrA: 570, ChrB: 160, which is within the rectangular region of ChrA: 478 - 672 * ChrB: 1 - 320. In some embodiments, another type of search tree (such as a quadtree, k - d tree, or B - tree) or any other suitable data structure for efficient search (such as a hash table) can be used to match read pairs with bin pairs.
[0094] At box 812, the spatial organization module 160 generates a density function based on the variation of read - pair density with genomic distance across the entire set of read pairs. In some embodiments, the density function can be a monotonically decreasing function. For a particular bin pair, the density function is integrated over the rectangular region of the bin pair (e.g., ChrA:478 - 672 * ChrB:1 - 320) to determine the expected density of this bin pair (box 814). The expected density can be determined for each bin pair in the set of bin pairs.
[0095] Then, the spatial organization module 160 can compare the expected density of a particular bin pair with the actual density of the particular bin pair. For example, the actual density can be the number of read pairs contained in the particular bin pair. Statistical analysis can be used to compare the actual density and the expected density to determine whether the difference between the expected densities differs by a statistically significant amount (the normalized interaction frequency) from the actual density (block 816). For example, the null hypothesis can be that the actual density of the bin pair is not greater than the expected density. The spatial organization module 160 can compare the expected density with the actual density according to a Poisson distribution or any other suitable distribution to generate a p-value. When the p-value is less than a threshold confidence level (e.g., a p-value of.05 corresponds to 95% confidence, a p-value of.01 corresponds to 99% confidence, etc.), the null hypothesis can be rejected, and the spatial organization module 160 can determine that the bin pair contains enriched contacts. In some embodiments, the spatial organization module 160 can apply a false discovery rate to the p-value, such as the Benjamini false discovery rate, or other statistical methods for multiple comparison control. In another example, the null hypothesis can be that the actual density of the bin pair is not less than the expected density. When the p-value is less than the threshold confidence level, the null hypothesis can be rejected, and the spatial organization module 160 can determine that the bin pair contains depleted contacts.
[0096] In some embodiments, the actual read counts of multiple sets of contacts corresponding to, for example, different biological cell systems or different physiological conditions can be analyzed together. For example, such systems can consist of two different tissues in the human body, tissue samples from two different individuals, a cell line that has been subjected to a medical treatment compared to a control sample, or multiple cell cycle conditions or cell differentiation states of the same tissue, cell line, or organism. This analysis can determine, for example, a set of differential contacts between a pair of sets of contacts by, for example, separately comparing the enriched and depleted contacts from each data set. Differential contacts can also be determined by, for example, using a multiple sampling distribution of a Poisson distribution or other statistical distribution to generate a p-value corresponding to the probability of an accidentally observed differential interaction frequency, which can then be corrected with a false discovery rate or other methods as described herein.
[0097] At block 818, the spatial organization module 160 can provide an indication of the bin pair and an indication of the normalized interaction frequency to the client device 106 of a researcher or healthcare professional. These indications can include a numerical indication of the normalized interaction frequency (e.g., the p-value), a graphical representation of the bin pair and the normalized interaction frequency (e.g., a spatial organization map), a list of bin pairs with enriched contacts, or any other suitable indication.
[0098] Throughout the specification, multiple instances may implement components, operations, or structures described as a single instance. Although the separate operations of one or more methods are shown and described as separate operations, one or more of the separate operations may be performed concurrently, and the operations need not be performed in the order shown. Structures and functions presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0099] Additionally, certain embodiments herein are described as including logic or a plurality of routines, subroutines, applications, or instructions. These may constitute software (e.g., code embodied on a machine-readable medium or in a transmitted signal) or hardware. In hardware, routines, etc. are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In an example embodiment, one or more computer systems (e.g., stand-alone client or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or a portion of an application) as hardware modules to perform certain operations described herein.
[0100] In various embodiments, the hardware modules may be implemented mechanically or electronically. For example, a hardware module may include dedicated circuitry or logic permanently configured to perform certain operations (e.g., a dedicated processor such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or custom silicon). A hardware module may also include programmable logic or circuitry temporarily configured by software to perform certain operations (e.g., as included in a general-purpose processor, a graphics processing unit (GPU), or other programmable processor). It should be understood that, driven by cost and time considerations, it may be determined to implement the hardware module mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software).
[0101] Thus, the term "hardware module" should be understood to encompass a tangible entity, referring to an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which the hardware module is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate every hardware module at any given moment. For example, in the case where a hardware module includes a general-purpose processor configured by software, the general-purpose processor may be configured into corresponding different hardware modules at different times. Thus, software can configure the processor, for example, to constitute a particular hardware module at one moment and different hardware modules at different moments.
[0102] Hardware modules can provide information to other hardware modules or receive information from other hardware modules. Thus, the hardware modules can be considered to be communicatively coupled. In cases where multiple such hardware modules are present simultaneously, communication can be achieved through signal transmission that connects the hardware modules (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, by storing and retrieving information in a memory structure accessible to the multiple hardware modules. For example, one hardware module can perform an operation and store the output of such operation in a memory device to which it is communicatively coupled. Then, another hardware module can access this memory device at a later time to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and can operate on resources (e.g., a collection of information).
[0103] The various operations of the example methods described herein can be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented modules that operate to perform one or more operations or functions. In some example embodiments, the modules referred to herein can include processor-implemented modules.
[0104] Similarly, the methods or routines described herein can be implemented, at least in part, by processors. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented hardware modules. The performance of certain operations can be distributed among one or more processors, which can reside not only within a single machine but also can be deployed across multiple machines. In some example embodiments, the processor or processors can be located in a single location (e.g., within a home environment, an office environment, or as a server farm), while in other embodiments, the processors can be distributed across multiple locations.
[0105] The performance of certain operations can be distributed among one or more processors, which can reside not only within a single machine but also can be deployed across multiple machines. In some exemplary embodiments, one or more processors or processor-implemented modules can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, one or more processors or processor-implemented modules can be distributed across multiple geographical locations.
[0106] Unless otherwise explicitly stated, discussions herein using terms such as "processing", "computing", "caculating", "determining", "presenting", "displaying", etc. can refer to actions or processes of a machine (e.g., a computer) to manipulate or transform data represented as physical (e.g., electrical, magnetic, or optical) quantities in one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0107] As used herein, any reference to "an embodiment" or "one embodiment" means that the particular element, feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. The phrase "in one embodiment" appearing in various places in the specification does not necessarily all refer to the same embodiment.
[0108] For example, some embodiments may use the term "coupled" to describe two or more elements that are in direct physical contact or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. Embodiments are not limited to these scopes.
[0109] As used herein, the terms "comprises / comprising", "includes / including", "has / having", or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or rather than an exclusive or. For example, any of the following satisfies the condition A or B: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0110] In addition, "a / an" is used to describe elements and components of embodiments herein. This is merely for convenience and to provide a general description. This description and the appended claims should be understood to include one or at least one, and the singular also includes the plural unless clearly indicated otherwise.
[0111] This detailed description should be construed as merely providing examples and not as describing every possible embodiment, as it would be impractical, if not impossible, to describe every possible embodiment. Many alternative embodiments can be implemented using current technology or technology developed after the filing date of this application.
[0112] The following list of aspects reflects various embodiments explicitly contemplated by this application. Those of ordinary skill in the art will readily understand that the aspects below are neither a limitation on the embodiments disclosed herein nor an exhaustive list of all embodiments that can be conceived from the foregoing disclosure, but are intended to be exemplary in nature.
[0113] 1. A computer-implemented method for analyzing the spatial and temporal organization of chromatin, the method being executed by one or more processors programmed to perform the method, the method comprising: obtaining, at one or more processors, a set of paired contacts of genomic elements; partitioning, by the one or more processors, the genomic elements into a plurality of bins, wherein the bin sizes of the plurality of bins are not consistent; identifying, by the one or more processors, a first set of the plurality of bins and a second set of the plurality of bins; generating, by the one or more processors n × m a matrix of bin pairs, wherein n corresponds to the first set of the plurality of bins, and m corresponds to the second set of the plurality of bins; identifying, by the one or more processors, a subset of the paired contacts within each bin pair of the bin pairs; determining, by the one or more processors, the interaction frequency of each bin pair of the bin pairs; normalizing, by the one or more processors, each interaction frequency of the interaction frequencies to generate a normalized interaction frequency for each bin pair; and providing, by the one or more processors, a map of chromatin interactions for display on a user interface, comprising an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies.
[0114] 2. The method according to aspect 1, wherein normalizing each of the interaction frequencies comprises: determining, by the one or more processors, a variation of the density of the set of paired contacts with genomic distance to generate a density function; for each of the plurality of bin pairs: integrating, by the one or more processors, the density function over the region of the bin pair to determine the expected density of the bin pair; comparing, by the one or more processors, the subset of paired contacts within the bin pair with the expected density of the bin pair by performing a statistical analysis using a Poisson statistical distribution to determine the likelihood that the actual density of the bin pair is significantly greater than the expected density of the bin pair; applying, by the one or more processors, a false discovery rate for multiple comparison control to the determined likelihood to determine an adjusted likelihood; and when the adjusted likelihood is less than a threshold likelihood, determining, by the one or more processors, that the bin pair has enriched contacts.
[0115] 3. The method according to any one of aspect 1 or aspect 2, further comprising: determining, by the one or more processors, a second likelihood that is significantly significant of the amount by which the actual density of the bin pair is less than the expected density of the bin pair by performing a statistical analysis using a Poisson distribution; applying, by the one or more processors, a false discovery rate for multiple comparison control to the determined second likelihood to determine an adjusted second likelihood; and when the adjusted second likelihood is less than a threshold likelihood, determining, by the one or more processors, that the bin pair has depleted contacts.
[0116] 4. The method according to any one of the preceding aspects, wherein the statistical analysis comprises a two-tailed test for determining a third likelihood that is statistically significant of the amount by which the actual density of the bin pair differs from the expected density; applying, by the one or more processors, a false discovery rate for multiple comparison control to the determined third likelihood to determine an adjusted third likelihood; and when the adjusted third likelihood is less than a threshold likelihood, determining, by the one or more processors, that the bin pair has enriched or depleted contacts.
[0117] 5. The method according to any one of the preceding aspects, wherein at least some of the paired contacts are cis contacts such that the two genomic elements in each of the at least some paired contacts correspond to the same chromosome; and wherein at least some of the paired contacts are trans contacts such that the two genomic elements in each of the at least some paired contacts correspond to different chromosomes.
[0118] 6. The method according to any one of the foregoing aspects, wherein the density function is generated from empirical data, and at least a portion of the density function decreases as the genomic distance increases.
[0119] 7. The method according to any one of the foregoing aspects, further comprising: identifying, by the one or more processors, a single locus in the DNA sequence that is associated with or causally related to one or more molecular phenotypes; identifying, by the one or more processors, a set of bins containing the single locus; obtaining, by the one or more processors, chromatin interaction data of a subject; comparing, by the one or more processors, the chromatin interaction data of the bins containing the single locus with the contact data of such bins in another biological cell system; and predicting, by the one or more processors, the molecular phenotype of the subject based on the comparison.
[0120] 8. The method according to any one of the foregoing aspects, further comprising: generating, by the one or more processors, a 3D or 4D model of the chromosomal structure based on the mapping of chromatin interactions.
[0121] 9. The method according to any one of the foregoing aspects, further comprising: generating, by the one or more processors, a spatial interaction network of a set of specific loci.
[0122] 10. The method according to any one of the foregoing aspects, wherein identifying the subset of paired contacts within each bin pair comprises identifying the subset of paired contacts within each bin pair using a binary search tree.
[0123] 11. The method according to any one of the foregoing aspects, wherein the first set of the plurality of bins and the second set of the plurality of bins are the same bins corresponding to the same chromosome.
[0124] 12. The method according to any one of the foregoing aspects, wherein each genomic element corresponds to a locus within the genome; and wherein each bin corresponds to a continuous segment of a deoxyribonucleic acid (DNA) sequence comprising at least one of the following: topologically associating domain (TAD), gene, chromatin state segment, loop domain, or chromatin domain.
[0125] 13. The method according to any one of the foregoing aspects, wherein identifying the first set of the plurality of bins and the second set of the plurality of bins comprises receiving a selection of the first set of bins and the second set of bins for a genome-wide search for long-range interactions, a genome-wide mapping of regulatory circuits, a comprehensive assessment of intercellular type variability in long-range interactions, or identifying a set of Hi-C-based diagnostic and prognostic biomarkers.
[0126] 14. The method according to any of the foregoing aspects, further comprising: for one or more of the pairs of bins, comparing, by the one or more processors, the actual density of the pair of bins from the first biological cell system or physiological condition with the actual density of the pair of bins from the second biological cell system or physiological condition to identify differential contacts.
[0127] 15. A computing device for analyzing the spatial and temporal organization of chromatin, the computing device comprising: a communication network, one or more processors; and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, the instructions, when executed by the one or more processors, cause the computing device to: obtain a set of paired contacts of genomic elements; partition the genomic elements into a plurality of bins, wherein the bin sizes of the plurality of bins are inconsistent; identify a first set of the plurality of bins and a second set of the plurality of bins; generate n × m a matrix of pairs of bins, wherein n corresponds to the first set of the plurality of bins, and m corresponds to the second set of the plurality of bins; identify a subset of paired contacts within each bin pair of the bin pairs; determine the interaction frequency of each bin pair of the bin pairs; normalize each interaction frequency of the interaction frequencies to generate a normalized interaction frequency for each bin pair; and provide, via the communication network, a map for displaying chromatin interactions on a user interface, including an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies.
[0128] 16. The computing device according to aspect 15, wherein, in order to normalize each interaction frequency of the interaction frequencies, the instructions cause the computing device to: determine the variation of the density of the set of paired contacts with genomic distance to generate a density function; for each bin pair of the plurality of bin pairs: integrate the density function over the region of the bin pair to determine the expected density of the bin pair; compare the subset of paired contacts within the bin pair with the expected density of the bin pair by performing a statistical analysis using a Poisson statistical distribution to determine the likelihood that the actual density of the bin pair is significantly greater than the expected density of the bin pair; apply a false discovery rate for multiple comparison control to the determined likelihood to determine an adjusted likelihood; and when the adjusted likelihood is less than a threshold likelihood, determine that the bin pair has enriched contacts, wherein the density function is generated from empirical data, and at least a portion of the density function decreases as the genomic distance increases.
[0129] 17. The computing device according to any one of aspects 15 or 16, wherein the instructions further cause the computing device to: identify a single locus in a DNA sequence that is associated with or causally related to one or more molecular phenotypes; identify a set of bins containing the single locus; obtain chromatin interaction data for a subject; compare the chromatin interaction data for the bins containing the single locus with contact data for such bins in another biological cell system; and predict the molecular phenotype of the subject based on the comparison.
[0130] 18. The computing device according to any one of aspects 15 to 17, wherein the instructions further cause the computing device to: generate a 3D or 4D model of a chromosomal structure based on the mapping of chromatin interactions; or generate a spatial interaction network for a set of specific loci.
[0131] 19. The computing device according to any one of aspects 15 to 18, wherein a binary search tree is used to identify the subset of paired contacts within each bin pair, the first set of the plurality of bins and the second set of the plurality of bins being the same bins corresponding to the same chromosome, wherein each genomic element corresponds to a locus within the genome, and wherein each bin corresponds to a contiguous segment of a deoxyribonucleic acid (DNA) sequence containing at least one of: a topologically associating domain (TAD), a gene, a chromatin state segment, a loop domain, or a chromatin domain.
[0132] 20. The computing device according to any one of aspects 15 to 19, wherein, in order to identify the first set of the plurality of bins and the second set of the plurality of bins, the instructions cause the computing device to receive a selection of the first set of bins and the second set of bins for a genome-wide search for long-range interactions, a genome-wide mapping of regulatory circuits, a comprehensive assessment of cell-type variability in long-range interactions, or the identification of a set of Hi-C-based diagnostic and prognostic biomarkers.
Claims
1. A computer-implemented method for analyzing the spatial and temporal organization of chromatin, the method being executed by one or more processors programmed to perform the method, the method comprising: Obtaining, at one or more processors, a set of paired contacts of genomic elements; Partitioning, by the one or more processors, the genomic elements into a plurality of bins, wherein the bin sizes of the plurality of bins are inconsistent; Identifying, by the one or more processors, a first set of the plurality of bins and a second set of the plurality of bins; Generating, by the one or more processors, a matrix of n×m bin pairs, where n corresponds to the first set of the plurality of bins and m corresponds to the second set of the plurality of bins; Identifying, by the one or more processors, a subset of paired contacts within each bin pair of the bin pairs; Determining, by the one or more processors, the interaction frequency of each bin pair of the bin pairs; Normalizing, by the one or more processors, each interaction frequency of the interaction frequencies to generate a normalized interaction frequency for each bin pair; And Providing, by the one or more processors, a map of chromatin interactions for display on a user interface, including an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies, wherein normalizing each interaction frequency of the interaction frequencies comprises: Determining, by the one or more processors, the variation of the density of the set of paired contacts with genomic distance to generate a density function; For each bin pair of the plurality of bin pairs: Integrating, by the one or more processors, the density function over the region of the bin pair to determine the expected density of the bin pair; Comparing, by the one or more processors, the subset of paired contacts within the bin pair with the expected density of the bin pair by performing a statistical analysis using a Poisson statistical distribution to determine the likelihood that the actual density of the bin pair is significantly greater than the expected density of the bin pair; Applying, by the one or more processors, a false discovery rate for multiple comparison control to the determined likelihood to determine an adjusted likelihood; And When the adjusted likelihood is less than a threshold likelihood, determining, by the one or more processors, that the bin pair has enriched contacts.
2. The method according to claim 1, further comprising: Performing, by the one or more processors, a statistical analysis using a Poisson distribution to determine a second likelihood that the amount by which the actual density of the bin pair is significantly less than the expected density of the bin pair; Applying, by the one or more processors, a false discovery rate for multiple comparison control to the determined second likelihood to determine an adjusted second likelihood; And When the adjusted second likelihood is less than a threshold likelihood, determining, by the one or more processors, that the bin pair has depleted contacts.
3. The method according to claim 2, wherein the statistical analysis comprises a two-tailed test for determining a third likelihood that the amount by which the actual density of the bin pair differs from the expected density is statistically significant; The one or more processors apply an error discovery rate for multiple comparison control to the determined third likelihood to determine an adjusted third likelihood; and when the adjusted third likelihood is less than a threshold likelihood, the one or more processors determine that the bin pair has enriched or depleted contacts.
4. The method according to claim 1, wherein at least some of the paired contacts are cis contacts such that the two genomic elements in each of the at least some paired contacts correspond to the same chromosome; and wherein at least some of the paired contacts are trans contacts such that the two genomic elements in each of the at least some paired contacts correspond to different chromosomes.
5. The method according to claim 1, wherein the density function is generated from empirical data and at least a portion of the density function decreases as the genomic distance increases.
6. The method according to claim 1, further comprising: the one or more processors identifying single loci in a DNA sequence that are associated with or causally related to one or more molecular phenotypes; the one or more processors identifying a set of bins containing the single locus; the one or more processors obtaining chromatin interaction data for a subject; the one or more processors comparing the chromatin interaction data of the bins containing the single locus with contact data of such bins in another biological cell system; and the one or more processors predicting the molecular phenotype of the subject based on the comparison.
7. The method according to claim 1, further comprising: the one or more processors generating a 3D or 4D model of the chromosomal structure based on the mapping of chromatin interactions.
8. The method according to claim 1, further comprising: the one or more processors generating a spatial interaction network for a set of specific loci.
9. The method according to claim 1, wherein identifying the subset of paired contacts within each bin pair comprises identifying the subset of paired contacts within each bin pair using a binary search tree.
10. The method according to claim 1, wherein the first set of the plurality of bins and the second set of the plurality of bins are the same bins corresponding to the same chromosome.
11. The method according to claim 1, wherein each genomic element corresponds to a locus within the genome; and wherein each bin corresponds to a contiguous segment of a deoxyribonucleic acid (DNA) sequence comprising at least one of: a topologically associating domain (TAD), a gene, a chromatin state segment, a loop domain, or a chromatin domain.
12. The method according to claim 11, wherein identifying the first set of the plurality of bins and the second set of the plurality of bins comprises receiving a selection of the first set of bins and the second set of bins for a genome-wide search for long-range interactions, a genome-wide mapping of regulatory circuits, a comprehensive assessment of intercellular type variability in long-range interactions, or identifying a set of Hi-C-based diagnostic and prognostic biomarkers.
13. The method according to claim 1, further comprising: For one or more of the bin pairs, comparing, by the one or more processors, the actual density of the bin pair from a first biological cell system or physiological condition with the actual density of the bin pair from a second biological cell system or physiological condition to identify differential contacts.
14. A computing device for analyzing the spatial and temporal organization of chromatin, the computing device comprising: A communication network; One or more processors; And A non-transitory computer-readable memory coupled to the one or more processors and storing instructions thereon, the instructions when executed by the one or more processors cause the computing device to: Obtain a set of paired contacts of genomic elements; Partition the genomic elements into a plurality of bins, wherein the bin sizes of the plurality of bins are inconsistent; Identify a first set of the plurality of bins and a second set of the plurality of bins; Generate a matrix of n×m bin pairs, where n corresponds to the first set of the plurality of bins and m corresponds to the second set of the plurality of bins; Identify a subset of paired contacts within each bin pair of the bin pairs; Determine the interaction frequency of each bin pair of the bin pairs; Normalize each of the interaction frequencies to generate a normalized interaction frequency for each bin pair; And Provide a map of chromatin interactions via the communication network for display on a user interface, including an indication of the bin pairs and a corresponding indication of the normalized interaction frequencies, wherein to normalize each of the interaction frequencies, the instructions cause the computing device to: Determine the variation of the density of the set of paired contacts with genomic distance to generate a density function; For each bin pair of the plurality of bin pairs: Integrate the density function over the region of the bin pair to determine the expected density of the bin pair; Compare the subset of paired contacts within the bin pair with the expected density of the bin pair by performing a statistical analysis using a Poisson statistical distribution to determine the likelihood that the actual density of the bin pair is significantly greater than the expected density of the bin pair; Apply a false discovery rate for multiple comparison control to the determined likelihood to determine an adjusted likelihood; and Determine that the bin pair has enriched contacts when the adjusted likelihood is less than a threshold likelihood, wherein the density function is generated from empirical data and at least a portion of the density function decreases as the genomic distance increases.
15. The computing device according to claim 14, wherein the instructions further cause the computing device to: Identify single loci in the DNA sequence that are associated with or causally related to one or more molecular phenotypes; Identify a set of bins containing the single locus; Obtain chromatin interaction data of a subject; Compare the chromatin interaction data of the bins containing the single locus with the contact data of such bins in another biological cell system; and Predict the molecular phenotype of the subject based on the comparison.
16. The computing device according to claim 14, wherein the instructions further cause the computing device to: generate a 3D or 4D model of the chromosomal structure based on the mapping of chromatin interactions; or generate a spatial interaction network of a set of specific loci.
17. The computing device according to claim 14, wherein a binary search tree is used to identify the subset of paired contacts within each bin pair, wherein the first set of the plurality of bins and the second set of the plurality of bins are the same bins corresponding to the same chromosome, wherein each genomic element corresponds to a locus within the genome, and wherein each bin corresponds to a contiguous segment of a deoxyribonucleic acid (DNA) sequence comprising at least one of the following: topologically associating domain (TAD), gene, chromatin state segment, loop domain, or chromatin domain.
18. The computing device according to claim 17, wherein, in order to identify the first set of the plurality of bins and the second set of the plurality of bins, the instructions cause the computing device to receive a selection of the first set of bins and the second set of bins for a genome-wide search for long-range interactions, a genome-wide mapping of regulatory circuits, a comprehensive assessment of intercellular type variability in long-range interactions, or the identification of a set of Hi-C based diagnostic and prognostic biomarkers.
Citation Information
Patent Citations
Method and apparatus for sequencing data samples
WO2009155443A2