Genome-scale imaging of chromatin 3D structure and transcriptional activity

Multiplexed FISH techniques like MERFISH overcome limitations of existing methods by enabling high-throughput imaging of chromatin and transcriptional activity in single cells, providing detailed insights into chromatin-nuclear interactions.

JP7849034B2Active Publication Date: 2026-04-21PRESIDENT & FELLOWS OF HARVARD COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PRESIDENT & FELLOWS OF HARVARD COLLEGE
Filing Date
2020-12-18
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current methods for studying the 3D structure of the genome and transcriptional activity in single cells are limited by low throughput and the inability to simultaneously measure chromatin composition and transcriptional activity, lacking tools for direct visualization of chromatin configuration in its native environment at both chromosome and genome scales.

Method used

Employing multiplexed FISH techniques, such as MERFISH, to image chromatin and nascent RNA in single cells, allowing for high-throughput imaging of multiple genomic loci and incorporating error-checking and error-correcting codes to enhance accuracy and efficiency.

Benefits of technology

Enables genome-scale imaging of chromatin structure and transcriptional activity, providing insights into chromatin-nuclear interactions and transcriptional relationships with high detection efficiency and spatial resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007849034000008
    Figure 0007849034000008
  • Figure 0007849034000009
    Figure 0007849034000009
  • Figure 0007849034000010
    Figure 0007849034000010
Patent Text Reader

Abstract

The present invention relates generally to genomics. Some embodiments are directed to imaging the 3D organization of a genome or a portion of a genome in sequence space with high throughput. Some embodiments are directed to imaging the 3D organization of a genome or a portion of a genome with respect to transcriptional activity and nuclear structure. In addition, certain embodiments are directed to chromatin structure, 3D chromatin organization, chromosomal trans-interactions, and chromatin-nuclear structure interactions, and their relationship to transcription. In addition, various embodiments are directed to imaging methods that enable mapping of the 3D organization of a genome or a portion of a genome with respect to nuclear structure and transcriptional activity. Some embodiments are directed to massively multiplexed fluorescence in situ hybridization methods for imaging chromatin loci and / or nascent RNA transcripts at the chromosomal or genome scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 954,720, filed Dec. 30, 2019, and entitled "Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin" by Zhuang et al.; and U.S. Provisional Patent Application No. 63 / 060,947, filed Aug. 4, 2020, and entitled "Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin" by Zhuang et al. Each of these is hereby incorporated by reference in its entirety.

[0002] The present invention generally relates to genomics. Some embodiments are directed to imaging the 3D organization of the genome with respect to transcriptional activity and nuclear structure. Additionally, certain embodiments are directed to chromatin composition and chromatin-nuclear structure interactions, and their relationship to transcription.

Background Art

[0003] The three-dimensional (3D) structure of the genome controls many essential cellular functions, from gene expression to DNA replication. Biochemical and imaging techniques have revealed complex chromatin structures across a wide range of scales. Recently, high-throughput chromosome conformation capture methods such as Hi-C and other sequencing-based methods have greatly deepened our understanding of 3D genome structure, revealing chromatin structures such as loops, domains, and compartments from a genome-wide perspective. However, these powerful sequencing-based techniques also have limitations. For example, these methods provide contact information between pairs of chromatin loci, but not direct spatial location information about these loci. Furthermore, most genome-wide insights into chromatin structure are constructed on population-average contact maps spanning millions of cells. Despite continuous improvements in single-cell Hi-C methods, the efficiency of capturing chromatin contacts in single-cell and / or cell throughput remains relatively low, and therefore, investigating 3D genome structure in single cells remains a challenging task. In addition, methods have emerged that combine Hi-C with other measurement methods to characterize chromatin contacts, for example, with respect to interacting proteins, nuclear structure, or DNA modifications, but multi-mode sequencing remains difficult. Notably, no method has emerged that allows for genome-scale measurement of both chromatin composition and transcriptional activity in the same cell, but such a method is desperately needed, as understanding how chromatin composition controls transcription, and in turn how transcription influences chromatin composition, is critically important.

[0004] On the other hand, imaging-based methods offer direct measurement of the spatial location of chromatin loci in individual cells with high detection efficiency. In particular, fluorescence in situ hybridization (FISH) enables highly specific detection of chromatin loci in fixed cells, and more recently, the CRISPR (clustered regularly interspersed short palindromic repeat) system has greatly improved our ability to image specific chromatin loci in living cells. Furthermore, chromatin imaging can be combined with RNA and protein imaging to reveal the interrelationships between chromatin composition and transcriptional activity or interacting protein factors. However, current imaging methods are limited in throughput in sequence space and traditionally only allow the study of a few different genomic loci at a time. Genome-scale imaging may require a significant increase in the number of genomic loci imaged in individual cells. Therefore, new improvements are needed. [Overview of the Initiative]

[0005] This invention, as a whole, relates to genomics. Some embodiments relate to imaging the 3D structure of a genome or a portion of a genome in sequence space at high throughput. Some embodiments relate to imaging the 3D structure of a genome or a portion of a genome with respect to transcriptional activity and nuclear structure. In addition, certain embodiments relate to chromatin structure, 3D chromatin structure, chromosome-trans interactions, and chromatin-nuclear structure interactions, as well as their relationship to transcription. The subject matter of this disclosure may include, depending on the context, multiple different uses of interrelated products, alternative solutions to specific problems, and / or one or more systems and / or items.

[0006] Certain embodiments, as a whole, cover systems and methods that use multiplexed FISH to image chromatin, for example, in cells, and in some cases, systems and methods that use multiplexed error-robust FISH (MERFISH). In addition, certain embodiments, as a whole, cover systems and methods for imaging and / or determining at least 100 or at least 500 distinct genomic loci within a single cell. Some embodiments, as a whole, cover systems and methods that use FISH to image chromatin, for example, in cells.

[0007] In one set of embodiments, the method includes associating multiple nucleic acid targets of a genome with multiple codewords, wherein the codewords include several positions and values ​​for each position; exposing a sample containing the genome to multiple nucleic acid probes; determining the binding of each nucleic acid probe in the sample for each of the multiple nucleic acid probes; generating codewords corresponding to the binding of the multiple nucleic acid probes in the sample; and determining the identification of nucleic acid targets based on the assigned codewords.

[0008] In another set of embodiments, the method includes determining the location of nascent RNA in the nucleus; applying an RNase to the nucleus; and determining the location of DNA in the nucleus.

[0009] In one set of embodiments, the method includes using MERFISH to image intracellular chromatin. In another set of embodiments, the method includes imaging at least 100 or at least 500 distinct genomic loci within a single cell.

[0010] According to one set of embodiments, the method involves: associating multiple nucleic acid targets of a genome with multiple codewords; exposing a sample containing cells believed to contain a genome to multiple nucleic acid probes, wherein at least a portion of the multiple nucleic acid probes comprises a first portion containing a target sequence and a second portion containing one or more readout sequences, each readout sequence representing a positional value within the multiple codewords; exposing a sample to one or more adapters in a round, each adapter comprising a first portion substantially complementary to one of the readout sequences and a second portion containing one identification sequence; exposing a sample to one or more readout probes in a round to determine one or more identification sequences, each readout Exposure of a readout probe comprising a first portion containing a sequence substantially complementary to one of the identification sequences and a second portion containing a signal-generating entity; determining the signal-generating entity at at least some locations in the sample; and inactivating the signal-generating entity at at least some locations in the sample; repeating the steps of exposing the sample to one or more adapters and one or more readout probes in a round, determining the signal-generating entity, and inactivating the signal-generating entity, wherein one or more distinct signal-generating entities are used in each round; determining a codeword at a location based on determining the signal-generating entity in the sample; and determining a nucleic acid target in the sample based on the codeword.

[0011] In yet another set of embodiments, the method involves associating multiple nucleic acid targets of a genome with multiple codewords; exposing a sample containing cells that are thought to contain a genome to multiple nucleic acid probes, wherein at least some of the nucleic acid probes comprise a first portion comprising a target sequence and a second portion comprising one or more readout sequences, each readout sequence representing a positional value within the multiple codewords; exposing a sample to one or more adapters in a round, each adapter comprising a first portion substantially complementary to one of the readout sequences and a second portion comprising one identification sequence; exposing a sample to one or more readout probes in a round to determine one or more identification sequences, wherein each readout sequence Exposure of a lobe comprising a first portion containing a sequence substantially complementary to one of the identification sequences and a second portion containing a signal-generating entity; determining the signal-generating entity at at least some locations in the sample; and inactivating the signal-generating entity at at least some locations in the sample; repeating the steps of exposing the sample to one or more adapters and one or more readout probes in one round, determining the signal-generating entity, and inactivating the signal-generating entity, wherein at least one of the signal-generating entity is used in more than one round; determining a codeword at a location based on determining the signal-generating entity in the sample; and determining a nucleic acid target in the sample based on the codeword.

[0012] In yet another set of embodiments, the method involves exposing a sample containing cells that are thought to contain a genome to a plurality of nucleic acid probes, wherein at least a portion of the plurality of nucleic acid probes comprises a first portion containing a target sequence and a second portion containing one or more readout sequences; exposing the sample to one or more adapters in a round, wherein each adapter comprises a first portion substantially complementary to one of the readout sequences and a second portion containing one identification sequence; and exposing the sample to one or more readout probes in a round to determine one or more identification sequences, wherein each readout probe comprises one of the identification sequences. Exposure to a sample comprising a first portion containing a sequence substantially complementary to one of the samples and a second portion containing a signal-generating entity; determining the signal-generating entity at at least some locations in the sample; and inactivating the signal-generating entity at at least some locations in the sample; repeating the steps of exposing the sample to one or more adapters and one or more readout probes in a round, determining the signal-generating entity, and inactivating the signal-generating entity, wherein one or more distinct signal-generating entities are used in each round; and determining nucleic acid targets in the sample based on the signal-generating entity determined in each round.

[0013] In another set of embodiments, the method includes: exposing a sample containing cells that are thought to contain a genome to a plurality of nucleic acid probes, wherein at least a portion of the plurality of nucleic acid probes comprises a first portion containing a target sequence and a second portion containing one or more readout sequences; exposing the sample to one or more readout probes in a round to determine one or more readout sequences, wherein each readout probe comprises a first portion containing a sequence substantially complementary to one of the readout sequences and a second portion containing a signal-generating entity; determining a signal-generating entity at at least a portion of the locations in the sample; and inactivating a signal-generating entity at at least a portion of the locations in the sample; repeating the steps of exposing the sample to one or more readout probes in a round, determining a signal-generating entity, and inactivating a signal-generating entity, wherein one or more distinct signal-generating entities are used in each of the rounds; and determining a nucleic acid target in the sample based on the signal-generating entity determined in each round.

[0014] In yet another set of embodiments, the method includes exposing a sample containing cells that are thought to contain a genome to a round of multiple nucleic acid probes, wherein at least a portion of the multiple nucleic acid probes comprises a first portion containing a target sequence and a second portion containing a signal-generating entity; determining the signal-generating entity at at least a portion of the locations in the sample; and inactivating the signal-generating entity at at least a portion of the locations in the sample; repeating the steps of exposing the sample to a round of multiple nucleic acid probes, determining the signal-generating entity, and inactivating the signal-generating entity, wherein one or more distinct signal-generating entities are used in each round; and determining a nucleic acid target in the sample based on the signal-generating entity determined in each round.

[0015] In one set of embodiments, the method is to associate multiple nucleic acid targets in a genome with multiple codewords, the codewords comprising several positions and values ​​for each position, the codewords forming an error-checking and / or error-correcting code space, and the multiple nucleic acid targets being separated by at least 100,000 nucleotides within the genome; to expose the nucleus of a cell containing a genome to multiple nucleic acid probes, at least a portion of the multiple nucleic acid probes comprising a first portion comprising a target sequence and a second portion comprising one or more read sequences, each read sequence representing a position value within a codeword; and each nucleic acid probe of the multiple nucleic acid probes The method includes determining the binding of nucleic acid probes within the nucleus; generating codewords corresponding to the binding of multiple nucleic acid probes within the nucleus, wherein the numerical value of the codeword is based on the read sequence present on the nucleic acid probe; matching at least some of the codewords to valid codewords, where, if no match is found, the codeword is rejected or error correction is applied to the codeword to form a valid codeword, wherein the valid codeword is a set of codewords assigned to multiple nucleic acid targets; and determining the abundance and / or spatial distribution of nucleic acids within the nucleus using the valid codewords corresponding to the binding of multiple nucleic acid probes within the nucleus.

[0016] In another set of embodiments, the method is to associate multiple nucleic acid targets of a genome with multiple codewords, wherein the codewords include several positions and values ​​for each position, the codewords form an error-checking and / or error-correcting code space, and the multiple nucleic acid targets of the genome are distributed such that each chromosome of the genome contains no more than 200 nucleic acid targets; to expose the nucleus of a cell containing a genome to multiple nucleic acid probes, wherein at least a portion of the multiple nucleic acid probes includes a first portion containing a target sequence and a second portion containing one or more read sequences, each read sequence representing a position value in the codeword; and each nucleic acid of the multiple nucleic acid probes The method includes determining the binding of a nucleic acid probe within the nucleus; generating codewords corresponding to the binding of multiple nucleic acid probes within the nucleus, wherein the numerical value of the codeword is based on the read sequence present on the nucleic acid probe; matching at least some of the codewords to valid codewords, where, if no match is found, the codeword is rejected or error correction is applied to the codeword to form a valid codeword, wherein the valid codeword is a set of codewords assigned to multiple nucleic acid targets; and determining the abundance and / or spatial distribution of nucleic acids within the nucleus using the valid codewords corresponding to the binding of multiple nucleic acid probes within the nucleus.

[0017] According to another set of embodiments, the method involves associating 500 to 1500 nucleic acid targets of a genome with a plurality of codewords, wherein the codewords include several positions and values ​​for each position, and the codewords form an error-checking and / or error-correcting code space; exposing the nucleus of a cell containing the genome to a plurality of nucleic acid probes, wherein at least a portion of the plurality of nucleic acid probes includes a first portion containing a target sequence and a second portion containing one or more read sequences, wherein each read sequence represents a positional value in the codeword; and for each nucleic acid probe of the plurality of nucleic acid probes, the nucleic acid probe in the nucleus The method includes determining binding; generating codewords corresponding to the binding of multiple nucleic acid probes in the nucleus, wherein the numerical value of the codeword is based on the read sequence present on the nucleic acid probe; matching at least some of the codewords to valid codewords, wherein if no match is found, the codeword is rejected or error correction is applied to the codeword to form a valid codeword, wherein the valid codeword is a set of codewords assigned to multiple nucleic acid targets; and determining the abundance and / or spatial distribution of nucleic acids in the nucleus using the valid codewords corresponding to the binding of multiple nucleic acid probes in the nucleus.

[0018] In yet another set of embodiments, the method comprises associating multiple nucleic acid targets in a genome with multiple codewords, the codewords comprising several positions and values ​​for each position, the codewords forming an error-checking and / or error-correcting code space, and the multiple nucleic acid targets being separated by at least 100,000 nucleotides within the genome; exposing the nucleus of a cell containing the genome to multiple nucleic acid probes; and determining the abundance and / or spatial distribution of nucleic acids in the nucleus by determining the binding of the multiple nucleic acid probes in the nucleus using an error-checking and / or error-correcting detection method.

[0019] In yet another set of embodiments, the method involves: associating multiple nucleic acid targets of a genome with multiple codewords; exposing a sample containing cells presumably containing a genome to multiple nucleic acid probes, wherein at least a portion of the nucleic acid probes comprises a first portion containing a target sequence and a second portion containing one or more read sequences, each read sequence representing a positional value within the multiple codewords; exposing the sample to multiple adapters, wherein at least a portion of the adapters comprises a first portion substantially complementary to one or more read sequences and a second portion containing one or more identification sequences; and exposing the sample to one or more readout probes in a round to determine one or more identification sequences. Exposure, wherein at least a portion of the readout probe comprises a first portion containing a sequence substantially complementary to one of the identification sequences and a second portion containing a signal-generating entity; determining the signal-generating entity at at least some locations in the sample; and inactivating the signal-generating entity at at least some locations in the sample; repeating the steps of exposing the sample to rounds, determining the signal-generating entity, and inactivating the signal-generating entity, wherein not more than 10 distinct signal-generating entities are used in all rounds; determining a codeword at a location based on determining the signal-generating entity in the sample; and determining a nucleic acid target in the sample based on the codeword.

[0020] According to yet another set of embodiments, the method involves: associating multiple nucleic acid targets of a genome with multiple codewords; exposing a sample containing cells presumably containing a genome to multiple nucleic acid probes, wherein at least a portion of the multiple nucleic acid probes comprises a first portion containing a target sequence and a second portion containing one or more read sequences, each read sequence representing a positional value within the multiple codewords; exposing the sample to multiple adapters, wherein at least a portion of the adapters comprises a first portion substantially complementary to one or more read sequences and a second portion containing one or more identification sequences; and exposing the sample to one or more readout probes in a round to determine one or more identification sequences. Exposure, wherein at least a portion of the readout probe comprises a first portion containing a sequence substantially complementary to one of the identification sequences and a second portion containing a signal-generating entity; determining the signal-generating entity at at least some locations in the sample; and inactivating the signal-generating entity at at least some locations in the sample; repeating the steps of exposing the sample to round 1, determining the signal-generating entity, and inactivating the signal-generating entity, wherein at least one of the signal-generating entity is used in more than one round; determining a codeword at a location based on determining the signal-generating entity in the sample; and determining a nucleic acid target in the sample based on the codeword.

[0021] According to another set of embodiments, the method includes determining the location of nascent RNA in the nucleus; determining the location of DNA in the nucleus; and determining the location of nuclear speckles in the nucleus.

[0022] In yet another set of embodiments, the method includes determining the location of nascent RNA in the nucleus; determining the location of DNA in the nucleus; and determining the location of proteins in the nucleus. In yet another set of embodiments, the method includes determining the location of nascent RNA in the nucleus; determining the location of DNA in the nucleus; and determining the location of nucleic acids in the nucleus, where the nucleic acid is neither nascent RNA nor DNA.

[0023] Some aspects include methods of manufacturing one or more of the embodiments described herein. Some aspects also include methods of using one or more of the embodiments described herein.

[0024] Other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying drawings.

[0025] Non-limiting embodiments of the present invention are described by way of example with respect to the accompanying drawings, which are schematic and are not intended to be drawn to scale. In the figures, each component shown that is identical or substantially the same is typically represented by a single number. For clarity, not every component may be shown in every figure, nor may every component of each embodiment of the invention be shown, if illustration is not necessary for one of ordinary skill in the art to understand the present disclosure.

Brief Description of the Drawings

[0026] Figures 1A-1I show genome-scale chromatin imaging according to certain embodiments. Figures 2A-2E show trans-chromosomal contact enrichment in another embodiment. Figures 3A-3H show genome-scale imaging of chromatin and transcriptional activity in the context of nuclear structure in yet another embodiment. Figures 4A-4F show trans-chromosomal interactions between active chromatin in another embodiment. Figures 5A-5E illustrate a saturation growth system in one embodiment. Figures 6A-6B show a contact frequency matrix in yet another embodiment. Figures 7A-7C show a comparison of sub-chromosomal structures and ensemble Hi-C data derived from genome-scale imaging in yet another embodiment. Figure 8 shows the reproducibility of chromatin imaging experiments between repeats in another embodiment. Figures 9A-9B show different spatial distributions in a single cell in a particular embodiment. Figures 10A-10B show nascent RNA transcript imaging in still other embodiments. Figure 11 shows the association of compartment B loci with the nuclear lamina in a particular embodiment. Figure 12 shows the association of compartment A loci with nuclear speckles in some embodiments. Figures 13A-13C show changes in the association of the nuclear lamina and nuclear speckles during transcriptional inhibition in yet another embodiment. Figure 14 shows the local density of trans-chromosomal A loci near each imaged locus in yet another embodiment. Figures 15A-15B show the enrichment of active-active trans-chromosomal interactions between chromatin loci in yet another embodiment. Figures 16A-16B show the enrichment of active-active trans-chromosomal interactions in yet another embodiment. Figures 17A-17M show high-resolution whole chromosome tracing by sequential hybridization and characterization of chromatin domains in a single cell in one embodiment. Figures 18A-18I show the relationship between compartment structure, transcriptional activity, and local chromatin content in a single chromosome in another embodiment. Figures 19A to 19H show the dependence of interdomain interactions on A / B composition and genomic distance in yet another embodiment. Figures 20A to 20H show genome-scale chromatin imaging by large-scale multiplexed combinatorial FISH in yet another embodiment. Figures 21A to 21E illustrate the enrichment of active-active chromatin interactions in trans chromosome interactions according to one embodiment. Figures 22A–22J show multimode genome-scale imaging of chromatin and transcriptional activity in the context of nuclear structure, according to another embodiment. Figures 23A to 23D show the correlation between the transcriptional activity of trans-chromosome active chromatin and local enrichment in yet another embodiment. Figures 24A to 24N show ensemble statistics of Chr21 structural properties compared to high-resolution whole-chromosome tracing by sequential hybridization and Hi-C in yet another embodiment. Figures 25A to 25G show an ensemble A / B compartment analysis relating to Chr21 and Chr2 in yet another embodiment. Figures 26A to 26J show the measurement of RNA and DNA FISH probe crosstalk in yet another embodiment. Figures 27A to 27J show genome-scale imaging in one embodiment, comparing combinatorial FISH: localization error, reproducibility, and Hi-C. Figures 28A and 28B show that, in another embodiment, the loci of compartment A and compartment B exhibit different spatial distributions within the nucleus. Figures 29A to 29F show the effect of transcriptional inhibition on trans chromosome chromatin interactions and the nuclear structure association rate of chromatin loci in yet another embodiment. Figures 30A to 30D illustrate the enrichment of trans-chromosome active chromatin interactions in different nuclear environments in yet another embodiment. [Modes for carrying out the invention]

[0027] This invention relates, as a whole, to genomics. Some embodiments focus on high-throughput imaging of the 3D structure of a genome or a portion of a genome in sequence space. Some embodiments focus on imaging the 3D structure of a genome or a portion of a genome with respect to transcriptional activity and nuclear structure. In addition, certain embodiments focus on chromatin structure, 3D chromatin structure, chromosome-trans interactions, and chromatin-nuclear structure interactions, as well as their relationship to transcription. In addition, a variety of embodiments focus on imaging methods that enable mapping of the 3D structure of a genome or a portion of a genome with respect to nuclear structure and transcriptional activity. Some embodiments focus on large-scale multiplexed fluorescence in situ hybridization methods for imaging chromatin loci and / or nascent RNA transcripts at the chromosome or genome scale. In some cases, simultaneous imaging of hundreds of genomic loci can be performed. In some cases, simultaneous imaging of the transcriptional activity of approximately 1000 genomic loci and / or approximately 1000 genes within these loci, with diverse nuclear structures, can be performed. In certain cases, chromatin domains and compartments can be observed. In certain cases, a wide range of chromosomal trans interactions enriched with active chromatin interactions in a manner correlated with transcription can be observed. In some cases, transcription-dependent chromatin interactions with nuclear speckles and nuclear laminas across the genome can be observed.

[0028] The three-dimensional (3D) chromatin configuration controls many genomic functions. Understanding 3D genomic configuration is hampered by the lack of tools that allow for direct visualization of chromatin configuration in its native environment at both the chromosome and genome scales. Therefore, in certain embodiments, a multiplexed FISH technique is described, involving sequential imaging across multiple hybridization rounds, for example, each round targeting one, two, or three genomic loci using one-color, two-color, or three-color imaging. In other embodiments, a combinatorial FISH technique is described, in which many chromatin loci are imaged simultaneously in each round, and distinct identification of loci is determined based on the combination of rounds in which the loci appear. This is based on MERFISH and other techniques, as discussed in their entirety, for example, in International Patent Application Publication WO2016 / 018960, entitled “Systems and Methods for Determining Nucleic Acids” and International Patent Application Publication WO2016 / 018963, entitled “Probe Library Construction,” each of which is incorporated herein by reference. The techniques, for example, those discussed herein, may be used to image distinct chromatin loci within a single cell and to provide insights into chromatin structure, their relationship to transcription, and their interactions with nuclear proteins.

[0029] Some embodiments relate to systems and methods that, in whole, use multiplexed FISH or other techniques, including the techniques described herein, for imaging chromosomes or chromatin in cells, for example, and in some cases use MERFISH. In addition, certain embodiments relate to systems and methods that, in whole, image and / or determine at least 100 distinct genomic loci, at least 500 distinct genomic loci, or at least 1,000 distinct genomic loci, etc., within a single cell. In some cases, other parts of the cell or the nucleus, such as RNA present in the nucleus, such as nascent RNA, nuclear speckles, nucleoli, nuclear lamina, other nuclear structures or proteins, etc., may be determined. In non-limiting examples, the location of chromosomes or chromatin, nascent RNA, nuclear speckles, nucleoli, and / or nuclear lamina may be determined for the nucleus of a cell.

[0030] Certain embodiments relate to determining a sample that may include cell cultures, cell suspensions, biological tissues, biopsies, organisms, etc. The sample may also be cell-free, but nevertheless, it may contain nucleic acids. If the sample contains cells, the cells may be human cells or any other suitable cells, such as mammalian cells, fish cells, insect cells, plant cells, etc. In some cases, more than one type of cell may be present.

[0031] In a sample, the targets to be determined may include nucleic acids, proteins, and the like. For example, these may be present in the nucleus of cells in the sample. In certain embodiments, intracellular chromatin may be determined for the nuclear structure of a cell, including, for example, nuclear speckles, nucleoli, nuclear lamina, or nuclear structures or nuclear proteins. In some cases, chromatin loci and / or RNA transcripts may be determined within a cell, for example, at the chromosomal or genome scale.

[0032] One example of such a method is discussed here. However, it should be understood that this method is presented as an example and is not limiting; other aspects and embodiments are also discussed herein. In one set of embodiments, nucleic acids within a cell, for example, within the cell nucleus, should be determined. These typically include DNA (e.g., genomic DNA, which may exist in the form of chromatin, for example, in the form of chromatin packaged with proteins such as histones) and RNA (e.g., RNA at the initiation of the transcription phase in which DNA is transcribed into RNA; this RNA in the nucleus is sometimes referred to as nascent RNA). In contrast to techniques for detecting RNA, which may exist anywhere within the cell, DNA is highly compressed within the cell nucleus, and therefore it is substantially more difficult to determine its structure. For example, DNA may be compressed within the cell as chromosomes or chromatin, and such DNA may often be tightly bound together, tangled, or compressed within the nucleus. Therefore, in certain embodiments, DNA targets may be selected to be spatially separated.

[0033] In some cases, the sample is subjected to multi-round hybridization with a nucleic acid probe, where one or more rounds target one or more target nucleic acids by monochromatic or multichromatic imaging. In some cases, the identification of target nucleic acids is determined based on which round and / or which color channel they are imaged in. In some cases, the location of the target nucleic acids is determined. In some cases, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 target nucleic acids are determined. In some cases, the target nucleic acids are genomic loci. In some cases, the target nucleic acids are genomic loci and / or nascent RNA transcripts. In some cases, the location of the genomic loci is used to determine the three-dimensional structure of chromatin within a cell or the three-dimensional structure of the genome.

[0034] In some cases, primary nucleic acid probes are designed that can target nucleic acids within cells, for example, within the cell nucleus. Each probe contains a target sequence that binds to one of the target nucleic acids. The probe may also contain a portion comprising one or more “readout sequences” that can be used to identify and locate the primary nucleic acid probe. In some embodiments, the primary nucleic acid probe may contain multiple readout sequences. These can be read individually using one or more rounds of secondary nucleic acid probes, referred to as readout probes, which can bind to the readout sequences of the primary nucleic acid probe. The readout probes may also contain signal-generating entities, such as fluorescent entities, which can be determined, for example, using various microscopy techniques. In some cases, multiple rounds of readout probes may be applied sequentially, such that one type of readout probe is applied to a sample, fluorescence in the sample is determined, then the readout probe or the signal-generating entity on the readout probe is inactivated or removed, and the next type of readout probe is applied. In some cases, the location in the sample may be associated with multiple readout probes, and this information may be quantified for analysis.

[0035] In some cases, multiple rounds of readout probes may be applied sequentially, such that one or more types of readout probes are applied to the sample in each round, and / or fluorescence in the sample is determined using multicolor imaging, after which the readout probes and / or signal-generating entities on the readout probes are inactivated or removed, and the next set of two or more types of readout probes is applied. In some cases, the position in the sample may be associated with multiple readout probes, and this information may be quantified for analysis.

[0036] In some cases, the positions of the primary nucleic acid probe and target nucleic acid may be determined using one or more rounds of readout probes. For example, there may be readout probes with at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 rounds. Thus, in some cases, the sample may be exposed to multiple rounds of applied readout probes to determine the probes in the sample (e.g., using a signal-generating entity as described herein) and remove or inactivate the secondary nucleic acid probes.

[0037] Furthermore, it should be understood that readout probes do not all need to be different. In some cases, for example, to determine whether any degradation and / or migration has occurred in the sample over time, more than one round of the same readout probe may be used to account for the effect of supplying multiple rounds of nucleic acids or other chemicals as controls.

[0038] In some cases, the sample is subjected to multi-round hybridization with a nucleic acid probe, with each round subjected to monochromatic or multichromatic imaging. In some cases, the identification of target nucleic acids is determined based on which combination of rounds and / or color channels they are imaged in. In some cases, the location of the target nucleic acids is determined. In some cases, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 target nucleic acids are determined. In some cases, the target nucleic acids are genomic loci. In some cases, the target nucleic acids are genomic loci and / or nascent RNA transcripts. In some cases, the location of genomic loci is used to determine the three-dimensional structure of chromatin within a cell or the three-dimensional structure of the genome.

[0039] In some cases, primary nucleic acid probes (also called coding probes) are designed that can target nucleic acids within cells, for example, within the cell nucleus. Each probe contains a target sequence that binds to one of the target nucleic acids. The probe may also contain a portion containing one or more "readout sequences" that can be used to identify and locate the primary nucleic acid probe or coding nucleic acid probe. In some embodiments, the primary nucleic acid probe or coding nucleic acid probe may contain multiple readout sequences. These can be read individually using one or more rounds of readout probes that can bind to the readout sequences of the primary nucleic acid probe or coding nucleic acid probe. The readout probes may also contain signal-generating entities, such as fluorescent entities, which can be determined, for example, using various microscopy techniques. In some cases, multiple rounds of readout probes may be applied sequentially, such that one type of readout probe is applied to a sample, fluorescence in the sample is determined, then the readout probe or signal-generating entity on the readout probe is inactivated or removed, and the next type of readout probe is applied. In some cases, the position in the sample may be associated with multiple readout probes, and this information may be quantified for analysis. In some cases, multiple rounds of readout probes may be applied sequentially, such that one or more types of readout probes are applied to the sample in each round, fluorescence in the sample is determined using multicolor imaging, then the readout probe or the signal-generating entity on the readout probe is inactivated or removed, and the next set of one or more types of readout probes is applied. In some cases, the position in the sample may be associated with multiple readout probes, and this information may be quantified for analysis.

[0040] In some cases, the position of the primary nucleic acid probe or encoding nucleic acid probe and the target nucleic acid may be determined using one or more rounds of readout probes. For example, there may be readout probes with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 16, at least 20, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000 rounds. Thus, in some cases, the sample may be exposed to a readout probe with multiple rounds applied to determine the probes in the sample (e.g., using a signal-generating entity as described herein) and remove or inactivate the secondary nucleic acid probes.

[0041] Primary nucleic acid probes or coding nucleic acid probes may be designed in some embodiments such that various targets in a sample can be determined using various combinations of readout sequences, without necessarily requiring each readout sequence to be unique. As a non-limiting example, if a set of primary nucleic acid probes or coding nucleic acid probes targeting each nucleic acid target contains only two readout sequences, such as four possible readout sequences A, B, C, and D, then up to six different targets corresponding to AB, AD, CB, CD, AC, and DB may be identified.

[0042] However, in some embodiments, not all possible combinations of readout sequences are used. Instead, some of the combinations may not be assigned to any target in the nucleus, and for example, primary nucleic acid probes or coding nucleic acid probes having those combinations may not be used. In some cases, valid combinations of readout sequences used in primary nucleic acid probes or coding nucleic acid probes may be arranged to form an error-checking and / or error-correcting code space. Using such a method, the determination of readout sequences in a sample that do not correspond to a valid primary nucleic acid probe may be determined to be in error using error checking, and in some cases, may be further corrected using error correction to correspond to a valid primary nucleic acid probe, for example.

[0043] Such methods have been previously described, for example, in International Patent Application Publication WO2016 / 018960, entitled "Systems and Methods for Determining Nucleic Acids"; and International Patent Application Publication WO2016 / 018963, entitled "Probe Library Construction," but such methods have not been applied to imaging DNA in the more constrained environment within the cell nucleus. As noted, unlike the rest of the cell, the cell nucleus contains a very large proportion of nucleic acids, including almost all of the cell's genomic DNA, and typically high concentrations of RNA (e.g., nascent RNA).

[0044] Therefore, to access DNA within the nucleus of a cell, the targets of a primary nucleic acid probe or coding nucleic acid probe may be selected so that binding in the nucleus occurs in a spatially separated manner. For example, targets may be selected so that they are separated in genomic space, e.g., separated by at least 10,000 bp, at least 30,000 bp, at least 100,000 bp, at least 300,000 bp, or at least 1,000,000 bp in the genome, or so that the genomic space contains nucleic acid targets not exceeding 100, 200, 300, 500, 1000, 5000, 10,000, 50,000, or 100,000. In some cases, more than one type of fluorescent probe or "color" may be used, also to enable the determination of more targets in the nucleus, for example.

[0045] In some embodiments, cells and / or nuclei may also be modified to allow such probes to reach nucleic acids within the cells and / or nuclei. For example, cells may be permeabilized or “fixed” to allow entry of nucleic acid probes. In addition, DNA may be denatured, for example by applying heat, to allow easier access to DNA by primary nucleic acid probes or encoding nucleic acid probes in some embodiments. This is not typically done for RNA determination, as DNA is usually double-stranded while RNA is single-stranded. In addition, in certain embodiments, RNA in the nucleus must be removed and / or inactivated, for example, to prevent probes targeting DNA from binding to RNA, before DNA can be studied. In some cases, enzymes such as RNases may be applied to the nucleus to prevent RNA from interfering with DNA determination.

[0046] In addition, it should be noted that in certain embodiments, nuclear RNA can also be determined. This can be particularly useful when studying, for example, the spatial locations of DNA and RNA in the nucleus, and how they relate to each other. Thus, in one set of embodiments, nuclear RNA can be determined before the removal or inactivation of RNA as described above, similar to how genomic DNA is determined above.

[0047] In addition, in certain embodiments, intracellular proteins, such as those within the cell nucleus, can also be determined. Examples include, but are not limited to, nuclear speckles, nucleolis, or histone proteins. A variety of methods can be used to determine proteins. For example, in one set of embodiments, an immunofluorescence assay may be used. In another set of embodiments, a “sandwich assay” may be used, in which a primary antibody capable of specifically binding to nuclear proteins is applied, followed by a secondary antibody capable of specifically binding to the primary antibody, the secondary antibody containing a signal-generating entity such as a fluorescent entity. The determination of such proteins may be performed on the same sample or the same nucleus as described above, for example, before or after the determination of nuclear nucleic acids. Thus, in some cases, proteins and nucleic acids within the cell nucleus can be determined, for example, spatially.

[0048] The above discussion is a non-limiting example of one embodiment that may be used to determine nucleic acids such as genomic DNA and / or nascent RNA within the nucleus of a cell. However, other embodiments are also possible. Therefore, more as a whole, diverse embodiments cover a variety of systems and methods for nucleic acids.

[0049] As mentioned, in certain embodiments, one, two, or more of the following can be determined within a cell, for example, within the cell nucleus: DNA, RNA, and proteins. The nucleic acids in the nucleus to be determined may include, for example, DNA (e.g., genomic DNA), RNA, or other nucleic acids present in the cell (or other sample). Nucleic acids may be endogenous in the cell or added to the cell. For example, nucleic acids may be viral or artificially created. Depending on the case, the nucleic acids to be determined may be expressed by the cell. In some embodiments, the nucleic acid is RNA. RNA may be coding RNA and / or non-coding RNA. For example, RNA may code for proteins. Non-limited examples of RNA that can be studied within a cell include mRNA, siRNA, rRNA, miRNA, tRNA, lncRNA, snoRNA, snRNA, exRNA, piRNA, etc.

[0050] In one set of embodiments, all or at least a substantial portion of a cell's genome can be determined. The determined genomic segments may be continuous or interrupted on the genome. For example, in some cases, at least four genomic segments within a cell may be determined, and in some cases, at least three, at least four, at least seven, at least eight, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 25 6. At least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 genomic segments may be determined.

[0051] In some cases, the entire genome of a cell may be determined. It should be understood that the genome, in general, includes not only chromosomal DNA but all DNA molecules produced within the cell. Therefore, for example, the genome may also include, in some cases, mitochondrial DNA, chloroplast DNA, plasmid DNA, etc., in addition to (or not in addition to) chromosomal DNA. In some embodiments, at least about 0.01%, at least about 0.1%, at least about 1%, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or 100% of the cell's genome may be determined.

[0052] In addition, in some embodiments, a significant portion of nucleic acids within cells or within the cell nucleus may be studied. For example, in some cases, nuclear RNA, such as nascent RNA, may be determined. In addition, in some cases, a sufficient amount of RNA present in the cell may be determined to produce a partial or complete transcriptome of the cell. In some cases, at least four types of RNA (e.g., mRNA, nascent RNA, etc.) may be determined within the cell or within the cell nucleus, and in some cases, at least three, at least four, at least seven, at least eight, at least twelve, at least fourteen, at least fifteen, at least sixteen, at least twenty, at least two twenty, at least thirty, at least three one, at least three twenty, at least sixty, at least thirty, at least thirty, at least fifty, at least sixty, at least sixty, at least seventy, at least seventy, at least seventy, at least one hundred, at least one twenty-seven, at least one twenty-eight, at least one fourteen, at least one fourteen, at least two fifty, and at least one twenty-five. At least 256, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 types of RNA can be determined within a cell or within the cell nucleus.

[0053] In some cases, the cellular transcriptome can be determined. It should be understood that the transcriptome, in general, encompasses not only mRNA but all RNA molecules produced within the cell. Therefore, for example, the transcriptome may also include rRNA, tRNA, siRNA, etc., in certain cases. In some embodiments, at least about 0.01%, at least about 0.1%, at least about 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the cellular transcriptome can be determined. In addition, in some cases, the nuclear transcriptome of the cell can be determined.

[0054] Furthermore, in some embodiments, other targets to be determined may include targets conjugated to nucleic acids, proteins, and the like. For example, in one set of embodiments, a binding entity capable of recognizing a target may be conjugated to a nucleic acid probe. The binding entity can be any entity capable of recognizing the target, for example, specifically or nonspecifically. Non-limiting examples include enzymes, antibodies, receptors, complementary nucleic acid chains, aptamers, and the like. For example, an antibody conjugated to an oligonucleotide may be used to determine a target. The target can bind to the antibody conjugated to the oligonucleotide, and the oligonucleotide may be determined as discussed herein.

[0055] The determination of targets such as nucleic acids within cells or in other samples may be qualitative and / or quantitative. In addition, the determination can be spatial; for example, the location of nucleic acids or other targets within cells or in other samples may be determined in two or three dimensions. In some embodiments, the location, number, and / or concentration of nucleic acids or other targets within cells or in other samples may be determined.

[0056] As mentioned, in one set of embodiments, DNA in the nucleus of a cell, e.g., cellular genomic DNA, may be studied using nucleic acid probes, such as those discussed herein, including, for example, using sequential imaging with error detection codes and / or error correction codes or using combinatorial imaging.

[0057] In certain embodiments, DNA targets or codes associated with DNA targets within a cell or cell nucleus may be selected to be spatially separated, for example, in genomic space or physical space, based on knowledge of chromatin structure, such as the target being composed of small, clustered regions, for example, of chromosomes, in each round of imaging. This can be useful for enabling the identification of various intracellular targets, for example, in the cell nucleus, using techniques such as those discussed herein.

[0058] Targets within the genomic space may be selected using any appropriate technique, for example, randomly or with a substantially uniform probability distribution. In certain embodiments, targets may be selected individually to ensure spatial separation. In addition, in some embodiments, targets may be selected to be targets of interest within the genome, for example, for a particular study.

[0059] For example, in some embodiments, the target may be selected such that the nucleus has no more than a certain number of nucleic acid targets within the genomic space. For example, the target may be selected such that the genomic space contains no more than 100,000, no more than 10,000, no more than 8,000, no more than 6,000, no more than 5,000, no more than 4,000, no more than 3,000, no more than 2,000, no more than 1,500, no more than 1,000, no more than 900, no more than 800, no more than 700, no more than 600, no more than 500, no more than 400, no more than 300, no more than 200, no more than 100 nucleic acid targets, no more than 30 nucleic acid targets, or no more than 10 nucleic acid targets. In addition, in some embodiments, the target may be selected such that the genomic space contains at least 10, at least 30, at least 50, at least 100, at least 200, at least 300, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 3,000, at least 5,000, at least 10,000, and at least 100,000 nucleic acid targets. Any combination of these is also possible in certain embodiments, for example, 30-100, 3,000-5,000, and 500-1,500 nucleic acid targets may exist. Such targets may be selected, for example, selectively or randomly, as discussed herein.

[0060] As another example, in some embodiments, the target may be selected so that the chromosomes in the genome have no more than a certain number of nucleic acid targets (e.g., genomic loci). For example, the target may be selected so that each chromosome has no more than 10,000, no more than 1,000, no more than 500, no more than 400, no more than 300, no more than 200, no more than 150, no more than 125, no more than 100, no more than 90, no more than 80, no more than 70, no more than 60, no more than 50, no more than 40, no more than 30, no more than 20, or no more than 10 nucleic acid targets. In some cases, the target may be selected so that the chromosomes have no more than 10, no more than 20, no more than 30, no more than 40, no more than 60, no more than 50, no more than 40, no more than 30, no more than 20, or no more than 10,000 nucleic acid targets. In some cases, these combinations may be selected, for example, chromosomes may have 30-50, 40-100, 50-60, or 30-80 nucleic acid targets. In addition, different chromosomes may independently have the same or different numbers of nucleic acid targets, including, for example, the ranges described herein.

[0061] Such targets may be selected, for example, selectively or randomly, as discussed herein. In a non-limiting example, nucleic acid targets within a genome may be selected to possess specific structural or functional characteristics, such as promoters, enhancers, and loci to which specific nuclear architecture proteins are bound. In some cases, some or all of the nucleic acid targets may be unique to each of their respective chromosomes.

[0062] In another embodiment, targets may be selected so as to be separated by a minimum of a certain number of nucleotides, for example, to promote the distribution of spatially separated targets. For example, targets may be selected so that no two targets are separated within the genome by at least 1,000, at least 3,000, at least 5,000, at least 10,000, at least 30,000, at least 50,000, at least 100,000, at least 300,000, at least 500,000, at least 1,000,000, at least 3,000,000, at least 5,000,000, at least 10,000,000 nucleotides, etc. In addition, in certain embodiments, targets may be selected such that no two targets are separated within the genome by more than 10,000,000 nucleotides, more than 5,000,000 nucleotides, more than 3,000,000 nucleotides, more than 1,000,000 nucleotides, more than 500,000 nucleotides, more than 300,000 nucleotides, more than 100,000 nucleotides, more than 50,000 nucleotides, more than 30,000 nucleotides, or more than 10,000 nucleotides. Any combination of these is also possible in certain embodiments, for example, targets may be separated by 30,000 to 100,000, 3,000,000 to 5,000,000, 500,000 to 1,000,000 nucleotides, etc. Such targets may be selected, for example, selectively, randomly, etc., as discussed herein.

[0063] In addition, in one set of embodiments, RNA in the cell nucleus, such as nascent RNA, may be studied instead of, or in addition to, the nuclear DNA described above. In some cases, for example, nuclear RNA may be determined, and then nuclear DNA may be determined.

[0064] In some cases, after RNA determination, the RNA may be removed or inactivated before DNA determination. This can facilitate the separation of DNA and RNA determinations, for example, by eliminating RNA signals that could complicate DNA determination. Examples of methods for removing or inactivating RNA include the use of RNases such as endoribonucleases or exoribonucleases. Specific non-limiting examples include RNase A, RNase H, RNase III, RNase L, RNase P, RNase PhyM, RNase T1, RNase T2, RNase U2, RNase V, PNPase, RNase PH, RNase R, RNase D, RNase T, oligoribonuclease, exoribonuclease I, exoribonuclease II, and others.

[0065] However, it should be understood that in other embodiments, DNA may be determined before RNA, and / or both may be determined simultaneously. For example, DNA may be removed or inactivated after determination using a DNase such as an exodeoxyribonuclease or endodeoxyribonuclease. Examples include, but are not limited to, deoxyribonuclease I (DNase I), deoxyribonuclease II (DNase II), DNase IV, UvrABC endonuclease, and others. Another example is that DNA may be degraded by exposure to a restriction endonuclease. Many such nucleases are commercially available.

[0066] RNA in the nucleus can be determined using any suitable technique, and may be determined using the same or different techniques used to determine DNA in the nucleus. In one embodiment, RNA can be determined using MERFISH. See, for example, International Patent Application Publication WO2016 / 018960, entitled "Systems and Methods for Determining Nucleic Acids"; and International Patent Application Publication WO2016 / 018963, entitled "Probe Library Construction," each of which is incorporated herein by reference in its entirety. In another embodiment, RNA can be determined using multiple nucleic acid probes, as discussed herein, for example. For example, in some embodiments, RNA can be determined using nucleic acids such as encoding nucleic acid probes, primary amplified nucleic acids, and secondary amplified nucleic acids, as described below. In some cases, the nucleic acid probes may define error detection and / or error correction codes, as discussed herein, for example.

[0067] In some embodiments, DNA, such as genomic DNA, may be determined using nucleic acids, such as coding nucleic acid probes, primary amplification nucleic acids, and secondary amplification nucleic acids, as described herein. In some cases, the nucleic acid probe may define error detection and / or error correction codes, as discussed herein, for example.

[0068] In addition, in one set of embodiments, proteins within the nucleus of a cell may be studied using techniques such as those described above, in addition to nucleic acids present in the nucleus. Examples of proteins that may be studied include, but are not limited to, nuclear speckles, nucleoli, nuclear laminas, or histone proteins. Speckles are structures that are enriched with premessenger RNA splicing factors and may be located in the interchromatin region of the nucleocytoplasm of mammalian cells. Nucleoli are structures that form around frequently transcribed genomic loci that encode ribosomal RNA (rRNA) and may be enriched with rRNA and the transcription machinery that associates with it. Each lamina is a protein structure associated with the inner nuclear membrane and may be enriched with intermediate filaments (lamins) and transcriptionally inactive chromatin. Histones are proteins used in the nucleus to wrap or fold DNA into smaller, more compact complexes and form chromatin.

[0069] A variety of methods can be used to determine proteins. For example, in one set of embodiments, an immunofluorescence assay may be used. In another set of embodiments, a “sandwich assay” may be used, in which a primary antibody capable of specifically binding to a nucleoprotein is applied, followed by a secondary antibody capable of specifically binding to the primary antibody, the secondary antibody containing a signal-generating entity such as a fluorescent entity, or an oligonucleotide that can be detected using a complementary oligonucleotide, for example, linked to a fluorescent entity. Such protein determination may be performed on the same sample or the same nucleus as described above, for example, before or after the determination of nucleic acids in the nucleus. Thus, in some cases, proteins and nucleic acids in the nucleus of a cell can be determined, for example, spatially.

[0070] As noted, in various embodiments, such as those described herein, various nucleic acid probes can be used to determine one or more targets in cells or other samples, for example, within the nucleus of a cell. Probes may include nucleic acids (or entities that can hybridize to nucleic acids, for example, specifically), such as DNA, RNA, LNA (locked nucleic acid), PNA (peptide nucleic acid), and / or combinations thereof. Examples of nucleic acid probes include, but are not limited to, International Patent Application Publication WO2016 / 018960, entitled "Systems and Methods for Determining Nucleic Acids" and International Patent Application Publication WO2016 / 018963, entitled "Probe Library Construction," each incorporated herein in whole by reference. In some cases, as discussed herein, for example, further components may also be present within the nucleic acid probe. In addition, nucleic acid probes can be introduced into cells, for example, the nucleus of a cell, using any suitable method.

[0071] For example, in some embodiments, cells are immobilized before introducing nucleic acid probes to preserve the location of nucleic acids or other targets, for example, within the cell, for example, in the nucleus. Techniques for immobilizing cells are known to those skilled in the art. As an example without limitation, cells may be immobilized using chemicals such as formaldehyde, paraformaldehyde, glutaraldehyde, ethanol, methanol, acetone, and acetic acid. In one embodiment, cells may be immobilized using Hepes-glutamate buffer-mediated organic solvent (HOPE).

[0072] In addition, in some cases, cells (or other samples) may be fixed more than once, for example, during a relatively long experiment. For example, a sample may be re-fixed after the start of the experiment, for example, after exposing the cell nuclei to multiple nucleic acid probes. For example, cells or other samples may be fixed at least once every 7 days, at least once every 4 days, at least once every 2 days, at least once daily, at least once every 12 hours, at least once every 6 hours, at least once every 3 hours, etc. In some cases, this may be done between different rounds, for example, exposure to nucleic acid probes (e.g., primary or secondary nucleic acid probes). In some cases, a sample may be fixed a certain number of times, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 times, or any other appropriate number of times. When multiple fixations are performed, the same or different fixation techniques may be used independently.

[0073] Nucleic acid probes can be introduced into cells (or other samples) using any suitable method. In some cases, cells may be thoroughly permeabilized so that the nucleic acid probe can be introduced into the cells by a fluid containing the nucleic acid probe in the vicinity of the cells. In some cases, cells may be thoroughly permeabilized as part of a fixation step, and in other embodiments, cells may be permeabilized by exposure to certain chemicals such as ethanol, methanol, or Triton. In addition, in some embodiments, techniques such as electroporation or microinjection can be used to introduce nucleic acid probes into cells or other samples.

[0074] Therefore, certain embodiments generally concern nucleic acid probes that are introduced into cells (or other samples). Depending on the application, probes may include any of a variety of entities that can typically hybridize with nucleic acids such as DNA, RNA, LNA, and PNA via Watson-Crick base pairing. Nucleic acid probes typically contain a target sequence capable of binding to at least a portion of a target, e.g., a target nucleic acid. In some cases, the binding may be specific (e.g., via complementary binding). Once introduced into cells or other systems, the target sequence may be capable of binding to a specific target (e.g., nascent RNA, genomic DNA, mRNA, or other nucleic acids discussed herein). Nucleic acid probes may also contain one or more readout sequences, as discussed below.

[0075] Depending on the case, more than one type of nucleic acid probe may be applied to the sample, for example, sequentially, or simultaneously. For example, there may be at least two, at least five, at least ten, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, at least 30,000, at least 100,000, at least 300,000, at least 300,000, and at least 1,000,000 identifiable nucleic acid probes that can be applied to the sample, for example, to cells to target the nucleus. Depending on the case, nucleic acid probes may be added sequentially. However, depending on the case, more than one nucleic acid probe may be added simultaneously.

[0076] A nucleic acid probe may contain one or more target sequences that can be positioned at any location within the nucleic acid probe. The target sequences may contain regions substantially complementary to a target that may be present in the nucleus, such as a portion of the target nucleic acid. For example, the portion may be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary to, for example, specific binding. Typically, complementarity is determined based on Watson-Crick type nucleotide base pairing.

[0077] Depending on the case, the target sequence may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. Depending on the case, the target sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the target sequence may have lengths between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0078] In some embodiments, nucleic acid targets or codes associated with nucleic acid targets within cells or the nucleus of cells may be selected to be spatially separated, for example, in genomic space or physical space, based on prior knowledge of chromatin structure, such as the target being composed of small, clustered regions of chromosomes in each round of imaging.

[0079] In addition, in some cases, the target sequence of a nucleic acid probe may be determined by referring to a target that is thought to be present in a cell or other sample, for example, in the nucleus of a cell. For example, the target nucleic acid for a protein (e.g., nuclear speckle, nuclear lamina, etc.) may be determined using the protein sequence, for example, by determining the nucleic acid expressed to form the protein. In some cases, for example, only a portion of the nucleic acid encoding a protein having the length discussed above may be used.

[0080] In a particular embodiment, more than one target sequence may be used to identify a specific target. For example, multiple probes may be used that can sequentially and / or simultaneously bind to the same or different regions of the same target, or that can sequentially and / or simultaneously hybridize with it. Hybridization typically refers to the annealing process in which complementary single-stranded nucleic acids associate via Watson-Crick type nucleotide base pairings (e.g., guanine-cytosine and adenine-thymine hydrogen bonds) to form a double-stranded nucleic acid.

[0081] In some embodiments, the nucleic acid probe may also include one or more “readout” sequences. The readout sequences may be used to identify the nucleic acid probe through association with signal-generating entities, for example, as discussed below. In some embodiments, the nucleic acid probe may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 24 or more, 32 or more, 40 or more, 48 or more, 50 or more, 64 or more, 75 or more, 100 or more, 128 or more readout sequences. The readout sequences may be placed anywhere within the nucleic acid probe. If there is more than one readout sequence, the readout sequences may be placed adjacent to each other and / or interrupted by other sequences.

[0082] A readout sequence can be of any length. If more than one readout sequence is used, the readout sequences may be the same or different independently. For example, a readout sequence may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides long. Depending on the case, the readout sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the readout sequence may have lengths between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0083] In some embodiments, the readout sequence may be arbitrary or random. In certain cases, the readout sequence is selected to reduce or minimize homology with other components of the cell or other sample, for example, so that the readout sequence itself does not bind or hybridize with other nucleic acids that are likely to be present in the cell or other sample. In some cases, the homology may be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1%. In some cases, homology may be less than 20 base pairs, less than 18 base pairs, less than 15 base pairs, less than 14 base pairs, less than 13 base pairs, less than 12 base pairs, less than 11 base pairs, or less than 10 base pairs. In some cases, such base pairs are continuous.

[0084] In addition, in some embodiments, some or all of the readout sequences may be selected so that they do not exhibit specific binding to each other and / or to the genome or other nucleic acids that are thought to be present in the sample. For example, a population of readout sequences may be "blasted" or tested for specific binding or complementarity. In some cases, the readout sequences may not exhibit specific binding to each other, and / or therefore, none of the readout sequences in the population of readout sequences have complementarity of more than 5, 6, 7, 8, 9, 10 nucleotides, etc., to another readout sequence in the population of readout sequences.

[0085] In one set of embodiments, the collection of nucleic acid probes may contain a certain number of readout sequences, which may be the same as the number of nucleic acid targets to be determined in the sample, for example, each unique readout sequence corresponds to a unique target. In another set of embodiments, the collection of nucleic acid probes may contain a certain number of readout sequences, which may be less than the number of nucleic acid targets to be determined in the sample. Those skilled in the art will know that when there is one signal-generating entity and n readout sequences, generally, 2 nIt will be noted that -1 different nucleic acid targets can be uniquely identified. However, not all possible combinations need to be used. For example, a nucleic acid probe population may target 12 different nucleic acid targets but may contain no more than 8 readout sequences. As another example, a nucleic acid probe population may target 140 different nucleic acid targets but may contain no more than 16 readout sequences. Different nucleic acid targets can be individually identified by using different combinations of readout sequences within each probe. For example, a population of nucleic acid probes may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, etc. or more readout sequences. In some cases, each nucleic acid probe population may contain the same number of readout sequences, but in other cases, different numbers of readout sequences may be present on different probes.

[0086] As a non-limiting example, a first nucleic acid probe may contain a first target sequence, a first readout sequence, and a second readout sequence, while a second, different nucleic acid probe may contain a second target sequence, the same first readout sequence, but a third readout sequence instead of a second readout sequence. Such probes can thus be identified by determining a variety of readout sequences present at or associated with a given probe or position, as discussed herein. For example, probes may be sequentially identified and encoded using “codewords,” as discussed below. The codewords may also be subjected to error detection and / or correction.

[0087] As another non-limiting example, a first population of nucleic acid probes may contain a first target sequence, a first readout sequence, and a second readout sequence, while a second different population of nucleic acid probes may contain a second target sequence, the same first readout sequence, but with a third readout sequence instead of the second readout sequence. Such probes can be distinguished by determining the diverse readout sequences present at or associated with a given probe or position, as discussed herein. For example, populations of probes can be sequentially identified and encoded using “codewords,” as discussed below. The codewords may also be used for error detection and / or correction, as appropriate.

[0088] In addition, in certain embodiments, nucleic acid probe populations may be constructed using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G" or all "C" from the probe population. In certain embodiments, sequences lacking "G" or "C" may form little secondary structure and contribute to more homogeneous and rapid hybridization. Therefore, depending on the case, nucleic acid probes may contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.

[0089] In one embodiment, a readout sequence on a nucleic acid probe may be able to bind (e.g., specifically) to a recognition sequence on the corresponding primary amplified nucleic acid. Therefore, if the nucleic acid probe recognizes a target in a biological sample, e.g., a DNA or RNA target, the primary amplified nucleic acid may also associate with the target via the nucleic acid probe through interaction between the readout sequence of the nucleic acid probe and the corresponding recognition sequence on the primary amplified nucleic acid, e.g., complementary binding. For example, the recognition sequence may be able to recognize a target readout sequence but substantially not recognize or bind to other non-target readout sequences. The primary amplified nucleic acid may also, depending on the application, include any of a variety of entities that can hybridize with nucleic acids, e.g., DNA, RNA, LNA, and / or PNA. For example, such entities may form part or all of the recognition sequence.

[0090] In some cases, the recognition sequence may be substantially complementary to the target readout sequence. In some cases, the sequences may be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary. Typically, complementarity is determined based on Watson-Crick type nucleotide base pairing. The structure of the target readout sequence may include structures already described.

[0091] Depending on the case, the recognition sequence may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. Depending on the case, the recognition sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the recognition sequence may have lengths between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0092] In some embodiments, the primary amplified nucleic acid may also include one or more readout sequences that can bind to the secondary amplified nucleic acid, as discussed below. For example, the primary amplified nucleic acid may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, 128 or more readout sequences. The readout sequences may be located anywhere within the primary amplified nucleic acid. If there is more than one readout sequence, the readout sequences may be located adjacent to each other and / or interrupted by other sequences. In one embodiment, the primary amplified nucleic acid includes a recognition sequence at a first end and includes multiple readout sequences at a second end.

[0093] Depending on the case, the readout sequence in the primary amplified nucleic acid may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. Depending on the case, the readout sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the readout sequence may have lengths between 10 and 20 nucleotides, between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0094] Any number of readout sequences can exist within a primary amplified nucleic acid. For example, a primary amplified nucleic acid may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more readout sequences. If more than one readout sequence is present in a primary amplified nucleic acid, the readout sequences may be the same or different. In some cases, for example, all readout sequences may be identical.

[0095] In some embodiments, the population of primary amplified nucleic acids may be made using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G" or all "C" from the nucleic acid population. In certain embodiments, sequences lacking "G" or "C" may form little secondary structure and contribute to more homogeneous and rapid hybridization. Therefore, depending on the case, the primary amplified nucleic acid may contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.

[0096] In some cases, more than one primary amplified nucleic acid may be applied to the sample, for example, sequentially or simultaneously. For example, at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 identifiable primary amplified nucleic acids may be applied to the sample. In some cases, the primary amplified nucleic acids may be added sequentially. However, in some cases, more than one primary amplified nucleic acid may be added simultaneously.

[0097] In one set of embodiments, a readout sequence on a primary amplified nucleic acid may be capable of binding (e.g., specifically) to a recognition sequence on a corresponding secondary amplified nucleic acid. Therefore, if a nucleic acid probe recognizes a target in a biological sample, e.g., a DNA or RNA target, the secondary amplified nucleic acid may also associate with the target via the primary amplified nucleic acid through interaction between the readout sequence of the primary amplified nucleic acid and the recognition sequence on the corresponding secondary amplified nucleic acid, e.g., complementary binding. For example, a recognition sequence on a secondary amplified nucleic acid may recognize a readout sequence on the primary amplified nucleic acid, but substantially not recognize or substantially bind to other non-target readout sequences. The secondary amplified nucleic acid may also include, depending on the application, any of a variety of entities capable of hybridizing with nucleic acids, e.g., DNA, RNA, LNA, and / or PNA. For example, such entities may form part or all of the recognition sequence.

[0098] In some cases, the recognition sequence on the secondary amplified nucleic acid may be substantially complementary to the readout sequence on the primary amplified nucleic acid. In some cases, the sequences may be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary.

[0099] Depending on the case, the recognition sequence on the secondary amplified nucleic acid may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. Depending on the case, the recognition sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the recognition sequence may have lengths between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0100] In some embodiments, the secondary amplified nucleic acid may include a signal-generating entity and / or one or more readout sequences capable of binding to the signal-generating entity, as discussed herein. For example, the secondary amplified nucleic acid may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, 128 or more readout sequences capable of binding to the signal-generating entity. The readout sequences may be located anywhere within the secondary amplified nucleic acid. If there is more than one readout sequence, the readout sequences may be located adjacent to each other and / or interrupted by other sequences. In one embodiment, the secondary amplified nucleic acid includes a recognition sequence at a first end and includes a plurality of readout sequences at a second end. This structure may also be the same as or different from the structure of the primary amplified nucleic acid.

[0101] Depending on the case, the readout sequence in the secondary amplified nucleic acid may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. Depending on the case, the readout sequence may be of a length not exceeding 500, 450, 400, 350, 300, 250, 200, 175, 150, 125, 100, 75, 60, 65, 60, 55, 50, 45, 40, 35, 30, 20, or 10 nucleotides. Any combination of these is also possible; for example, the readout sequence in a secondary amplified nucleic acid may have lengths between 10 and 20 nucleotides, between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, or between 10 and 300 nucleotides.

[0102] Any number of readout sequences can exist within a secondary amplified nucleic acid. For example, a secondary amplified nucleic acid may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more readout sequences. If more than one readout sequence is present in a secondary amplified nucleic acid, the readout sequences may be the same or different. In some cases, for example, all the readout sequences may be identical. In addition, the primary amplified nucleic acid and the secondary amplified nucleic acid may independently contain the same or different number of readout sequences.

[0103] In certain embodiments, the population of secondary amplified nucleic acids may be made using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G" or all "C" from the nucleic acid population. In certain embodiments, sequences lacking "G" or "C" may form little secondary structure and contribute to more homogeneous and rapid hybridization. Therefore, depending on the case, the secondary amplified nucleic acids may contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.

[0104] In some cases, more than one type of secondary amplified nucleic acid may be applied to the sample, for example, sequentially or simultaneously. For example, at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 identifiable secondary amplified nucleic acids may be applied to the sample. In some cases, the secondary amplified nucleic acids may be added sequentially. However, in some cases, more than one secondary amplified nucleic acid may be added simultaneously.

[0105] In addition, in certain embodiments, this pattern is repeated before the signal-generating entity, for example, by tertiary amplification nucleic acids, quaternary amplification nucleic acids, etc., as discussed above. Thus, the signal-generating entity can be bound to the final amplification nucleic acid. For example, in non-limiting examples, the target may be bound with an encoding nucleic acid probe, to which a primary amplification nucleic acid is bound, to which a secondary amplification nucleic acid is bound, to which a tertiary amplification nucleic acid is bound, and to which the signal-generating entity is bound; or the target may be bound with an encoding nucleic acid probe, to which a primary amplification nucleic acid is bound, to which a secondary amplification nucleic acid is bound, to which a tertiary amplification nucleic acid is bound, to which a quaternary amplification nucleic acid is bound, and to which the signal-generating entity is bound. Therefore, in all embodiments, the final amplification nucleic acid does not necessarily have to be a secondary amplification nucleic acid.

[0106] A non-limiting example of such a system is illustrated in Figure 5. Figures 5A to 5E illustrate the creation of a saturable system. Figure 5A shows an example of an encoding nucleic acid probe, where the encoding nucleic acid probe 15 is bound to the target RNA. Figure 5B shows the primary amplification nucleic acid used according to a particular embodiment. Figure 5C shows a secondary amplification nucleic acid that can bind to the primary amplification nucleic acid. Figure 5D shows multiple signal-generating entities bound to the readout sequence of the secondary amplification nucleic acid. Figure 5E shows that when amplification is not applied, the nucleic acid probe may be exposed to a suitable secondary nucleic acid probe containing a signal-generating entity.

[0107] Furthermore, in certain specific cases, other components may also be present within the nucleic acid probe or amplified nucleic acid. For example, in one set of embodiments, one or more primer sequences may be present, for example, to facilitate amplification by enzyme. Those skilled in the art will be familiar with primer sequences suitable for applications such as amplification (e.g., using PCR or other suitable techniques). Many such primer sequences are commercially available. Other examples of sequences that may be present within a primary nucleic acid probe or encoding nucleic acid probe include, but are not limited to, promoter sequences, operons, identification sequences, nonsense sequences, and others.

[0108] Typically, a primer is a single-stranded or partially double-stranded nucleic acid (e.g., DNA) used as a starting point for nucleic acid synthesis, allowing polymerase enzymes such as nucleic acid polymerase to extend the primer and replicate the complementary strand. The primer is complementary to and hybridizes with the target nucleic acid (e.g., is complementary to and designed to hybridize with it). In some embodiments, the primer is a synthetic primer. In some embodiments, the primer is a primer that does not exist in nature. Primers typically have a length of 10 to 50 nucleotides. For example, primers may have lengths of 10 to 40, 10 to 30, 10 to 20, 25 to 50, 15 to 40, 15 to 30, 20 to 50, 20 to 40, or 20 to 30 nucleotides. In some embodiments, the primer has a length of 18 to 24 nucleotides.

[0109] In some embodiments, as previously discussed, certain embodiments use a code space that encodes a variety of binding events and, as appropriate, uses error detection and / or correction to determine the binding of nucleic acid probes to their targets. In some cases, a group of nucleic acid probes may contain certain “readout sequences” that can bind to a particular amplified nucleic acid, as discussed above, and the position of the nucleic acid probe or target can be determined in a sample using a signal-generating entity that associates with the amplified nucleic acid, for example, within a certain code space, as discussed herein. See also International Patent Application Publications WO2016 / 018960 and WO2016 / 018963, respectively, which are incorporated herein by reference in their entirety. In some cases, the group of readout sequences in the nucleic acid probes can be combined in a variety of combinations, as discussed herein, for example, so that a relatively large number of different nucleic acid probes can be determined using a relatively small number of readout sequences.

[0110] Therefore, in some cases, a population of nucleic acid probes may each contain a certain number of readout sequences, some of which are shared among different nucleic acid probes so that the entire population of nucleic acid probes may contain a certain number of readout sequences. A population of nucleic acid probes can have any appropriate number of readout sequences. For example, a population of nucleic acid probes may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, etc. In some embodiments, more than 20 readout sequences may also be possible. In addition, in some cases, a population of nucleic acid probes may have in total one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, twenty or more, twenty-four or more, thirty-two or more, forty or more, fifty or more, sixteen or more, one hundred or more, one hundred-two or more, thirty-two or more, forty or more, fifty or more, sixteen or more, one hundred-two In addition, in some embodiments, the population of nucleic acid probes may have readout sequences that are not greater than 100, not greater than 80, not greater than 64, not greater than 60, not greater than 50, not greater than 40, not greater than 32, not greater than 24, not greater than 20, not greater than 16, not greater than 15, not greater than 14, not greater than 13, not greater than 12, not greater than 11, not greater than 10, not greater than 9, not greater than 8, not greater than 7, not greater than 6, not greater than 5, not greater than 4, not greater than 3, or not greater than 2.Furthermore, any combination of these is possible; for example, a population of nucleic acid probes may contain a total of 10 to 15 readout sequences.

[0111] As a non-limiting example of a method for combinatorially identifying a relatively large number of nucleic acid probes from a relatively small number of readout sequences contained within nucleic acid probes, in a population or group of six different types of nucleic acid probes or six different groups of nucleic acid probes (e.g., each probe group binds to a nucleic acid target), where each type or group of nucleic acid probes contains one or more readout sequences, the total number of readout sequences in the population may not exceed four. For the sake of clarity, four readout sequences are used in this example, but it should be understood that in other embodiments, a large number of nucleic acid probes can be realized using, for example, 5, 8, 10, 16, 32, or more readout sequences, or any other suitable number of readout sequences as described herein by application. For example, if each nucleic acid probe or each group of nucleic acid probes contains two different readout sequences, then by using four such read sequences (A, B, C, and D), up to six probes or six groups of probes can be identified separately. In this example, it should be noted that the order of the readout sequences in a nucleic acid probe or group of nucleic acid probes is not essential; that is, "AB" and "BA" can be treated as synonymous (although in other embodiments, the order of the read sequences may be essential, and "AB" and "BA" may not necessarily be synonymous). Similarly, if five readout sequences are used in a group of nucleic acid probes (A, B, C, D, and E), up to 10 probes or 10 groups of probes can be identified separately (e.g., AB, AC, AD, AE, BC, BD, BE, CD, CE, DE). For example, for k readout sequences in a group having n readout sequences in each probe or group of groups, assuming that the order of the readout sequences is not essential, at most...

[0112] [ka] Individual different probes may be produced; for all probes or groups of probes do not need to have the same number of readout sequences, and not all combinations of readout sequences need to be used in every embodiment, and more or fewer different probes than this number may be used in a particular embodiment, as those skilled in the art will understand. In addition, it should be understood that the number of readout sequences in each probe or group of probes does not need to be the same in some embodiments. For example, some probes or groups of probes may contain two read sequences, while other probes or groups of probes may contain three read sequences. In some embodiments, each group of probes binds to a nucleic acid target.

[0113] In some embodiments, the readout sequences and / or binding patterns of nucleic acid probes in a sample may be used to define error detection and / or error correction codes, for example, to reduce or prevent nucleic acid misidentification or errors. For example, if binding is indicated (determined, for example, using a signal-generating entity), the location may be identified by "1"; conversely, if binding is not indicated, the location may be identified by "0" (or, in some cases, the reverse). Then, using multiple rounds of binding determination, for example, with various readout probes complementary to the readout sequence, a "codeword" for, for example, its spatial location can be generated. In some embodiments, the codeword may be used for error detection and / or correction. For example, the codeword may be configured such that if no match is found for a given set of readout sequences or binding patterns of nucleic acid probes, the match may be identified as an error, and, as appropriate, error correction may be applied to the sequence to determine a precise target for the nucleic acid probe. In some cases, a codeword may have fewer "characters" or positions than the total number of nucleic acids it encodes, for example, where each codeword encodes a different nucleic acid.

[0114] Such error detection and / or error correction codes can take various forms. Various such codes, such as Goley codes or Hamming codes, have already been developed in other contexts, such as the telecommunications industry. In one set of embodiments, the readout sequence or binding pattern of the nucleic acid probe is assigned in such a way that not all possible combinations are assigned.

[0115] For example, if four readout sequences are possible and a nucleic acid probe or group of nucleic acid probes contains two readout sequences, then up to six nucleic acid probes or six groups of nucleic acid probes (e.g., each group of nucleic acid probes binds to a nucleic acid target) can be identified; however, the number of nucleic acid probes or groups of nucleic acid probes used may be less than six. Similarly, for k readout sequences in a population having n readout sequences in each nucleic acid probe or group of nucleic acid probes,

[0116] [ka] Individual different probes or different groups of probes may be produced, but the number of nucleic acid probes or groups of nucleic acid probes used is

[0117] [ka] This can be any number, more or less than 1. In addition, these may be assigned randomly or in a specific way that increases the ability to detect and / or correct errors.

[0118] As another example, when multiple rounds of nucleic acid probes are used (e.g., multiple rounds of readout probes that can therefore bind to readout sequences on a primary probe or coding probe), the number of rounds can be arbitrarily chosen. Within each round, if each target can produce two possible outcomes, such as detection or non-detection, then for n rounds of probes, at most 2 n While several different targets may be possible, the number of targets actually used is 2 n It can be any number less than 2. For example, if within each round each target can result in more than 2 possible outcomes, such as detection in different color channels, then for n rounds of probes, 2 n More than one (for example, 3 n , 4 n There may be different targets (…) in number. In some cases, the actual number of targets used may be any number less than this. In addition, they may be assigned randomly or in a specific way that increases the ability to detect and / or correct errors.

[0119] Codewords can be used to define diverse coding spaces. Each nucleic acid target is associated with a codeword. For example, in one set of embodiments, codewords may be assigned within a coding space such that assignments are separated by a Hamming distance, which measures the number of inaccurate "reads" in a given pattern that cause a codeword or associated nucleic acid target to be misinterpreted as a different, legitimate codeword or nucleic acid target. In a particular case, the Hamming distance may be at least 2, at least 3, at least 4, at least 5, at least 6, and so on. In addition, in one set of embodiments, assignments may be formed as Hamming codes, e.g., Hamming(7,4) code, Hamming(15,11) code, Hamming(31,26) code, Hamming(63,57) code, Hamming(127,120) code, and so on. In another set of embodiments, the assignment may form SECDED codes, such as SECDED(8,4) codes, SECDED(16,4) codes, SECDED(16,11) codes, SECDED(22,16) codes, SECDED(39,32) codes, SECDED(72,64) codes, and so on. In yet another set of embodiments, the assignment may form extended binary Goley codes, full binary Goley codes, or ternary Goley codes. In yet another set of embodiments, the assignment may represent a subset of possible values ​​taken from any of the codes described above.

[0120] For example, an error correction code may be formed by encoding a target using only binary words containing a fixed or constant number of "1" bits (or "0" bits). For example, the code space may contain only one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, tenth, eleventh, twelveth, thirteenth, fourteenth, fifteenth, sixteenth, and so on, e.g., all of the codes have the same number of "1" bits or "0" bits. In another set of embodiments, the assignment may represent a subset of possible values ​​taken from any of the codes described above for the purpose of addressing asymmetric readout errors. For example, a code in which the number of "1" bits may be fixed for all binary words used may eliminate bias measurements of a word by different numbers of "1"s when the proportion of "0" bits measured as "1" or the proportion of "1" bits measured as "0" are different.

[0121] Therefore, in some embodiments, once a codeword is determined (as discussed herein, for example), it can be compared to a valid nucleic acid codeword. If a match is found, the nucleic acid target can be identified or determined. If no match is found, an error in reading the codeword can be identified. In some cases, error correction can also be applied to determine a valid codeword, thereby resulting in the proper identification of the nucleic acid target. In some cases, the codewords can be selected such that, assuming only one error exists, only one possible and valid codeword is available, thereby allowing only one proper identification of the nucleic acid target. In some cases, this can also be generalized to larger codeword intervals or Hamming distances; for example, the codewords can be selected such that, if two, three, or four errors (or, in some cases, more errors) exist, only one possible and valid codeword is available, thereby allowing only one proper identification of the nucleic acid target.

[0122] Error correction codes may be binary error correction codes, or they may be based on other numbering systems, such as ternary or quaternary error correction codes. For example, in one set of embodiments, more than one type of signal-generating entity may be used and assigned to different numbers in the error correction code. Thus, as an unrestricted example, a first signal-generating entity (or, if applicable, more than one signal-generating entity) may be assigned as "1", a second signal-generating entity (or, if applicable, more than one signal-generating entity) may be assigned as "2" (where "0" indicates the absence of a signal-generating entity), and the codewords are distributed to define a ternary error correction code. Similarly, a third signal-generating entity may be assigned as "3" to create a quaternary error correction code, and so on.

[0123] In one set of embodiments, nucleic acid targets in a sample are each assigned a codeword. For example, these codewords may be selected from one of the coding spaces described herein. In some cases, the codewords form error detection and / or error correction codes. In some cases, the sample may be subjected to hybridization into a collection of primary nucleic acid probes or coding nucleic acid probes. Some or all of the primary or coding probes may contain target sequences that can bind to one of the nucleic acid targets and / or may contain one or more readout sequences. The readout sequences in the collection of primary or coding probes that bind to each nucleic acid target may form a unique codeword corresponding to the codeword assigned to the nucleic acid target. The sample is then subjected to one or more rounds of hybridization with the readout probes. The readout probes may be able to bind to readout sequences and / or associate with signal-generating entities. The collection of readout sequences may associate with nucleic acid targets, and therefore, the codewords assigned to the nucleic acid targets can be identified, for example, through the binding of the readout probes.

[0124] In some cases, multicolor imaging can be used in each round to enable simultaneous imaging and determination of multiple readout probes associating with different signaling entities. In some cases, the location of nucleic acid targets is determined. In some cases, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 nucleic acid targets are determined in this manner. In some cases, the target nucleic acid is a genomic locus. In some cases, the target nucleic acid is a genomic locus and / or a nascent RNA transcript. In some cases, the location of a genomic locus is used to determine the three-dimensional structure of chromatin within a cell or the three-dimensional structure of a genome. In some cases, primary amplification nucleic acids and / or secondary amplification nucleic acids and / or tertiary amplification nucleic acids and / or quaternary amplification nucleic acids are used to amplify the signal from each readout sequence. In some cases, adapters are used as described below.

[0125] In one embodiment, multiple adapters can be used to facilitate the detection of targets in a sample. Such adapters may be useful in allowing, for example, a relatively small number of identifiable signal-generating entities to be used, while still allowing a relatively large number of targets to be determined in the sample. For example, while using a small number of signal-generating entities, e.g., not more than 20, not more than 15, not more than 10, not more than 5, not more than 4, not more than 3, or not more than 2, while using at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 20, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, and at least A target can be determined to be at least 128, at least 140, at least 255, at least 256, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000.

[0126] In one set of embodiments, multiple adapters may be used. An adapter may comprise a first portion substantially complementary to one or more readout sequences on a nucleic acid probe (e.g., a primary nucleic acid probe) and a second portion containing one or more identification sequences. Thus, the adapter sequence can bind to a specific nucleic acid probe capable of binding to a target in the sample. The identification sequences are then available for binding via a readout probe or secondary nucleic acid probe, such as the probes discussed herein. Therefore, in some cases, the adapter may be located between a primary nucleic acid probe and a secondary nucleic acid probe. A non-limiting example of this is shown in Figure 24A.

[0127] In some cases, the adapter may be selected to allow the use of a relatively small number of signal-generating entities, as shown above. For example, the identification sequence may act as a readout sequence to which a secondary nucleic acid probe can bind. In one round of detection, a relatively small number of secondary nucleic acid probes may be used, for example, containing a sequence substantially complementary to one of the signal-generating entities and the identification sequence, and the signal-generating entities are determined, for example, as discussed herein. The secondary nucleic acid probes may then be removed and / or deactivated before the next detection round, for example, as described herein. Subsequent rounds may use the same or different signal-generating entities on secondary nucleic acid probes containing sequences substantially complementary to different identification sequences, for example.

[0128] In addition, in some embodiments, adapters used in previous rounds may be deactivated in some manner to reduce contamination or "crosstalk." For example, a blocking nucleic acid probe containing a sequence substantially complementary to the previously identified sequence may be added so that the blocking nucleic acid probe can bind to the previous adapter, because they are generally undetectable without the presence of a signal-generating entity. Thus, in subsequent detection rounds, the signal from previous rounds can be minimized.

[0129] Therefore, in some cases, a relatively large number of identification sequences can be determined using a relatively small number of signal-generating entities. For example, using no more than 20, no more than 15, no more than 10, no more than 5, no more than 4, no more than 3, or no more than 2 signal-generating entities, at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 20, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, and at least 14 Identification sequences such as 0, at least 255, at least 256, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 may be determined.

[0130] An identification sequence can be of any length. If more than one identification sequence is used, the identification sequences can have the same or different lengths independently. For example, an identification sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides long. In some cases, the identified sequence may be of a length not exceeding 500 nucleotides, not exceeding 450 nucleotides, not exceeding 400 nucleotides, not exceeding 350 nucleotides, not exceeding 300 nucleotides, not exceeding 250 nucleotides, not exceeding 200 nucleotides, not exceeding 175 nucleotides, not exceeding 150 nucleotides, not exceeding 125 nucleotides, not exceeding 100 nucleotides, not exceeding 75 nucleotides, not exceeding 60 nucleotides, not exceeding 65 nucleotides, not exceeding 60 nucleotides, not exceeding 55 nucleotides, not exceeding 50 nucleotides, not exceeding 45 nucleotides, not exceeding 40 nucleotides, not exceeding 35 nucleotides, not exceeding 30 nucleotides, not exceeding 20 nucleotides, or not exceeding 10 nucleotides. Any combination of these is also possible, for example, the identified sequence may have lengths of 10-30 nucleotides, 20-40 nucleotides, 5-50 nucleotides, 10-200 nucleotides, or 25-35 nucleotides, 10-300 nucleotides, etc.

[0131] The identification sequence may be arbitrary or random in some embodiments. In certain cases, for example, the identification sequence is selected to reduce or minimize homology with other components of the cell or other sample, so that the identification sequence does not bind or hybridize with other nucleic acids that are likely to be present in the cell or other sample. In some cases, the homology may be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1%. In some cases, homology may be less than 20 base pairs, less than 18 base pairs, less than 15 base pairs, less than 14 base pairs, less than 13 base pairs, less than 12 base pairs, less than 11 base pairs, or less than 10 base pairs. In some cases, such base pairs are sequential.

[0132] In addition, in some embodiments, some or all of the identification sequences may be selected so that they do not exhibit specific binding to each other and / or to the genome or other nucleic acids, such as readout sequences, that are thought to be present in the sample. For example, a population of identification sequences may be "blasted" or tested for specific binding or complementarity. In some cases, the identification sequences may not exhibit specific binding to each other, and / or therefore, none of the identification sequences in the population of identification sequences have complementarity of more than 5, 6, 7, 8, 9, 10 nucleotides, etc., to other sequences in the population of identification sequences and / or the population of readout sequences.

[0133] In some embodiments, the sample is first subjected to hybridization into a population of primary nucleic acid probes or coding nucleic acid probes. One or more of the primary or coding probes may contain a target sequence that can bind to one of the nucleic acid targets, and may also contain one or more readout sequences. The sample is then subjected to multiple rounds of hybridization with adapter probes and readout probes. The adapter probes may contain sequences that can bind to the readout sequences, and may also contain one or more identification sequences. The readout probes may be able to bind to the identification sequences and may also associate with signal-generating entities. In some cases, multicolor imaging can be used in each round to enable simultaneous imaging and determination of multiple readout probes associating with different signal-generating entities.

[0134] As discussed herein, in certain embodiments, signal-generating entities are determined, for example, by imaging to identify nucleic acid probes and / or generate codewords. Examples of signal-generating entities include those discussed herein. In some cases, signal-generating entities in a sample may be determined, for example, spatially using various techniques. In some embodiments, signal-generating entities may be fluorescent, and techniques for determining fluorescence in a sample, such as fluorescence microscopy or confocal microscopy, may be used to spatially identify the location of the signal-generating entity within the cell. In some cases, the location of the entity in the sample may be determined in two or three dimensions. In addition, in some embodiments, more than one signal-generating entity may be determined simultaneously (e.g., signal-generating entities with different colors or emission) and / or sequentially.

[0135] In addition, in some embodiments, a confidence level for identified targets, such as nucleic acid targets, may be determined. For example, the confidence level may be determined using the ratio of the number of exact matches to the number of matches with one or more one-bit errors. In some cases, only matches with a confidence ratio above a certain value may be used. For example, in a particular embodiment, a match may be acceptable only if the confidence ratio for the match is greater than about 0.01, greater than about 0.03, greater than about 0.05, greater than about 0.1, greater than about 0.3, greater than about 0.5, greater than about 1, greater than about 3, greater than about 5, greater than about 10, greater than about 30, greater than about 50, greater than about 100, greater than about 300, greater than about 500, greater than about 1000, or any other appropriate value. In addition, in some embodiments, a match may be acceptable only if the confidence ratio for the target exceeds that of the internal standard or false-positive control by about 0.01, about 0.03, about 0.05, about 0.1, about 0.3, about 0.5, about 1, about 3, about 5, about 10, about 30, about 50, about 100, about 300, about 500, about 1000, or any other appropriate value.

[0136] In some embodiments, the spatial location of the entity (and therefore, any nucleic acid probes that may be associated with the entity) can be determined with relatively high resolution. For example, the location can be determined with spatial resolutions such as greater than about 100 micrometers, greater than about 30 micrometers, greater than about 10 micrometers, greater than about 3 micrometers, greater than about 1 micrometer, greater than about 800 nm, greater than about 600 nm, greater than about 500 nm, greater than about 400 nm, greater than about 300 nm, greater than about 200 nm, greater than about 100 nm, greater than about 90 nm, greater than about 80 nm, greater than about 70 nm, greater than about 60 nm, greater than about 50 nm, greater than about 40 nm, greater than about 30 nm, greater than about 20 nm, or greater than about 10 nm.

[0137] For example, there are various techniques that use fluorescence microscopy to optically determine or image the spatial location of an object. In some embodiments, more than one color may be used. In some cases, the spatial location may be determined at ultra-high resolution, or at a resolution beyond the wavelength or diffraction limit of light. Non-limiting examples include STORM (stochastic optical reconstruction microscopy), STED (stimulated emission depletion microscopy), NSOM (Near-field Scanning Optical Microscopy), 4Pi microscopy, SIM (Structured Illumination Microscopy), SMI (Spatially Modulated Illumination) microscopy, RESOLFT (Reversible Saturable Optically Linear Fluorescence Transition Microscopy), GSD (Ground State Depletion Microscopy), SSIM (Saturated Structured-Illumination Microscopy), SPDM (Spectral Precision Distance Microscopy), Photo-activated localization microscopy (PALM), Fluorescence-activated localization microscopy (FPALM), LIMON (3D Light Microscopical Nanosizing Microscopy), SOFI (Super-resolution optical fluctuation imaging), expansion microscopy, and others.For example, see U.S. Patent No. 7,838,302; U.S. Patent No. 8,564,792, issued on 22 October 2013, issued by Zhuang et al., entitled “Sub-Diffraction Limit Image Resolution and Other Imaging Techniques,” each of which is incorporated herein by reference in its entirety; or International Patent Application Publication No. WO2013 / 090360, published on 20 June 2013, entitled “High Resolution Dual-Objective Microscopy,” by Zhuang et al.

[0138] As an exemplary, non-limiting example, in one set of embodiments, a sample may be imaged using an oil immersion objective lens with a high numerical aperture and a magnification of 100x, and light collected on an electron-multiplier CCD camera. In another example, a sample may be imaged using an oil immersion objective lens with a high numerical aperture and a magnification of 40x, and light collected on a wide-field academic CMOS camera. In various non-limiting embodiments, different combinations of objective lenses and cameras may correspond to a single field of view not exceeding 1 × 1 micron, 10 × 10 micron, 40 × 40 micron, 80 × 80 micron, 120 × 120 micron, 240 × 240 micron, 340 × 340 micron, or 500 × 500 micron, etc. Similarly, in some embodiments, a single camera pixel may correspond to a sample area not exceeding 10×10 nm, 20×20 nm, 40×40 nm, 80×80 nm, 120×120 nm, 160×160 nm, 240×240 nm, or 300×300 nm, etc. In another example, the sample may be imaged by light collected by an sCMOS camera and an air lens with a low numerical aperture and a magnification of 10x. In a further embodiment, the sample may be optically separated by focusing light through one or more pinholes and illuminating it through diffraction-limited focused light generated by one or more scans by a scanning mirror or rotating disk. In another embodiment, the sample may also be illuminated through a thin plate of light generated by any one of several methods known to those skilled in the art.

[0139] In one embodiment, the sample may be irradiated with a single Gaussian-mode laser beam. In some embodiments, the irradiation profile may be flattened by passing these laser beams through a multimode fiber vibrated via piezoelectric or other mechanical means. In some embodiments, the irradiation profile may be flattened by passing a single-mode Gaussian beam through various refractive beam shapers, such as a π-shaper, or a series of stacked Powell lenses. In yet another set of embodiments, the Gaussian beam may also be passed through various different diffuse elements, such as ground glass or an engineering diffuser, which may be spun at high speed to remove residual laser speckle. In yet another embodiment, the laser irradiation may be passed through a series of small lens arrays to produce an overlapping irradiation image that approximates a planar irradiation field.

[0140] In some embodiments, the centroid of the spatial position of an entity can be determined. For example, the centroid of a signal-generating entity can be determined in an image or in a series of images using an image analysis algorithm known to those skilled in the art. The algorithm may be selected to determine non-overlapping single emitters and / or partially overlapping single emitters in a sample. Non-limiting examples of suitable techniques include maximum likelihood algorithms, least-squares algorithms, Bayesian algorithms, and compressed sensing algorithms. Combinations of these techniques may also be used.

[0141] In some embodiments, one or more signal-generating entities may be determined. For example, a signal-generating entity may be bound to a readout probe or a recognition entity on a secondary amplified nucleic acid (or other encoding amplified nucleic acid). Non-limiting examples of signal-generating entities include, for example, fluorescent entities (fluorophores) or phosphorescent entities, as discussed herein. The signal-generating entities may then be determined, for example, to determine a nucleic acid probe or a target. The determination may be spatial, for example, in two or three dimensions. In addition, the determination may be quantitative, for example, determining the quantity or concentration of the signal-generating entity and / or target.

[0142] In one set of embodiments, signal-generating entities may be conjugated to a secondary amplified nucleic acid (or other final amplified nucleic acid). The signal-generating entities may be conjugated to the secondary amplified nucleic acid (or other final amplified nucleic acid) before or after its association with the target in the sample. For example, the signal-generating entities may be conjugated to the secondary amplified nucleic acid first, or they may be conjugated after the secondary amplified nucleic acid has been applied to the sample. In some cases, signal-generating entities are added and then reacted to conjugate them to the amplified nucleic acid.

[0143] In one set of embodiments, signal-generating entities may be attached to a nucleotide sequence via a bond that can be cleaved to release the signal-generating entity. For example, after determining the distribution of nucleic acid probes in a sample, the signal-generating entities may be released or inactivated before another round of nucleic acid probes and / or amplified nucleic acids. Thus, in some embodiments, the bond may be a cleavable bond, such as a disulfide bond or a photocleavable bond. Examples of photocleavable bonds are discussed in detail herein. In some cases, such bonds may be cleaved upon exposure to, for example, a reducing agent or light (e.g., ultraviolet light). See below for further details. In some cases, signal-generating entities are inactivated by photobleaching. Other examples of systems and methods for inactivating and / or removing signal-generating entities are discussed in detail herein.

[0144] In certain embodiments, the use of primary and secondary amplified nucleic acids may be used to create a maximum number of signal-generating entities that can bind to a given nucleic acid probe. For example, there may be a maximum number of signal-generating entities that can bind to a nucleic acid probe, depending on the maximum number of readout probes having signal-generating entities capable of binding to a finite number of secondary amplified nucleic acids, and / or the maximum number of primary amplified nucleic acids capable of binding to a finite number of read sequences on the nucleic acid probe. Each potential position does not need to be actually filled by a signal-generating entity, but this structure suggests that there is a saturation limit of signal-generating entities beyond which any further signal-generating entities that may exist by chance cannot associate with the nucleic acid probe or its target.

[0145] Therefore, certain embodiments relate to systems and methods for amplifying signals pointing to nucleic acid probes or their targets such that they are generally saturable, i.e., there exists a saturation limit, which is an upper limit on how many signal-generating entities can associate with the nucleic acid probe or its target. Typically, this number is greater than 1. For example, the upper limit on signal-generating entities could be at least 2, at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, and so on. Depending on the case, the upper limit may be less than 500, less than 400, less than 300, less than 250, less than 200, less than 175, less than 150, less than 125, less than 100, less than 75, less than 50, less than 40, less than 30, less than 25, less than 20, less than 15, less than 10, less than 5, etc. Depending on the case, the upper limit may be determined as the maximum number of signal-generating entities that can bind to secondary amplified nucleic acids multiplied by the maximum number of primary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of secondary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of primary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of secondary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of primary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of primary amplified nucleic acids multiplied by the maximum number of secondary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of

[0146] However, it should be understood that the average number of signaling entities that actually bind to the nucleic acid probe or its target does not necessarily have to be the same as its upper limit; that is, the signaling entities may not actually be fully saturated (although full saturation is possible). For example, the saturation level (or the number of bound signaling entities compared to the maximum number that can bind) may be less than 97%, less than 95%, less than 90%, less than 85%, less than 80%, less than 75%, etc., and / or at least 50%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, etc. In some cases, giving time for binding to occur and / or increasing the concentration of the reagent may increase the saturation level.

[0147] Due to a potential upper limit on the number of signal-generating entities that actually bind to a nucleic acid probe or its target, spatially distributed binding events in a sample, for example, may present a substantially uniform size and / or brightness, in contrast to uncontrolled amplification, such as the uncontrolled amplification discussed above. For example, due to a specific number of secondary-amplified nucleic acids that can bind to primary-amplified nucleic acids, secondary-amplified nucleic acids cannot be found beyond a fixed distance from the nucleic acid probe or its target, which can limit the "spot size," or the diameter of fluorescence from the signal-generating entity, that indicates binding.

[0148] In a particular embodiment, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the coupled events may exhibit substantially the same brightness, size (e.g., apparent diameter), color, etc., which may facilitate the identification of the coupled events from other events, such as nonspecific couplings or noise.

[0149] In addition, the signal-generating entities may be inactivated in some cases. For example, in some embodiments, a first secondary nucleic acid probe or readout probe that can associate with the signal-generating entities (e.g., using amplified nucleic acid) and recognize a first readout sequence (e.g., on a primary nucleic acid probe or coding nucleic acid probe) may be applied to the sample, and then, for example, a second secondary nucleic acid probe or readout probe that can associate with the signal-generating entities (e.g., using amplified nucleic acid) may be applied to the sample to inactivate the signal-generating entities. When multiple signal-generating entities are used, the same or different techniques may be used to inactivate the signal-generating entities, and some or all of the multiple signal-generating entities may be inactivated, for example, sequentially or simultaneously.

[0150] Inactivation may be caused by the removal of the signal-generating entity (e.g., removal from a sample or nucleic acid probe) and / or by chemically modifying the signal-generating entity in some way (e.g., by photobleaching the signal-generating entity, by fading the signal-generating entity, or by chemically modifying the structure of the signal-generating entity, for example, by reduction). For example, in one set of embodiments, a fluorescent signal-generating entity may be inactivated by chemical or optical techniques such as oxidation, photobleaching, chemical bleaching, rigorous washing, or reactions by enzymatic digestion or exposure to enzymes, dissociation of the signal-generating entity from other components (e.g., probes), or chemical reactions of the signal-generating entity (e.g., reactants that can alter the structure of the signal-generating entity). For example, exposure to oxygen or reducing agents may cause fading, and the signal-generating entity may be chemically cleaved from the nucleic acid probe and washed away through a fluid stream.

[0151] In some embodiments, for example, using amplified nucleic acids discussed herein, diverse nucleic acid probes can be associated with one or more signal-generating entities. When more than one nucleic acid probe (or secondary nucleic acid probe or readout probe) is used, the signal-generating entities may be the same or different. In certain embodiments, the signal-generating entity is any entity capable of luminescence. For example, in one embodiment, the signal-generating entity is a fluorescent entity. In other embodiments, the signal-generating entity may be a phosphorescent entity, a radioactive entity, an absorbent entity, etc. Depending on the circumstances, the signal-generating entity may be any entity that can be determined in a sample at relatively high resolution, for example, at a resolution beyond the wavelength of visible light or the diffraction limit. The signal-generating entity may be, for example, a dye, a small molecule, a peptide, or a protein. Depending on the circumstances, the signal-generating entity may be a single molecule. When multiple secondary nucleic acid probes or readout probes are used, the nucleic acid probes may associate with the same or different generating entities.

[0152] Non-limiting examples of signal-generating entities include fluorescent entities (fluorophores) or phosphorescent entities, such as cyanine dyes (e.g., Cy2, Cy3, Cy3B, Cy5, Cy5.5, Cy7, etc.), Alexa Fluor dyes, Atto dyes, photoswitchable dyes, photoactivatable dyes, fluorescent dyes, metal nanoparticles, semiconductor nanoparticles, or "quantum dots."

[0153] In one set of embodiments, a signal-generating entity may be conjugated to an oligonucleotide sequence via a bond that can be cleaved to release the signal-generating entity. In one set of embodiments, a fluorophore may be conjugated to an oligonucleotide via a cleavable bond, such as a photocleavable bond. Non-limiting examples of photocleavage bonds include 1-(2-nitrophenyl)ethyl, 2-nitrobenzyl, biotin phosphoramidite, acrylic phosphoramidite, diethylaminocoumarin, 1-(4,5-dimethoxy-2-nitrophenyl)ethyl, cyclododecyl(dimethoxy-2-nitrophenyl)ethyl, 4-aminomethyl-3-nitrobenzyl, [4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl]methyl-S-acetylthioate, [4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl]methyl-3-(2-pyridyldithiopropionic acid) ester [(4-nitro-3-(1-thlorocarbonyloxyethyl)phenyl)methyl-3-(2-pyridyldithiopropionic acid] This includes, but is not limited to, acid)ester, 3-(4,4'-dimethoxytrityl)-1-(2-nitrophenyl)-propane-1,3-diol-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-trifluoroacetylcaproamidomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(4,4'-dimethoxytrityloxy)butylamidomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(N-(4,4'-dimethoxytrityl))-biotinamidecaproamido-methyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, or similar linkers. The oligonucleotide sequence may be a primary or secondary (or other) amplified nucleic acid, such as the amplified nucleic acid discussed herein.

[0154] In another set of embodiments, the fluorophore may be conjugated to an oligonucleotide via a disulfide bond. Disulfide bonds can be cleaved by a variety of reducing agents, including but not limited to dithiothreitol, dithioerythritol, beta-mercaptoethanol, sodium hydrogen boride, thioredoxin, glutarredoxin, trypsinogen, hydrazine, diisobutylaluminum hydride, oxalic acid, formic acid, ascorbic acid, phosphoric acid, tin chloride, glutathione, thioglycolate, 2,3-dimercaptopropanol, 2-mercaptoethylamine, 2-aminoethanol, tris(2-carboxyethyl)phosphine, bis(2-mercaptoethyl)sulfone, N,N'-dimethyl-N,N'-bis(mercaptoacetyl)hydrazine, 3-mercaptopropionate, dimethylformamide, thiopropyl agarose, tri-n-butylphosphine, cysteine, ferrous sulfate, sodium sulfite, phosphates, hypophosphates, phosphorothioates, and / or any combination thereof. Oligonucleotides can be primary nucleic acid probes, encoding nucleic acid probes, readout probes, or primary or secondary (or other) amplified nucleic acids, such as the probes discussed herein.

[0155] In another embodiment, the fluorophore may be conjugated to an oligonucleotide via one or more phosphorothioate-modified nucleotides in which sulfur modification replaces bridging and / or unbridging oxygen atoms. In a particular embodiment, the fluorophore may be cleaved from the oligonucleotide by adding a compound such as, but not limited to, iodine mixed in ethanol, iodoethanol, silver nitrate, or mercury chloride. In yet another set of embodiments, the signal-generating entity may be chemically inactivated via reduction or oxidation. For example, in one embodiment, a chromophore such as Cy5 or Cy7 may be reduced to a stable, non-fluorescent state using sodium boride. In yet another set of embodiments, the fluorophore may be conjugated to an oligonucleotide via an azo bond, which may be cleaved by 2-[(2-N-arylamino)phenylazo]pyridine. In yet another set of embodiments, the fluorophore may be conjugated to an oligonucleotide via a suitable nucleic acid segment that can be cleaved upon appropriate exposure to a DNase, such as an exodeoxyribonuclease or an endodeoxyribonuclease. Examples include, but are not limited to, deoxyribonuclease I or deoxyribonuclease II. In one set of embodiments, cleavage may occur via a restriction endonuclease. Non-limited examples of potentially suitable restriction endonucleases include BamHI, BsrI, NotI, XmaI, PspAI, DpnI, MboI, MnlI, Eco57I, Ksp632I, DraIII, AhaII, SmaI, MluI, HpaI, ApaI, BclI, BstEII, TaqI, EcoRI, SacI, HindII, HaeII, DraII, Tsp509I, Sau3AI, PacI, and the like. More than 3,000 restriction enzymes have been studied in detail, and more than 600 of these are commercially available. In yet another set of embodiments, fluorophores may be conjugated to biotin, and oligonucleotides may be conjugated to avidin or streptavidin.The interaction between biotin and avidin or streptavidin causes the fluorophore to bind to the oligonucleotide, while sufficient exposure to an excess addition may cause free biotin to "override" the binding, thereby resulting in the dissociation of the fluorophore from the oligonucleotide. In addition, in another set of embodiments, the probe may be removed using a corresponding "toe-hold probe" containing an extra number of bases (e.g., 1 to 20 extra bases, e.g., 5 extra bases) that has the same sequence as the secondary probe or readout probe, as well as homology to the primary probe or coding probe. These probes may remove the labeled secondary probe or readout probe via a strand displacement interaction. The oligonucleotide may be a primary nucleic acid probe, coding nucleic acid probe, readout probe, or primary or secondary (or other) amplified nucleic acid, such as the probes discussed herein.

[0156] As used herein, the term “light” generally refers to electromagnetic radiation having any suitable wavelength (or, synonymously, frequency). For example, in some embodiments, light may include wavelengths in the optical or visible range (e.g., light having wavelengths between about 400 nm and about 700 nm, i.e., “visible light”), infrared wavelengths (e.g., light having wavelengths between about 300 micrometers and 700 nm), ultraviolet wavelengths (e.g., light having wavelengths between about 400 nm and about 10 nm), and so on. In certain cases, as discussed herein, more than one entity, i.e., chemically different or significantly different entities, for example, structurally different, may be used. However, in other cases, the entities may be chemically identical or at least substantially chemically identical.

[0157] In one set of embodiments, the signal-generating entities are “switchable,” meaning they can be switched between two or more states in which at least one of them emits light of a desired wavelength. In other states, the entities may not emit light, or they may emit light of a different wavelength. For example, an entity can be “activated” to a first state in which it is capable of producing light of a desired wavelength, and “inactivated” to a second state in which it is not capable of emitting light of the same wavelength. An entity is “photoactivated” if it is activated by incident light of an appropriate wavelength. As an unrestricted example, Cy5 or Alexa647 can be switched between a fluorescent state and a non-fluorescent state in a controlled, reversible manner by light of different wavelengths; that is, red light at 633 nm (or 642 nm, 647 nm, 656 nm) can switch Cy5 or Alexa647 to a stable non-fluorescent state or inactivate it, while green light at 405 nm can switch Cy5 or Alexa647 to a fluorescent state or revert it back to activate it. In some cases, the entity can be reversibly switched between two or more states when exposed to a suitable stimulus, for example. For example, a first stimulus (e.g., light of a first wavelength) can be used to activate a switchable entity, while a second stimulus (e.g., light of a second wavelength) can be used to inactivate the switchable entity, for example, to a non-fluorescent state. Any suitable method can be used to activate the entity. For example, in one embodiment, incident light of a suitable wavelength may be used to activate and cause the entity to emit light, i.e., the entity is "light-switchable". Thus, a light-switchable entity can be switched between different luminescent and non-luminescent states by incident light of different wavelengths, for example. The light may be monochromatic (e.g., provided by a laser) or multicolored. In another embodiment, the entity may be activated when stimulated by an electric and / or magnetic field. In yet another embodiment, the entity may be activated when exposed to a suitable chemical environment, for example, by adjusting the pH or inducing a reversible chemical reaction involving the entity.Similarly, any suitable method may be used to inactivate an entity, and the method of activating an entity does not have to be the same as the method of inactivating an entity. For example, an entity may be inactivated by exposure to incident light of the appropriate wavelength, or it may be inactivated by waiting for a sufficient amount of time.

[0158] Typically, a "switchable" entity can be identified by a person skilled in the art by determining the conditions under which an entity in a first state can emit light when exposed to an excitation wavelength, switching the entity from the first state to a second state, for example, by exposure to light of a switched wavelength, and then demonstrating that when the entity is in the second state, it no longer emits light (or emits light of reduced intensity) when exposed to the excitation wavelength.

[0159] As discussed, in one set of embodiments, a switchable entity can be switched upon exposure to light. The light used to activate the switchable entity may originate from an external light source, such as a laser light source, or from another light-emitting entity adjacent to the switchable entity. The second light-emitting entity may be a fluorescent entity, and in certain embodiments, the second light-emitting entity may also be a switchable entity itself.

[0160] In some embodiments, the switchable entity includes a first light-emitting portion (e.g., a fluorophore) and a second portion that activates or “switches” the first portion. For example, when exposed to light, the second portion of the switchable entity may activate the first portion, causing the first portion to emit light. Examples of activating portions include, but are not limited to, Alexa Fluor 405 (Invitrogen), Alexa Fluor 488 (Invitrogen), Cy2 (GE Healthcare), Cy3 (GE Healthcare), Cy3B (GE Healthcare), Cy3.5 (GE Healthcare), or other suitable dyes. Examples of luminescent parts include, but are not limited to, Cy3B (GE Healthcare), Cy5, Cy5.5 (GE Healthcare), Cy7 (GE Healthcare), Alexa Fluor 647 (Invitrogen), Alexa Fluor 680 (Invitrogen), Alexa Fluor 700 (Invitrogen), Alexa Fluor 750 (Invitrogen), Alexa Fluor 790 (Invitrogen), DiD, DiR, YOYO-3 (Invitrogen), YO-PRO-3 (Invitrogen), TOT-3 (Invitrogen), TO-PRO-3 (Invitrogen), or other suitable dyes. See, for example, U.S. Patent No. 7,838,302, which is incorporated herein by reference in its entirety. The first luminescent part may subsequently be inactivated by any suitable technique (e.g., by directing 647 nm red light to the Cy5 portion of the molecule).

[0161] In some embodiments, multiple nucleic acid probes having different sequences are used, and the distribution of each nucleic acid probe is sequentially analyzed and used to generate a “codeword” for each position based on the binding pattern of each nucleic acid probe. By selecting nucleic acid probes that define an appropriate code space, obvious errors in the observed binding pattern can be identified and / or rejected and / or corrected to identify accurate codewords, and thus the precise target of the nucleic acid probe in the sample can be identified. This robustness and error correction system for errors was first introduced for MERFISH (multiplexed error-robust fluorescence in situ hybridization) and has since been used in a variety of related techniques. See, for example, International Patent Application Publications WO2016 / 018960 and WO2016 / 018963, each incorporated herein in whole by reference.

[0162] As mentioned, in certain embodiments, such techniques may be combined with error correction, as used in MERFISH or other similar techniques. For example, the codeword may be based on the binding (or unbinding) of multiple readout probes that may bind to a readout sequence on a primary nucleic acid probe or coding nucleic acid probe, and in some cases, the codeword may define an error correction code to help reduce or prevent misidentification of nucleic acid probes. In some cases, a relatively small number of readout probes may be used to identify a relatively large number of different targets, for example, by using a variety of combinatorial techniques. Fluorescence microscopes, wide-field fluorescence microscopes, epifluorescence microscopes, confocal microscopes, or light-sheet microscopes can be used for image acquisition. Additionally, image acquisition techniques such as STORM or other super-resolution imaging methods may be used to image such samples and facilitate the determination of nucleic acid probes. For further details regarding techniques such as MERFISH, see, for example, U.S. Patent No. 9,712,805 or No. 10,073,035, or International Patent Application Publication WO2008 / 091296 or WO2009 / 085218, each of which is incorporated herein by reference in its entirety. In some cases, expansion microscopy, in which the sample is expanded before imaging, may also be used. See, for example, International Patent Application Publication WO2018 / 089445, entitled "Matrix Imprinting and Clearing," or International Patent Application Publication WO2018 / 089438, entitled "Multiplexed Imaging Using MERFISH and Expansion Microscopy," each of which is incorporated herein by reference in its entirety.

[0163] Another aspect relates to computer implementations. For example, a computer and / or automated system may be provided that can automatically and / or repeatedly perform any of the methods described herein. As used herein, an “automated” device means a device that can operate without human instruction; that is, an automated device can function after any human has finished taking any action to advance its function, for example, by inputting a command to start a process into a computer. Typically, an automated device can perform a repeating function after this point in time. In some cases, the processing steps may also be recorded on a computer-readable medium.

[0164] For example, in some cases, a computer may be used to control the imaging of a sample, such as using fluorescence microscopy, wide-field fluorescence microscopy, epifluorescence microscopy, confocal microscopy, light-sheet microscopy, diffraction-limited optical microscopy, STORM, or other super-resolution techniques as described herein. Depending on the context, the computer may also control operations in image analysis, such as drift correction, physical registration, hybridization, and cluster alignment, cluster decoding (e.g., decoding of fluorescence clusters), error detection or correction (e.g., as discussed herein), noise reduction, and identification of foreground features from background features (e.g., noise or debris in the image). As an example, a computer may be used to control the activation and / or excitation and / or deactivation of signal-generating entities in a sample, and / or the acquisition of images of the signal-generating entities. In one set of embodiments, a sample may be excited using light of varying wavelengths and / or intensities, and a computer may be used to correlate the sequence of wavelengths of light used to excite the sample with images acquired for a sample containing signal-generating entities. For example, a computer may illuminate a sample with light of varying wavelengths and / or intensities to produce an average number of different signal-generating entities within each region of interest (e.g., one activated entity per location, two activated entities per location, etc.). In some cases, this information may be used to construct an image of the signal-generating entities and / or determine their locations at high resolution, as mentioned above.

[0165] In some embodiments, the sample is placed on a microscope. In some cases, the microscope may include one or more channels, such as fluid channels or microfluidic channels, for directing or controlling a fluid to or from the sample. For example, in one embodiment, a nucleic acid probe, such as the nucleic acid probes discussed herein, may be introduced and / or removed by the fluid to or from the sample via one or more channels. Optionally, there may also be one or more chambers or reservoirs, for example, fluid-connected to the channels and / or the sample, for holding the fluid. Those skilled in the art will be familiar with channels, including fluid or microfluidic channels, for moving a fluid to or from a sample.

[0166] The following documents: International Patent Application Publication No. WO2018 / 218150, titled "Systems and Methods for High-Throughput Image-Based Screening"; International Patent Application Publication No. WO2016 / 018960, titled "Systems and Methods for Determining Nucleic Acids"; International Patent Application Publication No. WO2016 / 018963, titled "Probe Library Construction"; International Patent Application Publication No. WO2018 / 089445, titled "Matrix Imprinting and Clearing"; International Patent Application Publication No. WO2018 / 089438, titled "Multiplexed Imaging Using MERFISH and Expansion Microscopy"; and U.S. Patent Application No. 62 / 836,578, titled "Imaging-Based Pooled CRISPR Screening" and "Amplification Methods and Systems for MERFISH and Other U.S. Patent Application No. 62 / 779,333, titled “Applications,” is incorporated herein by reference. Furthermore, U.S. Patents No. 2017 / 0220733 and 2017 / 0212986 are incorporated herein by reference in their entirety.

[0167] In addition, U.S. Patent Application Publication No. 62 / 954,720, entitled "Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin," and U.S. Patent Application Publication No. 63 / 060,947, entitled "Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin," are incorporated herein by reference in their entirety. [Examples]

[0168] The following examples are intended to illustrate certain embodiments of the present invention and not to illustrate the entire scope of the invention. [Example 1]

[0169] The following examples demonstrate large-scale multiplexed FISH techniques for imaging the 3D chromatin structure at genome scale in single cells, both of which further demonstrate the ability to place the 3D genomic structure in its native structural and functional context by combining genome-scale chromatin and nascent transcript imaging with nuclear structure identification.

[0170] This first example reports a large-scale multiplexed FISH technique that enables genome-scale imaging of chromatin composition in single cells. Using this technique, imaging and identification of over 1,000 distinct genomic loci (approximately 2,000 chromatin loci totaling homologous pairs of chromosomes) across the human genome in single cells were demonstrated. Furthermore, simultaneous imaging of these genomic loci was shown using nascent RNA transcripts of over 1,000 genes present at these loci in the context of various nuclear structures, including nuclear speckles, nucleoli, and nuclear laminas. This technique was used to explore the relationships between chromatin composition, transcriptional activity, and nuclear context in single cells.

[0171] To achieve genome-scale chromatin imaging, we devised a combinatorial FISH method that was conceived from multiplexed, error-robust FISH methods previously developed for transcriptome imaging, but with significant modifications specifically designed for chromatin imaging by considering both the polymeric nature of chromatin (i.e., adjacent loci in the genome sequence are spatially close) and the regional organization of chromosomes (i.e., different chromosomes tend to occupy separate spatial regions). See, for example, WO2016 / 018960, titled "Systems and Methods for Determining Nucleic Acids," and WO2016 / 018963, titled "Probe Library Construction," respectively, which are incorporated herein by reference in their entirety. To enable combinatorial imaging, each genomic locus was assigned a unique 100-bit binary code with 2 Hamming weights. That is, each barcode contains 2 "1" bits and 98 "0" bits (Figure 1A). The bit values ​​in these barcodes determined the presence (1) or absence (0) of a signal for each locus across sequential imaging rounds. From these 100-bit Hamming-weighted 2 barcodes, a subset was further selected to encode target genomic loci to avoid simultaneously imaging spatially close chromatin regions with the same bit, optimizing the assignment of these barcodes so that loci with "1" bits at the same barcode location were maximally separated in genomic space. This strategy minimized detection errors caused by overlapping signals originating from neighboring chromatin loci. Furthermore, since the vast majority of possible 100-bit binary codes were invalid (i.e., not assigned to any target locus), this design allowed for the identification and discarding of detection errors, further improving measurement accuracy.

[0172] The barcodes were physically imprinted onto target genomic loci using a highly diverse library of encoding probes, each containing a 40nt target region for binding to one of the target loci and a 20nt readout sequence selected from 100 pre-designed readout sequences (Figure 1A). Each readout sequence corresponds to one of 100 bits, and the encoding probe set for each genomic locus (approximately 400 probes per locus) contains only two different readout sequences, corresponding to the two bits read as "1" in the barcode assigned to that locus. After encoding probe binding, the barcodes imprinted on the chromatin loci were detected by sequential hybridization of fluorescently labeled readout probes complementary to one of the 100 readout sequences (Figure 1A). Two different readout probes were introduced per hybridization round, and after 50 rounds of hybridization, approximately 1000 genomic loci were imaged in two color channels to image and identify a total of approximately 1000 genomic loci (Figures 1A-1C). In contrast, a linear sequential approach to imaging 1000 genomic loci would have instead required 500 rounds of hybridization along with two-color imaging. Since each chromosome in diploid cells had two homologs, homolog identification of the loci being imaged was further assigned using a clustering algorithm that leveraged the tendency of chromosomes to occupy different regions within each nucleus.

[0173] In this example, 1,041 genomic loci, each approximately 30 kb in size and uniformly encompassing 22 autosomes and X chromosomes, were selected for imaging in human lung fibroblasts (IMR90). Each chromosome also needed to contain at least 30 target loci, and therefore, the number of loci imaged per chromosomal homolog needed to range from 30 to 80, depending on the chromosome length. These 1,041 genomic loci in approximately 5,400 individual cells across five biological repeats were imaged with a detection efficiency of approximately 80% for each locus. Considering two homologs per chromosome, approximately 1,700 chromatin loci were detected in each cell (Figures 1D-1E).

[0174] To obtain a population-averaged figure of chromatin composition within each cell, the spatial distance between all pairs of imaged chromatin loci was calculated, and then both the median distance and the contact frequency between all pairs of loci were determined across all imaged cells (Figure 1F and Figure 6A). The contact frequencies between pairs of chromatin loci within the same chromosome, determined from the imaging data, showed a high correlation with the contact frequencies detected by Ensemble Hi-C, with a Pearson correlation coefficient of 0.91 (Figure 6B). The imaging data captured chromatin structure at multiple scales, from chromosome composition within regions (Figure 1F and Figure 6A) to the formation of A and B compartments within chromosomal arms (Figure 7A), and showed good agreement with the compartments identified by Ensemble Hi-C measurements (Figures 7B-7C). Furthermore, the imaging results showed high reproducibility between independent biological replicates (Figure 8).

[0175] By exploring the chromatin composition within individual cells, we found that chromosomes occupied different regions within each cell (Figure 1F-1G) while also exhibiting substantial overlap with one another (Figure 1G-1H). On average, approximately 80% of the convex hull volume occupied by any given chromosome was shared with other chromosomes within the same cell (Figure 1I), suggesting a high degree of trans-chromosome interaction. Since these interactions were not sufficiently explored, the following analysis focused on these trans-chromosome interactions.

[0176] Figures 1A–1I show genome-scale chromatin imaging. Figure 1A shows the imaging scheme. An error-robust barcode, e.g., a 100-bit binary barcode with a Hamming weight of 2, was assigned to the target genomic locus (i.e., two of the 100 bits are read as "1"). The barcode was imprinted onto the genomic locus using encoding oligonucleotide probes that recognize the locus, associate two different readout sequences with each locus, and correspond to the two bits read as "1" in the barcode assigned to the locus. Each locus was labeled with a total of 400 encoding probes, but only four are shown. By sequentially adding fluorescent readout probes complementary to the readout sequences and imaging, it was possible to determine the bits read as "1" at each locus and therefore the barcode identification of that locus. Approximately 1000 genomic loci were imaged. Figure 1B shows representative images from multiple imaging rounds within the nucleus of a single cell. The fluorescent signals of chromatin loci from readout probes are shown in the lighter shadow, while the signal of 4',6-diamidino-2-phenylindole (DAPI), used as a nuclear marker, is shown in the darker shadow. The scale bar is 5 micrometers. Figure 1C is a magnified image of a small spatial region (enclosed in Figure 1B) centered on one chromatin locus across all imaging rounds. Locus identification was determined based on two signaling readout probes (1 and 13). The scale bar is 300 nm. Figure 1D is a 3D rendering of all detected chromatin loci in a single cell, shown in grayscale according to the chromosome to which they belong. Adjacent loci in the genome sequence are connected by thin lines. Figure 1E shows the chromatin loci of the same cell as in Figure 1D, but with two homologs of the shown chromosomes shown in grayscale, distinct from all other loci. Figure 1F is the median distance matrix calculated from approximately 5,400 single cells. For each pair of loci, the median of all observed 3D spatial distances between the loci is presented.Figure 1G shows an illustrative image illustrating the locations of multiple chromosomal regions within a single cell. Shaded regions represent the convex hull surrounding each chromosome, which was used as an operational definition of the chromosomal region. Figure 1H shows a distance matrix for the same cell shown in Figure 1G. Spatial distances between each pair of chromatin loci are shown. Chromosome order is as shown below the heatmap, with two homologs of each chromosome shown separately. Figure 1I quantifies the percentage of volume of each chromosomal region shared by at least one other chromosome within the same cell. The median (centerline), 25th–75th percentile (box), and 5th–95th percentile (whiskers) are shown. n = 10,910 copies of chromosomes (5,455 cells and two homologous copies per cell for each chromosome).

[0177] Figures 6A-6B show a comparison of the contact frequency matrix derived from genome-scale imaging with ensemble Hi-C data. Figure 6A shows the contact frequency matrix for all 1041 genomic loci imaged in this example. The contact frequency between locus pairs was calculated by dividing the number of occurrences where the measured distance between loci was shorter than 500 nm by the total number of measured distances between the two loci. Figure 6B shows a correlation plot of contact frequency between chromosomal locus pairs derived from imaging data and those derived from ensemble Hi-C experiments, binned at 500 kb and centered on the target locus. The Pearson correlation coefficient was 0.91.

[0178] Figures 7A–7C show a comparison of subchromosome structures derived from genome-scale imaging with ensemble Hi-C data. Figure 7A shows a contact frequency matrix generated from imaging data for one arm of chromosome 22. The assignment of each locus to either the A or B compartment based on this matrix is ​​shown in the bars below the matrix. Figure 7B shows a contact frequency matrix for the same arm of chromosome 22, calculated from Hi-C data binned at 500kb and centered on the target locus. The bars below the matrix show the assignment of each locus to the A and B compartments based on this matrix, assigned using the same procedure as in Figure 7A. The A / B compartment assignments derived from imaging data and Hi-C data are identical. Figure 7C shows a correlation plot for contact frequencies between locus pairs in chromosome 22 derived from imaging data and those derived from ensemble Hi-C experiments. The Pearson correlation coefficient was 0.91.

[0179] Figure 8 shows the reproducibility of chromatin imaging experiments between replicates. The plot shows the correlation of pairwise distances between chromatin loci observed in two independent biological replicates of imaging experiments for 1041 genomic loci. The Pearson correlation coefficient between replicates was 0.98. The upper right cloud represents pairwise distances of trans chromosomes, and the lower left cloud represents pairwise distances within chromosomes. [Example 2]

[0180] In this embodiment, we investigated how trans-chromosome interactions depend on the epigenetic properties of chromatin. Previous Hi-C and imaging analyses have shown that chromatin is separated into compartments A and B, enhanced by active and inactive chromatin, respectively. Different mechanisms may mediate active-active and inactive-inactive chromatin interactions, such as HP1-mediated heterochromatin condensation and transcription factor and cofactor-mediated active chromatin condensation. In this embodiment, each imaged genomic locus was classified into compartments A and B using an established calling method based on published ensemble Hi-C data. 38% of the imaged loci belonged to compartment A, which was relatively gene-rich and tended to be enhanced by active chromatin markers such as H3K27Ac, ​​while 62% belonged to compartment B and tended to be enhanced by inactive chromatin markers such as H3K9me3. To examine whether the degree of trans-chromosome interaction differs for active and inactive chromatin, genomic loci were ordered in a trans-chromosome contact frequency matrix, placing all A loci next to each other and all B loci together. This matrix showed that compartment A loci had a substantially stronger tendency to interact trans-chromosome with compartment A loci than compartment B loci (Figures 2A-2B). In contrast, compartment B loci did not show similar trans-chromosome affinity toward each other, but instead showed a slightly higher probability of interacting trans-chromosome with compartment A chromatin (Figures 2A-2B). In other words, trans-chromosome AA interactions appeared with a substantially stronger tendency than AB interactions, and then with a slightly stronger tendency than BB interactions. This is in striking contrast to cis interactions within the same chromosome, where A and B compartments tended to condense, resulting in enrichment of both AA and BB interactions more than AB interactions.

[0181] Next, the epigenetic dependence of trans chromosome interactions was examined at the single-cell level. Within individual cells, compartment A and compartment B loci adopted different spatial distributions, with A loci tending to be more centrally localized in the nucleus than B loci (Figure 2C and Figures 9A-9B). There was also a substantial degree of mixing between A and B loci (Figure 2C and Figures 9A-9B). For each imaged locus within each chromosome, its local density of A and B loci derived from all other chromosomes was calculated, and the ratio of these two densities was determined (hereinafter referred to as the trans chromosome A / B density ratio) (Figure 2C). This amount provided a measure of local enrichment of trans chromosome active chromatin near the locus. The majority (62%) of the imaged loci belonged to compartment B, and the overall bias regarding the A / B ratio was less than 1. To control for this bias, the distribution of trans chromosome A / B density ratios observed for loci A and B was compared to the distribution obtained in randomized controls, where the number of A and B loci remained unchanged, and the A and B discrimination of the imaged loci was randomly shuffled among the imaged loci. In particular, the trans chromosome A / B density ratio observed for loci A was substantially higher than the value observed for loci B, and subsequently higher than the value derived from the randomized controls (Figure 2D), and this trend was observed in many single cells (Figure 2E). These single-cell analyses again supported the view that trans chromosome interactions preferentially enhance interactions between active chromatins.

[0182] Figures 2A–2E show that trans chromosome contacts are preferentially enhanced for interactions between active chromatin. Figure 2A shows the normalized trans chromosome contact frequency matrix. The contact frequencies between each trans chromosome locus pair (a pair of loci on different chromosomes) are shown. The loci are rearranged so that loci in compartment A appear first, followed by loci in compartment B, so that the upper left block represents interactions between pairs of loci in compartment A, and the lower right represents interactions between pairs of loci in compartment B. Each entry in the matrix is ​​normalized by the median contact frequency of all locus pairs originating from the same pair of chromosomes that constitute the basic level of changing interactions between pairs of chromosomes. Figure 2B shows the distribution of trans chromosome contact frequencies for pairs of A loci (AA, right; n=72,771 locus pairs), pairs of B loci (BB, left; n=193,753 locus pairs), and pairs containing one A and one B locus (AB, n=237,986 locus pairs), derived from the matrix shown in Figure 2A. The distribution is shown as a histogram in the upper panel and as a box plot in the lower panel, showing the median (centerline), 25th–75th percentile (box), and 5th–95th percentile (whiskers). Figure 2C shows the distribution of compartment A and compartment B loci in a single cell. The left panel shows the location of all detected loci in a single z-plane within a single nucleus. Compartment A loci are shown above the scale, while compartment B loci are shown below. In the right panel, the shadows for each locus represent the ratio of local densities of trans chromosome A and B loci, according to the scale bar of the shadow shown on the right. Figure 2D shows the distribution of local trans chromosome A / B density ratios for the imaged genomic loci. For each locus, the median A / B density ratio across all cells was determined, and the distribution of different loci is shown along with the loci in compartment A (n=382 loci) and compartment B (n=623 loci).Of the 1041 imaged loci, 36 were not assigned A / B classification due to the different versions of genome assembly and Hi-C datasets used for compartment calling in this study. The dark gray histogram represents the randomized control, where the A and B compartment classifications were randomly shuffled while the total number of A and B loci remained unchanged. Figure 2E shows the distribution of enrichment in trans chromosome A / B density ratios compared to the randomized control. For each imaged cell, the median A / B density ratio across all A loci is divided by the median A / B density ratio of the randomized control, as shown in Figure 2D, and the distribution of this value across all imaged cells is presented (n=5,455 cells). The same enrichment distribution is shown for B loci (n=5,455 cells). The line marks a value of 1, i.e., no enrichment.

[0183] Figures 9A and 9B show that the loci of compartment A and compartment B exhibit different spatial distributions within single cells. The left panel of Figure 9A shows illustrative images of the compartment A and compartment B loci in a single z-plane within a single cell. The right panel shows the distribution of distances to the nuclear periphery for the compartment A and compartment B loci within these single cells. The nuclear periphery is identified as the convex hull surrounding all detected chromatin loci. The histogram shows the distribution of distances from the nuclear periphery for points uniformly sampled within the convex hull surrounding the detected chromatin loci. Figure 9B shows the population mean distribution of distances to the nuclear periphery for the compartment A and compartment B loci. There are n=382 A loci; n=623 B loci. [Example 3]

[0184] To place the 3D structure of chromatin in the context of its functional activity and other nuclear structures, the imaging method in this embodiment was extended to enable the simultaneous measurement of chromatin structure together with the transcriptional activity of several genomic loci and nuclear landmark structures. Specifically, 1,041 genomic loci were imaged together with the nascent RNA transcribed from each of the 1,137 genes located at these loci, as well as simultaneously with important nuclear structures, including nuclear speckles and nucleoli (Figure 3A).

[0185] To enable imaging of DNA, RNA, and nuclear structures within the same cell, we performed multiplexed imaging of intronic RNAs of 1,137 genes by employing a combinatorial imaging strategy similar to that used for chromatin (Figure 3A). Considering that not all genes are transcribed in individual cells, and therefore the density of transcription foci cannot be the same as that of chromatin loci, we encoded the RNA with a 54-bit Hamming weight 2 code and selected 1,137 possible barcodes to encode the genes in a manner similar to how barcodes for chromatin imaging are selected to minimize the opportunity to image spatially close genes within the same bit. After RNA imaging was complete, the RNA transcripts were enzymatically digested (a step also performed in our single-mode chromatin imaging experiments), and multiplexed DNA FISH was performed as described above to image 1,041 genomic loci (Figure 3A). Decoding of genomic loci and nascent RNA transcripts was performed largely independently, with additional constraints regarding transcripts co-localized with the genomic loci in which they reside. This procedure further improved the detection accuracy for RNA transcription and allowed for the estimation of the detection efficiency (approximately 90%) of transcriptional bursts at each genomic locus. Finally, nuclear speckles and nucleoli were imaged using immunofluorescence against known molecular components of these structures (Figure 3A). The location of nuclear laminas was estimated by calculating the convex hull encompassing all imaged genomic loci and determining the boundaries of the convex hull. Simultaneously, these multimode measurements enabled an integrated single-cell view of 3D genomic structure, transcriptional activity, and nuclear composition (Figure 3B). These multimode imaging experiments were performed on approximately 3700 individual cells in two biological replicates. Chromatin imaging data from these multimode experiments were also included in the above five replicates (approximately 5400 cells) for 3D genomic composition analysis.

[0186] From the measurement of nascent RNA transcripts in these multimode experiments, both transcriptional burst frequency (Figure 3C), as the proportion of cells actively transcribing the gene, and median burst size (Figure 3D), derived from the brightness of the RNA intron signal, were quantified for each gene. These measures showed high correlation across replicated experiments (Figures 10A-10B). Burst frequency exhibited two-pattern behavior: genes with high burst frequencies were mainly located in compartment A, while genes with low burst frequencies were located in both compartments (Figure 3C). Furthermore, using a 250 nm cutoff spatial distance, we estimated whether specific chromatin loci associated with nuclear structures and found that compartment B loci had a higher association frequency with nuclear lamina (Figure 11), and compartment A loci had a higher association frequency with nuclear speckles (Figure 12). These results were consistent with previous observations of preferential association of inactive and active chromatin with lamina and nuclear speckles, respectively. For individual gene loci, the local trans chromosome A / B density ratio showed a negative correlation with lamina association frequency (Figure 3E) and a positive correlation with nuclear speckle association frequency (Figure 3F). Finally, nucleoli showed preferential association with centromeres, telomeres of certain chromosomes, and chromosomes containing genes encoding ribosomes (Figure 3G). These biological results provided further justification for multimodal measurements.

[0187] In particular, for virtually all imaged loci, lamina association reduced observed transcriptional activity (Figure 3H), while nuclear speckle association correlated with higher transcriptional activity for many imaged loci (Figure 3H). In addition, upon treatment with alpha-amanitin to inhibit transcription, the association rate with laminas increased overall for almost all loci, while the association rate with nuclear speckles decreased overall (Figures 13A-13C). These results build upon previous imaging studies of nuclear rearrangement of single or several genomic regions during transcriptional activation or inhibition, providing a genomic-scale view of the relationship between transcriptional activity and its interaction with nuclear structure.

[0188] Figures 3A–3H show genome-scale imaging of chromatin and transcriptional activity in the context of nuclear structure. Figure 3A is a diagram of a multimode imaging scheme that combines imaging of chromatin (left panel), nascent RNA transcripts (center panel), and intranuclear structures (right panel) to generate an integrated view of chromatin composition in the context of nuclear structure and functional activity. Approximately 1000 genomic loci, nascent RNA transcripts of approximately 1100 genes in target loci, and two types of intranuclear structures (nuclear speckles and nucleoli) are imaged. Below are representative raw images for each imaging modality: chromatin loci (left) across multiple imaging rounds, nascent RNA transcripts (center) and intranuclear structures (right: nuclear speckles; left: nucleoli) across multiple imaging rounds. Scale bar is 5 micrometers. Figure 3B is a 3D rendering of chromatin loci, transcriptional bursts, and intranuclear structures in a single cell. Left: All detected chromatin loci are shown in grayscale by chromosome (based on chromosome indices shown below). Center: All detected intron RNAs are shown as spheres, with shadows indicating identification of the imaged gene, and the size of the sphere representing the size of the transcription burst. Chromatin loci are shown in the background. Right: Volume-filled representation of detected intranuclear structures. Nucleoli and nuclear speckles are shown with different shadows. Nuclear laminas are identified as the surface of the convex hull surrounding all detected chromatin loci. Figures 3C–3D show the distribution of transcription burst frequencies (Figure 3C) and burst sizes (Figure 3D) for genes present in compartment A loci (n=494 genes) and compartment B loci (n=625 genes). Figures 3E–3F are scatter plots of local trans chromosome A / B density ratios for each imaged genomic locus as a function of frequency found when the locus associates with nuclear laminas (Figure 3E) and nuclear speckles (Figure 3F). A gene locus is considered to be associated with a nuclear structure if its measured distance to that structure is less than 250 nm. The trans chromosome A / B density ratio values ​​shown in the plot are median values ​​across all imaged cells.Figure 3G shows the nucleolar association frequencies for all imaged genomic loci, ordered by genomic location. Vertical lines represent centromere locations, and brackets highlight chromosomes containing ribosome-coding genes (rDNA). Figure 3H shows the effect of nuclear structure association on transcription. Circles represent the multiplier change in transcriptional burst frequency for each locus when comparing cell populations where the locus associates with / does not associate with lamina (left) and where it associates with / does not associate with speckle (right). Dotted lines highlight no change, and solid lines represent the median multiplier change in each case.

[0189] Figures 10A–10B show the reproducibility of nascent RNA transcription imaging experiments between repeats. Figures 10A–10B show the correlation between RNA imaging repeats regarding burst frequency (Figure 10A) and burst size (Figure 10B) for each gene. The Pearson correlation coefficients were 0.94 and 0.81, respectively.

[0190] Figure 11 shows the preferred association between compartment B loci and nuclear lamina. The distribution of association rates between compartment A loci (n=382) and compartment B loci (n=623) and nuclear lamina is shown. A locus is operationally defined as associated with lamina if its distance to the nuclear periphery is less than 250 nm.

[0191] Figure 12 shows the preferred association between compartment A loci and nuclear speckles. The distribution of association rates between compartment A loci (n=382 loci) and compartment B loci (n=623 loci) and nuclear speckles is shown. A locus is operationally defined as associated with a speckle if its distance to the nearest speckle is less than 250 nm.

[0192] Figures 13A–13C show changes in the association of nuclear lamina and nuclear speckle during transcriptional inhibition. Figures 13A–13B show representative images of individual nuclei, including imaged chromatin loci, nucleoli, and nuclear speckle, for untreated cells (Figure 13A) and cells treated with the transcription inhibitor alpha-amanitin (Figure 13B). Figure 13C shows the multiplier changes in the association rates of each locus with lamina (left) and nuclear speckle (right) during transcriptional inhibition by alpha-amanitin. Data points for each genomic locus are indicated by circles, the solid line represents the median multiplier change for all loci in each case, and the dotted line represents no change. [Example 4]

[0193] In this example, these multimode single-cell measurements were used to further characterize trans chromosome interactions in the context of transcriptional activity and nuclear structure. Considering the observation that trans chromosome interactions were preferentially enhanced for interactions between compartment A loci, we tested whether these interactions correlated with chromatin transcriptional activity. To this end, local densities of A and B chromatin derived from other chromosomes, as well as the trans chromosome A / B density ratio, were calculated for each locus in each cell. The median values ​​of these amounts were determined for two populations of cells: (i) cells in which the locus under consideration exhibited transcriptional activity (i.e., RNA burst signaling), and (ii) cells in which the locus appeared transcriptionally silent, at least momentarily (Figure 4A). In particular, in addition to the observation that compartment A loci showed a higher local trans chromosome A / B density ratio than compartment B loci (Figures 2D-2E), a consistent trend of higher trans chromosome A density and A / B density ratio was observed for the same locus when it was actively transcribed (Figures 4B and 14). The observed correlation between transcriptional activity and trans-chromosome interactions was consistent with both of the following interpretations: higher epigenetic or transcriptional activity of a chromatin locus increases its rate of trans-chromosome interactions, or the arrangement of a locus in an environment enhanced by active chromatin enhances its transcriptional activity.

[0194] Figures 4A–4F demonstrate that preferential trans-chromosome interactions between active chromatin correlate with transcription and are disrupted by treatments that disturb condensate formation. Figure 4A shows single-cell images of chromatin loci and transcriptional activity. Left: Location of all imaged compartment A (top of scale) and B (bottom of scale) loci in a single z-plane derived from a single nucleus. Center: Local trans-chromosome A / B density ratio for the same locus, based on the grayscale scale bar. Right: Same as the center panel, with detected transcription bursts superimposed and shown as circles. Figure 4B shows a comparison of local trans-chromosome A / B density ratios for each locus in transcriptional and silent states. For each genomic locus containing at least one imaged gene, the trans-chromosome A / B density ratio was calculated for cells in which it is actively transcribed (designated as transcriptional) and cells in which it is not transcribed (designated as silent). Median values ​​across cells are shown for each state. The loci were ordered by their A / B density ratio in the silent state, and the A / B density ratios were plotted for both the silent and transcriptional states. Figure 4C shows the normalized trans chromosome contact frequency matrix for cells treated with alpha-amanitin to inhibit transcription. The matrix is ​​ordered and normalized as shown in Figure 2A. Figure 4D shows the distribution of AA, BB, and AB contact frequencies, shown as box plots, as shown in Figure 2B. There are n=72,771 locus pairs for AA, n=193,753 locus pairs for BB, and n=237,986 locus pairs for AB. Figures 4E-4F are the same as Figures 4C-4D, but for cells treated with 1,6-hexanediol.

[0195] Figure 14 shows the local density of trans chromosome A loci near each imaged locus when the locus is in an active transcriptional state or a silent state. For each locus, cells were divided into two groups depending on whether the locus was actively transcribed or silent. The median local density of A loci is shown for the cells in these two groups (transcriptional and silent). Loci are ordered based on their local trans chromosome A locus density in the silent state.

[0196] Since nuclear speckles are one of the most prominent intranuclear structures that concentrate actively transcribed loci, association with nuclear speckles may provide a simple explanation for the observed preferential development of trans-chromosomal active-active chromatin interactions. Interestingly, when the analysis was limited to loci not associated with nuclear speckles, the same trend was observed for AA enrichment compared to AB and BB interactions between trans-chromosomal contacts (Figure 15A), as well as for actively transcribed loci showing a higher local A / B density ratio compared to silent loci (Figure 15B). Notably, these trends were also observed when considering only loci associated with laminas and thus in an environment enhanced by compartment B chromatin (Figures 16A-16B). This latter result also indicated that the observed enrichment of active-active trans-chromosomal interactions cannot be simply composed of the fact that active chromatin is more concentrated toward the center of the nucleus.

[0197] Figures 15A–15B show enrichment of active-active trans chromosome interactions between chromatin loci not associated with nuclear speckle. Figure 15A shows the trans chromosome contact frequencies between A locus pairs (AA), B locus pairs (BB), and pairs containing one A and one B locus (AB), considering only cells in which both loci are not associated with nuclear speckle. Contact frequencies were normalized as described for Figure 2A. The distribution is represented as a box plot, as described for Figure 2B. There are n=72,771 locus pairs for AA (left), n=193,753 locus pairs for BB (right), and n=237,986 locus pairs for AB (center). For comparison, the median of all data, regardless of speckle association status, is shown as a triangle. Figure 15B shows the multiplicative change in local trans chromosome A / B density ratio between transcriptional status and silent status for loci not associated with nuclear speckle. For each genomic locus, considering only cells in which the locus is not associated with nuclear speckle, the multiplicative change in the local trans chromosome A / B density ratio between the transcriptional and silent states of the locus was calculated. The median A / B density ratio in each state (transcribed or silent) was determined for each locus, and the multiplicative change between the two states is shown on the left (each circle corresponds to a genomic locus). For comparison, the corresponding multiplicative change derived from all data, regardless of the nuclear speckle association state of the locus, is shown on the right. In all cases, the dotted line represents no change, and the solid line represents the median multiplicative change across all loci.

[0198] Figures 16A-16B show the enrichment of active-active trans chromosome interactions between chromatin loci associated with nuclear laminas. Figure 16A shows the trans chromosome contact frequencies between A locus pairs (AA), B locus pairs (BB), and pairs containing one A and one B locus (AB), considering only pairs of loci associated with laminas (within 250 nm). The contact frequencies were normalized as described for Figure 2A. The distribution is shown as a box plot as described for Figure 2B. For AA, there are n=72,771 locus pairs (left), for BB, n=193,753 locus pairs (right), and for AB, n=237,986 locus pairs (center). For comparison, the median of all data, regardless of lamina association status, is shown as a triangle. Because only a relatively small proportion of loci associate with lamina, the variance is relatively large in these cases; nevertheless, the differences between different types of AA, BB, and AB pairs are statistically significant (P-value < 10⁻¹⁰). Figure 16B shows the multiplicative change in local trans chromosome A / B density ratio between transcriptional and silent states for loci associated with nuclear lamina. For each genomic locus, the multiplicative change in local trans chromosome A / B density ratio between transcriptional and silent states was calculated, considering only cells in which the locus associated with nuclear lamina. The median A / B density ratio in each state (transcribed or silent) was determined for each locus, and the multiplicative change between the two states is shown on the left (each circle corresponds to a locus). Outliers (33 loci above and 18 below the presented scale) have been omitted to allow for a clearer visualization of the median multiplicative change. For comparison, the multiplicative change derived from all data, regardless of lamina association status, is shown on the right. In all cases, the dotted line represents no change in the ploidy level, and the solid line represents the median change in ploidy level across all loci.

[0199] The results indicated that trans-chromosome interactions preferentially occur between active chromatin loci, and that this behavior was consistently observed across multiple different nuclear environments. The next question explored was what might potentially trigger these preferential, broad-spectrum active-chromatin interactions. Since RNA polymerase II (Pol II) contains a low-complexity domain (LCD) and can form condensates, we tested whether Pol II-mediated transcription could be responsible for these preferential trans-chromosome interactions by using alpha-amanitine, a transcription inhibitor that causes Pol II dissociation and degradation. Despite inactivating transcription and altering its association with nuclear structure and chromatin (Figures 13A-13C), treatment with alpha-amanitine did not substantially reduce preferential trans-chromosome interactions between active chromatin (Figures 4C-4D), suggesting that additional or other active chromatin-binding factors were involved in these trans-chromosome interactions. Several other proteins associated with active chromatin were shown to contain LCDs that may potentially mediate condensate formation. Therefore, the objective was to disrupt condensate formation more generally by using 1,6-hexanediol, a drug known to disrupt hydrophobic interactions between LCDs. In particular, preferential enrichment of active chromatin interactions in trans-chromosome contacts was largely neutralized upon treatment of cells with 2% 1,6-hexanediol for 45 minutes (Figures 4E-4F), suggesting a potential role of condensate formation in the establishment or maintenance of these interactions.

[0200] In summary, these examples developed large-scale multiplexed FISH methods for imaging the 3D chromatin composition at genome scale in single cells, and both further demonstrated the ability to place 3D genomic composition in its innate structural and functional context by combining genome-scale chromatin and nascent transcript imaging with nuclear structure identification. This provides an integrated view of nuclear composition in single cells. Here, target loci were selected uniformly across all chromosomes to provide an unbiased view of the overall 3D genomic composition, but this method can also be used to target genomic loci with specific structural and functional characteristics, such as loci to which promoters, enhancers, and certain nuclear structural proteins are bound, and to test the interactions between these loci as well as their relationship to transcription and other chromatin functions. The broad application of this method to various questions related to genomic composition can reveal both the mechanisms governing chromatin composition and the role of chromatin structure in regulating genomic function. [Example 5]

[0201] This embodiment illustrates various materials and methods that can be used in the above embodiment.

[0202] Target genomic regions. For chromatin imaging, genomic loci were selected for imaging using the following method: For each human autosome and X chromosome, 30kb segments were selected at intervals of approximately 3Mb. If this interval resulted in fewer than 30 selected loci on a given chromosome, the interval was reduced for that chromosome until all chromosomes had at least 30 selected loci. This resulted in a total of 1,041 target genomic loci for imaging, with the number of loci on individual chromosomes ranging from 30 to 80. Then, encoding probes were designed for each 30kb segment for combinatorial FISH imaging (approximately 400 oligonucleotide probes).

[0203] For imaging of nascent RNA transcripts, all intron-containing genes that completely or partially overlap with target genomic loci were selected. Then, coding probes were designed for the introns of all these RNAs, with each RNA having approximately 20 coding probes, ensuring that the targeting sequences of the coding probes were preserved as close as possible to the transcription start site (TSS). A total of 1,137 genes were targeted.

[0204] Barcode design for combinatorial FISH imaging. Binary barcodes for imaging 1,041 genomic loci were selected in the following manner. First, all possible 100-bit binary barcodes with a Hamming weight of 2 (i.e., each barcode contains 2 "1" bits and 98 "0" bits) were generated, and 1,041 barcodes were randomly selected from this list. Next, the selected barcodes were arbitrarily assigned to the first 1,041 genomic loci. Then, barcodes were randomly swapped between the used and unused code pools, and between loci originating from different chromosomes, to minimize the variance in the number of loci appearing across different bits (i.e., read as "1") for each chromosome. This resulted in approximately equal numbers of loci being imaged per bit for each chromosome. To optimize the association of barcodes to loci within each chromosome, loci within the same chromosome were barcode-swapped and optimized for the maximum minimum genomic distance between loci with barcodes read as "1" at the same coding position. When comparing coding assignments with the same minimum genome distance, we selected the one that minimized the coefficient of variation in genome distance (so that the genome distance has both a larger mean and a smaller standard deviation).

[0205] Barcodes for imaging the nascent RNA transcripts of 1,137 genes were selected using 54-bit codes with a Hamming distance of 2, instead of similar codes with a 100-bit code with a Hamming distance of 2.

[0206] Encoding probe design. Encoding probes for chromatin imaging were synthesized from a pool of oligonucleotides purchased from Twist Biosciences. Each oligo in this pool used the following subsequence (from 5' to 3'): 1. 20-nucleotide (nt) forward priming region for PCR amplification and reverse transcription (RT) 2. A 20nt readout sequence corresponding to one of the bits in which the genomic locus targeted by the probe is imaged. 3. A 40nt target sequence designed to uniquely bind to a single target genomic locus. 4. Further copies of the above 20nt readout sequence 5. A 20nt reverse priming sequence for PCR amplification.

[0207] Forward and reverse priming sequences were selected from a previously generated list of randomly generated 20nt sequences optimized for PCR.

[0208] The readout sequences were selected using the following process: First, a list of 30nt sequences showing minimal homology to the human genome was created. Next, a subset of these sequences was ranked by the observed signal-to-noise ratio (SNR), and the top 100 were selected as DNA readout probes. Finally, the readout sequences were selected by reverse-interpolating the last 20nt of each of these sequences.

[0209] 40nt target sequences were similarly selected. Briefly, the following procedure was repeated for each target genomic region (see the discussion of "Target Genomic Regions" above): First, a list of all 40nt sequences complementary to the target genomic region was created (starting with each possible base in the target region). Next, the sequences were filtered by requiring that they fall within specified ranges for melting point and GC content. Then, the remaining sequences were further filtered by limiting the acceptable degree of homology with databases containing the human genome, human transcriptome, and repetitive sequences. Homology was determined by creating a table of all possible 17nt sequences and the number of times they appear in the target database (e.g., human genome, human transcriptome) and calculating the total number of exact 17nt matches that a given candidate sequence has with it. Finally, after a final filtering step such that no genomic duplication exists between any pair of target sequences, target sequences were selected from the remaining sequences.

[0210] To generate full-length probes, each selected 40nt target sequence for each target genomic locus was alternately assigned to two groups covering all target loci. Each of these groups was associated with a single readout sequence corresponding to one of the two bits from which the locus was imaged. Each target sequence was then ligated to two identical copies of the readout sequences assigned to that group, and subsequently ligated to forward and reverse PCR primers.

[0211] The probes for RNA imaging were designed similarly, except that each probe contained three identical copies of the readout sequence: one from the 5' end of the target region and two from the 3' end. The readout sequences for RNA imaging were orthogonal to those used for DNA imaging and were selected from the same ranked list of readout sequences tested.

[0212] Encoded probe synthesis. Encoded probes were amplified from the template library described above (see "Encoded Probe Design" above). This was done using an amplification protocol that included the following steps: 1. The initial oligopool was expanded using limited-cycle PCR over approximately 20 cycles. The reverse primers used in this step also introduced the T7 promoter sequence via primer extension. 2. The obtained product was purified by column chromatography and further amplified and converted to RNA by a high-yield in vitro transcription reaction. 3. The RNA product was converted back to single-stranded DNA by reverse transcription. 4. The product from the previous step was subjected to alkaline hydrolysis (to remove residual RNA and primer DNA) and then purified by column. 5. If necessary, the products from the previous steps were dried under reduced pressure and resuspended in water to obtain the desired concentration of the primary probe.

[0213] All primers were purchased from Integrated DNA Technologies (IDT).

[0214] Cell culture and hybridization of encoding probes. Cell preparation was performed as follows: IMR-90 cells were purchased from the American Type Culture Collection (ATCC, CCL-186) and grown according to the recommended protocol. To avoid potential alterations to chromatin structure, all cells in this study were seeded within 6 weeks of the start of culture at the densities specified below.

[0215] To prepare for DNA imaging, cells were seeded on 40 mm round #1.5 coverslips (Bioptechs, 0420-0323-2) at a density of approximately 500,000 cells per coverslip. Cells were grown at 37°C and 5% CO2 for approximately 2 days until density was reached. For transcription inhibition experiments, the cell medium was replaced with fresh medium containing 100 micrograms / mL of alpha-amanitin (Sigma-Aldrich, A2263) 6 hours before cell fixation. For experiments using 1,6-hexanediol (Sigma-Aldrich, 240117), the coverslips were coated with 10 micrograms / mL of fibronectin (Sigma-Aldrich, F1141) before cell seeding, and the medium was replaced with fresh medium containing 2% w / v of 1,6-hexanediol over 45 minutes. Next, the cultures were fixed with 4% paraformaldehyde (PFA) in PBS for 10 minutes at room temperature and washed 2-3 times in PBS. Then, the cells were permeabilized in two steps: first, they were treated with 0.5% v / v Triton-X (Sigma-Aldrich, T8787) in PBS for 10 minutes at room temperature. Next, the cells were treated with 0.1 M hydrochloric acid (HCl) for 5 minutes at room temperature and washed 2-3 times in PBS. After HCl treatment, the cells were treated with a solution of 0.1 mg / mL RNase A (ThermoFisher, EN0531) dissolved in PBS for 30-45 minutes at 37°C to remove potential sources of off-target binding to RNA. Following this treatment, the cells were incubated for approximately 10 minutes in a pre-hybridization buffer consisting of 2x saline-sodium citrate buffer (SSC; Ambion, AM9763) and 50% formamide (Ambion, AM9342). Next, the cell coverslips were inverted and placed on droplets of 50 microliters of hybridization buffer (containing a mixture of encoding probes at a total concentration of approximately 25 micromoles, with or without 10 micrograms of human Cot-1 DNA (ThermoFisher, 15279011), consisting of 2x SSC, 50% formamide, and 10% dextran sulfate (Sigma-Aldrich, D8906)) in a 60 mm Petri dish.The dishes were partially immersed in a water bath at approximately 90°C for 3 minutes and incubated in a humidified chamber at 47°C for 16–36 hours. After incubation with the encoding probe, the samples were washed in 2x SSC and 40% formamide for 30 minutes and post-fixed in 4% PFA in 2x SSC for 10 minutes at room temperature. The samples were then incubated with standard beads (either ThermoFisher F8805 or ThermoFisher F8792) in 2x SSC for 2–3 minutes, stained with 1 micromolar of 4',6-diamidino-2-phenylindole (DAPI; ThermoFisher D1306) in 2x SSC for 5–10 minutes, and stored in 2x SSC until imaging.

[0216] For experiments including RNA imaging, all buffers were used starting from cells fixed with 1:10 to 1:1,000 dilutions of either an RNAse inhibitor (NEB M0314 or Fisher Scientific N2615). The procedure for RNA staining was the same as the protocol above up to the HCl treatment. After this step, the cells were incubated in pre-hybridization buffer for 10 minutes, then the cell coverslips were inverted and placed on droplets of hybridization buffer containing an RNA intron-targeting encoding probe at a total concentration of approximately 1 micromol, as described for DNA staining. In this case, however, thermal denaturation at 90°C was not performed, and the cells were immediately incubated at 47°C for 16–36 hours in a humidified chamber. After incubation with the encoding probe, the samples were washed in formamide solution and post-fixed with PFA as described above for DNA. They were then incubated with standard beads, stained with 1 micromol of DAPI, and stored in 2x SSCs until imaging. After RNA imaging, the sample was removed from the microscope, the cells were treated with RNase A, and then DNA hybridization was initiated in the same manner as described above for DNA imaging, without using RNA imaging.

[0217] Sequential hybridization of readout probes for FISH imaging. All fluid exchange in this part of the protocol was achieved by using a custom fluid system with coverslip mounted in an FCS2 flow chamber (Bioptechs, 060319-2). The fluid system used 3-4 computer-controlled eight-way valves (Hamilton, MVP and HVXM8-5) and a computer-controlled peristaltic pump (Gilson, MINIPLUS3). Together, these components allowed control of both the velocity and type of fluid flow at any given time.

[0218] The hybridization in each round used the following steps: 1. Flow in a hybridization buffer with a set of oligonucleotide probes specific to each round, as described below. 2. Incubate at room temperature for 10 minutes. 3. Pour out the cleaning buffer. 4. Incubate for approximately 200 seconds. 5. Pour in the imaging buffer.

[0219] An imaging buffer was prepared containing 60 mM Tris pH 8.0, 10% w / v glucose, a 1% glucose oxidase oxygen scavenger solution (containing approximately 100 mg / mL of glucose oxidase (Sigma-Aldrich, G2133) and a 1:3 dilution of catalase (Sigma-Aldrich, C3155)), 0.5 mg / mL of 6-hydroxy-2,5,7,8-tetramethylchroman-2-carboxylic acid (Trolox; Sigma-Aldrich, 238813), and 50 micromoles of Trolox quinone (produced by UV irradiation of the Trolox solution). Trolox was dissolved in methanol and then added to the solution. After preparation, the imaging buffer was covered with a layer of mineral oil approximately 0.5 cm thick to prevent exposure to oxygen.

[0220] The hybridization buffer and washing buffer consisted of 35% and 30% formamide in 2x SSCs, respectively, and the hybridization buffer also contained 0.01% v / v Triton-X. The hybridization buffer was kept separately for each hybridization round and contained 2 or 3 sets of readout probes (for DNA and RNA imaging, respectively). Fluorescence signals were introduced using one of two methods:

[0221] 1. For DNA imaging, each round of hybridization buffer contained two fluorescent readout probes, one labeled with Cy5 or Alexa647 and the other with Alexa750. The fluorescent readout probes were either 1) a fluorescently labeled oligo complementary to the readout sequence common to all encoding probes imaged in a given bit, added at a concentration of 100 nM, or 2) a combination of an adapter oligo having a sequence complementary to the readout sequence, which is linked to a further readout sequence (referred to as a secondary readout sequence) common to all adapters (more precisely, common to all adapters in each color channel) and orthogonal to all other readout sequences used, and a fluorescently labeled oligo probe complementary to this secondary readout sequence. The adapters and secondary readout probes were pre-mixed in a ratio of 1:1.5 and added at a final concentration of approximately 100 nM. For several experiments, the adapters and readout probes were sequentially hybridized to the sample. This enabled the use of lower-concentration, more expensive secondary readout probes.

[0222] 2. For RNA imaging, each round of hybridization buffer contained three adapter oligos (detectable in three different color channels), each bound to a different readout sequence and each containing a further secondary readout sequence. All adapters corresponding to the same color channel shared the same secondary readout sequence. Each round involved two separate hybridization steps: first, the adapters were hybridized and flushed to wash away excess material. Then, three fluorescent readout probes, each labeled with Cy3, Cy5, and Alexa750 and complementary to the secondary readout sequences on the adapters, were flushed sequentially. The fluorescent readout probes used for RNA imaging contained disulfide bonds linking the fluorophores to the secondary oligos, allowing for efficient signal removal between rounds. After hybridizing the fluorescent readouts, the imaging buffer was flushed to collect the signal.

[0223] Prior to the next round of readout probe or adapter probe hybridization, the fluorescence signal from the current round's readout probe or secondary readout probe was removed as described in "Signal Removal Between Hybridization Rounds" below.

[0224] Prior to the first round of hybridization, an imaging round was performed to acquire DAPI signals and identify nuclear boundaries. The entire set of 1,041 genomic loci was then imaged in 50 rounds of hybridization with two color channels per round. In each round, the genomic loci were imaged in 3D by stepping through stages in the z-dimension. Similarly, nascent RNA transcripts for 1,137 genes were imaged in 3D in 3 colors over 18 rounds. Further rounds were used to relabel the set of genomic loci and evaluate chromatic aberration and pass-through between color channels. Imaging approximately 60 fields of view containing a total of approximately 1,000–2,000 cells took approximately 3 days.

[0225] The 3-4 valve system allowed for the loading of up to 20-30 different hybridization solutions. As a result, after exhausting all the channels in the fluid system, the sample chamber was bypassed and all channels used for hybridization were washed with 30% formamide in water. The chamber was then reconnected to perform the next set of hybridization and imaging rounds.

[0226] Antibody labeling and imaging. Antibody imaging was performed immediately after RNA or DNA imaging. After the completion of imaging according to the protocol described above, the sample underwent the following steps: 1. The samples were incubated for 30 minutes with a blocking solution (PBS containing 0.1% v / v Tween-20 (Sigma-Aldrich P9416) and 1% w / v bovine serum albumin (BSA; Jackson Immunoresearch 001-000-162)). 2. The sample was incubated with the primary antibody diluted in blocking solution for 1 hour. 3. The samples were washed three times in PBS containing 0.1% Tween-20 for 5 minutes each time. 4. Steps 2 and 3 were repeated for fluorescently tagged secondary antibodies.

[0227] All buffer exchanges were performed on a microscope using the microfluidic system described above. A Cy5 color channel was used for imaging, and photobleaching was used to eliminate signals between sequential antibody labelings.

[0228] The following sets of primary and secondary antibodies were used: 1. To image nuclear speckles, we used a primary antibody against SC35, a splicing factor commonly used as a marker for nuclear speckles (Abcam, ab11826), at a 1:200 dilution from stock, and a donkey anti-mouse secondary antibody labeled with Cy5 dye (Jackson Immunoresearch, 715-175-150), at a 1:1,000 dilution from stock concentration. 2. To image the nucleoli, we used an anti-fibrillarin antibody (Abcam, ab5821) diluted 1:200 from stock, and a donkey anti-rabbit secondary antibody (Jackson Immunoresearch, 711-605-152) labeled with Alexa657 dye, diluted 1:1,000 from stock concentration.

[0229] Signal removal between hybridization rounds. Before imaging each round, the signal from the previous round (or, in the case of the first round, the endogenous background) was eliminated. This was achieved by photobleaching of the signal. Photobleaching was performed by changing the buffer to 2xSSC and irradiating each field for 10 seconds at the maximum available power of the 647 and 750 lasers (as well as the 560 laser when imaging RNA). In RNA imaging experiments, the buffer used for bleaching also contained 50 mM tris(2-carboxyethyl)phosphine (TCEP; Sigma-Aldrich, C4706) to cleave the disulfide bonds connecting the fluorophores to the readout probe. As a result of the high formamide concentration in the hybridization and washing buffers, the DAPI signal was eliminated.

[0230] Image acquisition. Image acquisition was performed using a custom-made microscope system. The system was constructed around a Nikon Ti-U microscope body, along with a Nikon CFI Plan Apo Lambda 60x oil immersion objective lens with a numerical aperture of 1.4. Irradiation was based on one of two options: 1. Solid-state single-mode lasers with the following wavelengths: 405nm (Coherent, Obis 405nm LX 200mW), 560nm (MPB Communications, 2RU-VFL-P-2000-560-B1R), 647nm (MPB Communication, 2RU-VFL-P-1500-647-B1R), and 750nm (MPB Communication, 2RU-VFL-P-500-750-B1R). In this case, the output of the 560nm, 647nm, and 750nm lasers was controlled by an acousto-optic tunable filter (AOTF), while the 405nm laser was directly controlled by its laser control box. Custom dichroic filters (Chroma, zy405 / 488 / 561 / 647 / 752RP-UF1) and emission filters (Chroma, ZET405 / 488 / 461 / 647-656 / 752m) were used to separate the excitation and emission irradiation. 2. The following wavelengths were used: Lumencor CELESTA optical engine (a solid-state laser-based irradiation system bonded to a fiber). This system was used in conjunction with a pentabandpass dichroic (IDEX, FF421 / 491 / 567 / 659 / 776-Di01-25x36) and a pentabandpass filter (IDEX, FF01-441 / 511 / 593 / 684 / 817-25).

[0231] Academic CMOS cameras (Hamamatsu FLASH4.0 or Hamamatsu C13440, factory calibrated for single-molecule imaging) were used for image acquisition. The three-dimensional sample position was controlled using an XYZ stage (Ludl). A custom autofocus system was used to maintain a constant focal plane over extended periods. This was achieved by comparing the relative positions of two IR laser beams (Thorlabs, LP980-SF15) reflected from the glass-liquid interface, and imaging them on separate CMOS cameras (Thorlabs, uc480).

[0232] For each experiment, approximately 60 fields of view (FOV) were selected for imaging, avoiding areas with sparse cells (the inventors typically identified 10 to 50 cells per FOV). Each camera's FOV had either 1,000x1,000 pixels, with camera pixels corresponding to 153 nm in each dimension of the imaging plane, or 2,048x2,048 pixels, with camera pixels corresponding to 108 nm in each dimension of the imaging plane.

[0233] After hybridization in each round (see "Sequential Hybridization of Readout Probes for FISH Imaging" above), z-stack images of each FOV were acquired in 3 or 4 colors: FISH images were acquired using 647 nm and 750 nm irradiation (or 560 nm, 647 nm, and 750 nm irradiation for RNA imaging when DNA and RNA imaging were combined), and standard beads were imaged using 560 nm irradiation (or 405 nm irradiation when RNA and DNA imaging were combined). For imaging in the first round, the DAPI signal was imaged using 405 nm irradiation, while for antibody imaging, the 647 nm excitation channel was used after RNA or DNA imaging. Serial z-sections were separated by 85, 100, or 150 nm, encompassing the entire nuclear volume of all imaged cells. After acquiring images in all channels at each z-position, the stage was moved and images were acquired at a speed of approximately 10 Hz.

[0234] Image analysis and spot fitting for DNA and RNA imaging. The following analysis pipeline was applied to each imaged FOV to obtain the three-dimensional (3D) location of all target gene loci: 1. The criteria were applied to all rounds of imaging and used for image alignment. 2. In the first imaging round (preceding the hybridization in the first round), DAPI signals were used to identify the boundaries of individual nuclei and for image registration between RNA and DNA imaging. 3. The diffraction-limited spots within each identified nucleus were fitted to a 3D Gaussian function to identify their centers of mass and brightness above the local background. 4. The fitted spots were compared to other localizations within the same nucleus across all rounds of hybridization, and the loci from which they arose were identified using custom algorithms and software (detailed in the sections “Decoding Algorithms for Fitted DNA Spots” and “Decoding Algorithms for Fitted RNA Spots”).

[0235] Spot fitting for DNA and RNA imaging. Signals from individual FISH imaging rounds were fitted using a 3D Gaussian distribution. To make the analysis more manageable, the number of fitted spots per image retained for decoding was fixed at 125 (approximately 3 times more than the number of different loci expected to be noise-free).

[0236] Drift correction. Standard bead spot fitting was performed in the same manner as above. Then, the set of standard bead positions was compared between hybridization rounds, and rigorous transformation was applied to minimize the sum of the squared differences in the relative positions of the beads.

[0237] Color effect correction. Pass-through and chromatic aberration in multicolor imaging were corrected by independently labeling the same set of genomic loci in each imaging channel and comparing the signals of the same loci in different channels.

[0238] Nuclear segmentation. Using DAPI images from the first round of imaging, we identified the volume of individual nuclei and enabled cell segmentation. This was achieved using a convolutional neural network, which was built and trained, taking the maximum projection of DAPI images onto the xy plane as input.

[0239] Further image analysis: Image registration between DNA and RNA imaging. In experiments involving imaging of both DNA and RNA, DAPI signals were initially used for rough image registration (to camera pixel precision) across two sets of images using 2D image correction (all images in each set were aligned to the DAPI image using standard beads). After an initial round of RNA decoding (see "Decoding Algorithm for Fitted RNA Spots" below), finer alignments were calculated by assuming that the displacement between the nascent RNA localization and the DNA locus in which they reside should, on average, be zero when considered across all imaged genes and cells in a given field of view. Therefore, further rigorous transformations were calculated to minimize the mean displacement between the imaged nascent RNA and its corresponding DNA locus, and this was used as the final alignment.

[0240] Identification of intranuclear structures from immunofluorescence imaging. The locations of intranuclear structures (nuclear speckles and nucleoli) were extracted from immunofluorescence signals by applying a threshold to the intensity of the immunofluorescence signal to obtain a pixelated mask that identifies high immunofluorescence signals. This was then processed as the locations of a pixelated set "containing" the intranuclear structures.

[0241] Decoding algorithm for fitted DNA spots. Identification and 3D localization of each genomic locus were achieved through the following steps: 1. A list was generated showing the drift-corrected and aberration-corrected positions of all identified spots in each round of imaging. 2. For each detected spot in all imaging rounds, it was found that all spots originating from other rounds were within a set cutoff distance (approximately 150 nm in x, y, and z) from its location. All such spot pairs were retained for further analysis, regardless of whether the barcodes produced by the spot pairs (based on the round and color channel in which they appear) corresponded to legitimate barcodes (barcodes assigned to genomic loci). 3. For each spot pair, three quality metrics were calculated: A. Displacement between the 3D localization of the two spots B. Difference in brightness between the two spots C. The average brightness of the two spots. The brightness of each spot was normalized by the median brightness of all spots in the corresponding bit. 4. Next, the spot pairs were separated into two groups based on whether they corresponded to valid barcodes (and therefore potentially genomic loci). Within each group, the distribution of quality metrics was calculated. For convenience, the distribution of spot pair quality metrics from invalid barcodes was referred to as the "invalid distribution," and the distribution from all valid barcodes was referred to as the "valid distribution." 5. For each spot pair, the three quality metrics from step 3 were combined into a single measure by calculating the combined Fisher p-value for all candidate spot pairs relative to the “legitimate distribution” in step 4. This was considered the overall quality score for each spot pair and was calculated per pair as follows: For each of the three metrics, the ratio of other spot pairs in the “legitimate distribution” was calculated using the lower quality metric, and these ratios were multiplied together. Then, using an expectation maximization procedure, the two spot pairs with the highest quality scores corresponding to each targeted chromatin locus were sequentially selected to re-update the “legitimate distribution,” and this optimization procedure was repeated until convergence. After convergence, the 3D spatial location of the locus was determined using the final set of spot pairs, each corresponding to a chromatin locus. 6. Following step 5, a modified K-means algorithm was used to separate chromatin loci belonging to the same chromosome into two homologs. Contrary to the standard K-means clustering algorithm, which divides points into two groups and minimizes the radius of rotation within each group, points were progressively switched between groups to first maximize the ratio of assigned points in each homolog, and then minimize the radius of rotation of each homolog. 7. After separating the two homologs, the center of mass and distance of each spot pair from Step 2 were calculated relative to the center of mass of the parent chromosome. The distance to the chromosome center was added as another quality metric in addition to the three metrics considered in Step 3, and Steps 3-6 were repeated. 8. Finally, spot pairs from process 7 were filtered to remove pairs whose quality scores still resembled the “invalid distribution”. The remaining spot pairs after step 8 were used to determine the final location of the chromatin locus and track the chromatin structure.

[0242] Decoding algorithm for adapted RNA spots. Signals from RNA imaging rounds were decoded using the following procedure. 1. A list was generated showing the drift-corrected and aberration-corrected positions of all identified spots in each round of imaging. 2. For each detected spot in all imaging rounds, it was found that all spots from other rounds were within a set cutoff distance from its location, and these spot pairs were retained as candidate RNA bursts if they formed a valid barcode. 3. Next, the location of each of these candidate RNA bursts was compared to the location of the DNA locus containing the associated gene after initial image registration (based on DAPI images) and drift and aberration correction. If they were within a set threshold distance, they were retained. 4. The registration between DNA imaging and RNA imaging was refined based on the displacement between the initial decoded RNA localization (from step 3) and the location of the DNA locus containing them, as described in the section "Image Registration Between DNA Imaging and RNA Imaging" above. 5. The locations of all candidate RNA bursts were re-compared to the locations of the DNA loci containing the genes they decode, using the refined image registry used in this study. If the nascent RNA localization was within the cutoff distance from the DNA locus at this stage, it was considered a detected transcription burst.

[0243] Further analysis: Identification of nuclear lamina. The location of nuclear lamina was estimated by generating minimal 3D convex hulls around the location of all decoded chromatin loci in a given cell (using the scipy package in Python).

[0244] Spatial distance. The spatial distance between any pair of gene loci was simply calculated as the Euclidean distance between their fitted 3D Gaussian distribution centers, multiplied by an appropriate ratio of camera pixels and z-step to physical distance. For distances to nuclear structures, the minimum Euclidean distance to the "location" of all identified nuclear structures or the minimum distance to the surface of the convex hull defining the nuclear lamina was calculated.

[0245] Contact frequency matrix from imaging. To calculate the contact frequency between any given pair of loci, the number of measured distances between that locus pair that were less than a set threshold was counted. This number was then divided by the total number of measured distances for that locus pair.

[0246] Local density analysis. To calculate the trans chromosome local density of compartment A and compartment B loci at each decoded location, the spatial distance between each pair of chromatin loci for each cell was calculated. For each locus, the local A / B density ratio was calculated as follows: 1. The density contribution of each other locus was calculated from different chromosomes by evaluating the Gaussian function value using a standard deviation of 500 nm (adjusted to account for variations in cell size) of the distance between the two loci. 2. Next, the total A density at each locus was calculated as the sum of Gaussian function values ​​for all trans chromosome A loci, and the total B density was calculated in the same manner. 3. The total density of the A locus in trans chromosome compartment A was divided by the density of the B locus in trans chromosome compartment to determine the A / B density ratio at that locus.

[0247] Estimation of detection efficiency in multiplexed RNA imaging. The detection efficiency of transcriptional burst events was estimated as follows: 1. We considered all target genomic loci that have genes whose RNA introns are being imaged. For any of these genomic loci, the corresponding RNA signal appeared in two predefined bits when the gene was transcribed. Knowing the rate (p) at which each of these two bits was not detected allowed us to induce RNA detection efficiency. We identified a set of genomic loci that localized (within approximately 150 nm) along with the RNA signal to at least one of the two expected bits of the corresponding gene. 2. From the total set of chromatin loci identified in step 1, the ratio (f) of loci colocalizing with RNA signals was determined from exactly one of the corresponding bits (not both bits) of the gene.

[0248]

number

[0249]

number

[0250] Hi-C data analysis. Hi-C data for IMR-90 cells were obtained and loaded using straw. Established and published protocols were followed for the identification of A / B compartments in individual chromosomes. To compare contact frequencies derived from imaging data with Hi-C, bins centered on target regions were created, and Hi-C data for these bins were obtained by summing the number of reads in higher-resolution Hi-C data. [Example 6]

[0251] The three-dimensional (3D) structure of chromatin regulates many genomic functions. However, understanding 3D genomic structure is hindered by the lack of means to directly visualize chromatin conformation in its native context. Reported herein is an imaging platform for visualizing chromatin structure across multiple scales in a single cell with high genome throughput. First, we demonstrate multiplexed imaging of hundreds of genomic loci by sequential hybridization, enabling high-resolution conformational tracing of the entire chromosome. Next, we develop a combinatorial imaging method for genome-scale chromatin tracing, demonstrating simultaneous imaging of over 1000 genomic loci and over 1000 nascent transcripts of genes, along with landmark nuclear structures. Using this platform, we characterize chromatin domains, compartments, and trans-chromosome interactions and their relationship to transcription in a single cell. This highly efficient, multi-scale, and multi-mode imaging technique has broad applications, providing an integrated view of chromatin structure in its native structural and functional context.

[0252] The 3D structure of the genome regulates many essential cellular functions, from gene expression to DNA replication. Biochemical and imaging measurements have elucidated the complex chromatin structure at various scales. In particular, highly efficient chromosome conformational capture methods, such as Hi-C and other sequencing-based methods, have revealed chromatin structures such as domains and compartments, along with a genome-wide view. Chromatin is divided into self-interacting genomic regions called topologically related domains (TADs), which appear as block-like structures on Hi-C contact maps. These TADs, ranging in size from several hundred kilobases (kb) to several megabases (Mb), often contain genes that are simultaneously regulated and have boundaries that coincide with regulatory epigenetic elements. On a larger scale, chromatin is divided into two main compartments, called A and B compartments, which are enhanced for active and inactive chromatin, respectively, as indicated by alternating "grid" patterns in Hi-C maps, consistent with previous imaging-based observations that chromatin segments with high and low gene content tend to be spatially separated. Recent imaging experiments have shown that chromatin of compartment A and compartment B does indeed tend to be spatially separated within a single cell. The physiological significance of A / B compartmentalization is implied by its changes during development and between cell types.

[0253] Overall, highly efficient sequencing-based methods have greatly enhanced our knowledge of 3D genome architecture. Nevertheless, these powerful methods also have limitations. For example, these methods provide contact information about pairs of chromatin loci, but not direct spatial location information about these loci. Furthermore, much of the insight into chromatin architecture is built on population-averaged contact maps across millions of cells. Despite the continuous improvements in single-cell Hi-C methods, the capture efficiency of chromatin contacts in single cells and / or the cell throughput of these methods remain low, and therefore, scrutinizing 3D genome architecture in single cells remains a challenging task. Moreover, while methods have emerged that combine Hi-C with other measurement modalities to provide characterization of chromatin contacts in the context of interacting proteins, nuclear structure, or DNA modification, multimode measurements by sequencing remain challenging. In particular, methods that enable the measurement of both chromatin composition and transcriptional activity at the genome scale within the same cell have not emerged despite the demand for such methods to further understand how chromatin composition regulates transcription, and subsequently how transcription influences chromatin composition.

[0254] On the other hand, imaging-based methods offer a direct measure of the spatial location of chromatin loci in individual cells, along with high detection efficiency. In particular, fluorescence in-situ hybridization (FISH) enables highly specific detection of chromatin loci in fixed cells, and more recently, the CRISPR (clustered regularly interspersed short palindromic repeat) system has substantially enhanced the ability to image specific chromatin loci in living cells. Chromatin imaging can also be combined with RNA and protein imaging to examine chromatin composition and its interactions with transcriptional activity or interacting protein factors. However, current imaging methods limit throughput in genomic (sequence) space, traditionally allowing only a small number of different genomic loci to be tested at a time. Recently, chromatin tracing techniques using sequential rounds of FISH imaging, where each round targets one or two genomic loci using one- or two-color imaging, have been developed. This technique has been used to enable imaging of dozens of different chromatin loci within a single cell, providing insights into chromatin structure and its relationship to transcription. However, because the number of genomic loci that can be imaged simultaneously within an individual cell remains limited, a high-resolution view of the entire chromosome in a single cell, let alone a genome-scale view of chromatin composition within an individual cell, is still not available.

[0255] A multiscale, multiplexed FISH imaging platform enabling simultaneous imaging of hundreds to over 1,000 different genomic loci in a single cell at varying resolutions and genomic coverages is reported herein. First, a sequential imaging technique was substantially evolved to enable imaging of hundreds of genomic loci, and this method was applied to provide a high-resolution view of the entire chromosome, elucidating chromatin domain and compartment structure in single cells, their relationships to each other, and the relationship between chromatin composition and transcription. Next, a large-scale multiplexed FISH technique was developed based on combinatorial labeling and imaging, enabling imaging of many more genomic loci using far fewer hybridization rounds. Using this technique, simultaneous imaging of over 1,000 genomic loci in individual cells, as well as simultaneous imaging of these loci together with nascent RNA transcripts of over 1,000 genes present in landmark nuclear structures, including nuclear speckles and nucleoli, allowing chromatin composition to be placed in its native structural and functional context. Using this method, we explored the relationships between trans chromosome interactions, transcriptional activity, and nuclear structure in single cells.

[0256] To enable a systematic view of chromatin structure across multiple scales, we developed an imaging platform using a custom-built microscope and fluid engineering setup (see Example 19) for exceptionally efficient direct visualization of chromatin in sequence space down to the genome scale. This platform included two complementary techniques (Figure 17A). First, for imaging chromatin structures that were relatively small, such that the different loci contained within were difficult to resolve in any single image, we extended a previously reported sequential imaging strategy to enable tracing of hundreds of chromatin loci in a single cell. This technique imaged chromatin one locus at a time (or two to three loci at a time using two- or three-color imaging) over many imaging rounds (Figure 17A, left). This technique was demonstrated by using it to track the conformation of the entire chromosome in a single cell at high resolution. Secondly, to image chromatin structures dispersed over areas substantially larger than diffraction-limited resolution, such as structures extending throughout the entire nucleus, we developed a more efficient combinatorial strategy in which many chromatin loci are imaged simultaneously in each round, and their different identification is determined based on the different combinations of rounds in which they appear (Figure 17A, right). This latter approach enabled imaging of numerous genomic loci in far fewer imaging rounds. Using this method, we provided a genome-scale view of chromatin structures in single cells, in the context of transcriptional activity and important nuclear structures.

[0257] Figures 17A–17M show high-resolution whole-chromosome tracing by sequential hybridization and characterization of chromatin domains in single cells. Figure 17A shows a schematic diagram of a multiscale chromatin tracing platform. Left: Schematic diagram of whole-chromosome chromatin tracing by sequential hybridization and imaging. When the target chromatin structure is equivalent to or less than diffraction-limited resolution, a single chromatin locus is imaged in each color channel per imaging round. After all rounds of imaging, a chromatin trace can be generated in 3D for each copy of the target chromosome. Right: Schematic diagram of genome-scale imaging by combinatorial FISH. When the target locus is expected to be spread across a space substantially larger than diffraction-limited resolution, such as when the locus is dispersed throughout the entire nucleus, multiple loci can be imaged and degraded in each round, and the identification of each locus can be derived from a barcode based on the combination of imaging rounds in which the locus is detected. This method significantly reduces the number of rounds required to image the same number of gene loci compared to sequential imaging methods. [Example 7]

[0258] High-resolution chromatin tracing of the entire chromosome. This section describes high-resolution whole-chromosome tracing using sequential imaging techniques (Figure 17A, left; Figure 24A). We first focused on human chromosome 21 (Chr21) and divided the non-repeatable portion of the chromosome (Chr21: 10.4–46.7 Mb) into more than 600 continuous segments (i.e., more than 600 genomic loci) each 50 kb in length. We designed a library of primary oligonucleotide probes, each containing a variable target sequence for hybridizing to the chromosome and a readout sequence unique to each 50 kb locus (Figure 24A). All primary probes bound to a particular 50 kb locus shared the same readout sequence, and therefore, each locus could be identified by hybridization of a fluorescently labeled complementary readout probe using that readout sequence (Figure 24A). However, identifying these genomic loci using over 600 different fluorescent readout probes is prohibitively expensive due to the high cost of dye-labeled oligonucleotides. To overcome this challenge, we devised a two-step labeling strategy to detect different readout sequences using a common set of three dye-labeled oligonucleotide probes (called readout probes, one for each color channel) mediated by an unlabeled adapter probe that converts each locus-specific readout sequence to one of three common readout sequences (Figure 24A). Using this strategy, we sequentially imaged over 600 chromatin loci in Chr21 in human lung fibroblasts (IMR-90) using over 200 rounds of hybridization of adapters and readout probes, along with three-color imaging in each round.To enable stable imaging over such a large number of hybridization rounds, the imaging protocol was further optimized as follows: (i) the samples were re-fixed with formaldehyde periodically during imaging to maintain sample integrity and the binding stability of the primary probes; (ii) a combination of chemical cleavage and photobleaching techniques was used to remove the fluorescence signal of the readout probes and unlabeled readout probes were added to block the unoccupied binding sites on the adapter probes after each round of imaging, ensuring complete removal of the fluorescence signal after each imaging round and minimizing the accumulation of residual signals over hundreds of labeling rounds; (iii) the duration of each hybridization round and the flow rate of the fluid system were optimized to minimize perturbation and experimental time.

[0259] After imaging, the center of mass of each chromatin locus was determined in 3D, and the conformation of each homologous copy of Chr21 in each cell was reconstructed (Figure 17B). To estimate how stable the samples and imaging instruments were across many hybridization rounds, the same chromatin locus was re-imaged after a different number of hybridization / imaging rounds, and the displacement between the original locus and that of the corresponding re-imaged locus was used as a measure of measurement accuracy. The median displacement between the original and re-imaged loci increased from approximately 70 nm when 11 rounds of hybridization separated the two imaging examples to approximately 120 nm when the initial and re-imaging examples were separated by approximately 250 rounds of hybridization (Figures 24B-24C), and loci showing greater displacement during re-imaging also had lower fluorescence signal intensity (Figure 24D). The median displacement error was found to be substantially smaller than the median distance between neighboring chromatin loci (approximately 250 nm), even after more than 250 rounds of hybridization (Figure 24B). Furthermore, the median pairwise distances between imaged loci were highly reproducible across biological replicates (Figure 24E). The detection efficiency of chromatin loci in these experiments exceeded 90% (i.e., more than 90% of target chromatin loci were detected in each chromosome).

[0260] To obtain a population-average view of the chromatin conformation of Chr21, pairwise interactions between imaged loci were quantified over approximately 3,500 imaged cells by calculating the median spatial distance and the probability of the loci coming close together (Figs. 17C and 24F–24I). High correlations (Pearson correlation of 0.89; Figs. 24F–24G) were observed between the median pairwise distances from the imaging data and previously published Hi-C data across all length scales present in Chr21, consistent with particularly high correlations for shorter genomic distances (Pearson correlation of 0.97; Figs. 24H–24I). To select a cutoff distance below which two loci were considered to be in proximity, Pearson correlation coefficients were calculated between the Hi-C data and the proximity frequencies (i.e., the ratio of examples where two genes were in proximity) derived from the imaging data over a range of cutoff distances. The Pearson correlation coefficient remained high for a wide range of cutoff distances but peaked at 0.88 when the cutoff distance was approximately 400–500 nm (Fig. 24J). Thus, 500 nm was selected as the cutoff distance to generate the proximity frequency map through this operation (see Example 19 for more detailed rationale regarding the selection of the cutoff distance).

[0261] Both the median distance and the proximity frequency map showed block-like TAD structures (Figs. 17C and 24F and 24H). TAD boundaries identified from both the distance and the proximity frequency map from the imaging data were very similar to those determined from the ensemble Hi-C data (Fig. 24K). Furthermore, it was confirmed that the locus localization specific error (approximately 100 nm) in chromatin trace had little effect on the identification and accuracy of domain boundaries (Fig. 24K).

[0262] Figure 17B shows 3D structural renderings and spatial distance matrices of two copies of Chr21 in a single IMR90 cell, imaged using a sequential hybridization technique. Left: Two copies of Chr21 in a single cell superimposed on a nuclear DAPI image. Scale bar is 5 micrometers. Top right: 3D renderings of all detected chromatin loci (spheres) in the two Chr21 copies by their genomic candidate between chromosomes (genomic locations shown on the right). Flexible lines connect adjacent chromatin loci. Scale bar is 1 micrometer. Bottom right: Pairwise spatial distance matrix corresponding to the chromosomal copies shown above (genomic regions that do not contain a suitable reference genome or contain highly repetitive sequences are not imaged).

[0263] Figure 17C shows the ensemble proximity frequency matrix for Chr21 and the preferred arrangement of single-cell domain boundaries at the CTCF / RAD21 binding site. Top: Proximity frequency matrix for Chr21 derived from imaging data. Each matrix element is defined as the frequency in which the measured distance between locus pairs is shorter than the 500 nm cutoff distance. Center: Zoomed-in version of the proximity frequency matrix for a 10 Mb portion of Chr21. Bottom: Probability of single-cell domain boundary formation in each of the imaged 50 kb genomic segments. Triangles indicate ChIP-seq peaks for CTCF and RAD21.

[0264] Figures 24A–24N show ensemble statistics of Chr21 structural characteristics compared to high-resolution whole-chromosome tracing by sequential hybridization and Hi-C. Figure 24A shows the labeling and imaging scheme for sequential hybridization with adapter probes. First, a sample containing target sequences and readout sequences, respectively, enabling specific binding to target genomic loci, is hybridized with a primary probe. Each locus is labeled with a total of 350–500 primary probes, but only one is shown. Each target genomic locus is assigned a unique readout sequence (shown in various colors) common to all primary probes that bind to the locus. The readout sequences are then detected using sequential round hybridization. During each round of hybridization, a readout sequence corresponding to the target locus (one for each of the three color channels: Alexa750, Alexa647, and Cy3) is labeled with an oligonucleotide adapter probe consisting of two segments: a segment complementary to the locus-specific readout sequence and a segment containing a color channel-specific common readout sequence. Each color channel contains a unique common readout sequence shared by all adapters visualized in the same color channel. The common readout sequence is then hybridized to a complementary readout probe conjugated with a dye in the corresponding color channel. This procedure allows imaging of three genomic loci in three color channels during each round of hybridization. After imaging of each round, the fluorescent dye attached to the readout probe by a disulfide bond is cleaved from the common readout probe by TCEP, and the unoccupied readout sequence on the adapter is blocked with the unlabeled common readout probe to prevent crosstalk between hybridization rounds. This process is repeated hundreds of times until all readout sequences, and therefore all genomic loci, have been detected.

[0265] Figure 24B shows the displacement of a gene locus over the course of a single experiment. Contiguous 50kb segments within a 900kb region of Chr21 (chr21: 32.45–33.35 Mb) were imaged at both the beginning and end of the experiment, separated by more than 250 rounds of hybridization. The distribution of displacement between the re-imaged spots and their original imaged counterparts is shown. For comparison, the distribution of distances between adjacent 50kb segments within the same 900kb region, measured in the original imaging rounds, is shown.

[0266] Figure 24C shows box plots of chromatin locus displacements between the original imaging round and the re-imaging round, separated by a different number of hybridization rounds. The median (centerline), 25th–75th percentile (box), and 10th–90th percentile (whiskers) are shown.

[0267] Figure 24D shows box plots of fluorescence signals at chromatin loci exhibiting low (<500 nm) and high (<500 nm) displacement errors between the original imaging and re-imaging experiments, separated by more than 250 rounds of hybridization. The median (centerline), 25–75th percentile (box), and 10–90th percentile (whiskers) are shown.

[0268] Figure 24E shows a comparison of median interlocusal distances between two replicate experiments. Median interlocusal distances between pairs of imaged chromatin loci were calculated separately for two biological replicate experiments of Chr21 and plotted relative to each other. The Pearson correlation coefficient between the data measured in the two replicates is ρ = 0.98.

[0269] Figure 24F shows a comparison of the imaging-derived spatial distance median matrix (left), imaging-derived proximity frequency matrix (center), and ensemble Hi-C contact matrix for Chr21 (right). For imaging data, two chromatin loci are considered close if the spatial distance between the two loci is less than a 500 nm cutoff distance. The Hi-C contact matrix is ​​binned at 50 kb and centered on the target region.

[0270] Figure 24G shows a log-log scatter plot of the number of contacts derived from ensemble Hi-C and the median pairwise distance derived from imaging of individual pairs of chromatin loci. The straight line represents the linear regression of the data (slope = -4.43). The Pearson correlation coefficient between the imaging and the Hi-C data is ρ = 0.89.

[0271] Figure 24H is the same as Figure 24F, but for a 3Mb region within Chr21 (chr21: 30.30~33.38Mb). The TAD boundary is marked with a straight line.

[0272] Figure 24I is the same as Figure 24H, but for the 3Mb region shown in (H). The gradient is -4.51 and ρ = 0.97.

[0273] Figure 24J shows the Pearson correlation between the Hi-C contact map and the imaging-derived proximity frequency map generated using a varying cutoff distance. To generate the proximity frequency map, a cutoff distance was selected, and two loci with a distance smaller than this cutoff value were considered to be in close proximity. The proximity frequency between pairs of loci was then calculated by dividing the number of occurrences where the measured distance between loci was shorter than the cutoff distance by the total number of measured distances between the two loci.

[0274] Figure 24K shows the normalized isolation score as a function of genomic location on Chr21, calculated from 1) median pairwise distance from imaging (top), 2) proximity frequency from imaging (center), and 3) Hi-C contact reads (bottom). To calculate the isolation score, a fixed-length (250kb) genomic segment upstream and a segment of the same length downstream of the location of interest were first selected. The normalized isolation score is then defined as the difference between the median pairwise distance between segments and the median pairwise distance within a segment, normalized by the sum of these two median distances. The TAD boundary is defined as the local maximum of the normalized isolation score along the chromosome, identified by a standard peak calling algorithm (see Example 19). The vertical dotted line is the ensemble TAD boundary called from Hi-C data. Furthermore, after disturbing the locus locations with a 3D Gaussian noise term exhibiting a standard deviation of 100 nm, equivalent to the estimated localization measurement error, the median distance (black line, upper panel) and proximity frequency (black line, middle panel) are shown in the upper and middle panels.

[0275] Figure 24L shows the chromatin domain in the 10 Mb region of Chr21 (chr21: 28.2–38.1 Mb) in two exemplary single cells. The pairwise distances from two individual copies of Chr21 in a single cell (top, center) are shown together with the median population pairwise distance from all imaged cells (bottom). [Example 8]

[0276] Chromatin domains within a single chromosome. At the single-cell level, chromosomes were observed to be divided into domains that appear as block-like features in the single-cell spatial distance matrix (Figure 24L). These domains and interlocus distances showed high intercellular variability, consistent with substantial intercellular variability in chromatin contacts observed in single-cell Hi-C data (Figures 24L-24M). Similar domain structures within single cells have been previously observed when imaging small (approximately 2 Mb) regions of chromosomes at similar resolution. However, within these previously measured small regions, a significant proportion of cells did not exhibit clear single-cell domain boundaries, and it remained unclear whether domains formed within these cells or whether all imaged regions were within a single domain. Furthermore, due to the small size of these previously imaged regions, many domains were artificially truncated at the ends of the imaged genomic regions, thus hindering the accurate characterization of certain fundamental domain characteristics, such as their physical size and genomic size. The high genome throughput in this study provided a whole-chromosome view of these single-cell domain structures, essentially demonstrating their prominent presence across all chromosomes in all imaged cells, thus enabling a more systematic characterization of their properties.

[0277] We first identified the genomic locations of these single-cell domain boundaries and quantified the probability of boundary formation at each 50kb genomic locus. While a non-zero probability of boundary formation was observed at all imaged genomic loci, domain boundaries were preferentially located near CTCF and cohesin binding sites (Figures 17C–17D).

[0278] In addition to intercellular variability at the location of domain boundaries, substantial heterogeneity was observed in other features of these single-cell domains, ranging from the physical size of the domains to the degree of isolation or interaction between domains (Figures 17E-17H). Specifically, single-cell domains were observed to be variable in both their genomic size (Figure 17I) and their physical size as measured by the radius of gyration (Figures 17E and 17J). Neither the distribution of genomic size nor the distribution of physical size of these domains was sensitive to an estimated locus localization error of approximately 100 nm (Figures 17I-17J). In particular, domains bounded by the same genomic region or having the same genomic size varied considerably in their physical size between cells (Figures 17E and 24N). Interestingly, domains bounded by interacting CTCF / cohesin binding sites tended to be smaller in physical size than domains not bounded by such genomic loci (Figure 17K). In addition, the degree of physical separation between neighboring domains also varied substantially (Figures 17F and 17L), with some neighboring domains being completely separated and connected only by linker regions, while others showed partial overlap and less distinct boundaries (Figure 17F). Furthermore, even domains completely separated from their neighbors could partially overlap in space, while non-neighboring domains were separated by small or large genomic distances (Figure 17G). Finally, the two chromatin loci at the ends of these single-cell domains also showed varying distances from each other and did not tend to be closer to each other compared to chromatin loci separated by similar genomic distances within the domain, regardless of whether the domains were bounded by CTCF / cohesin binding sites (Figures 17H and 17M).

[0279] Figure 17D shows the average probability of domain boundary formation in a single cell at genomic locations centered around the CTCF / Rad21 binding site or ensemble TAD boundary (gray).

[0280] Figure 17E shows examples of two single-cell chromatin domains with the same genomic coordination that occupy large (top) or small (bottom) volumes in physical space. Left: A 3D rendering of the chromatin domain, where the green spheres represent the imaged genomic loci within the domain, and the flexible linkers connect adjacent loci in the genomic sequence. The gray spots represent the imaged loci in the remaining chromosomes. The scale bar is 1 micrometer. Right: A pairwise distance matrix (marked with lines) of the chromatin domains shown on the left, together with adjacent regions.

[0281] Figure 17F shows examples of two pairs of chromatin domains with high (top) and low (bottom) insulation scores. Left: A 3D rendering of the chromatin domain, similar to Figure 17E. The scale bar is 250 nm. Right: A pairwise distance matrix of the chromatin domains shown on the left, together with the rendered domains marked with corresponding colors.

[0282] Figure 17G shows two examples of long-range contacts between chromatin domains with partially overlapping volumes. Left: A 3D rendering of the chromatin domain, similar to (E). The shadows represent different domains. The scale bar is 250 nm. Right: A pairwise distance matrix of the chromatin domains, together with the rendered domains marked with corresponding colors. The gray space indicates a gap at a genomic distance of 22.85 Mb.

[0283] Figure 17H shows an example of a chromatin domain flanked by CTCF binding sites, showing small (top) and large (bottom) distances between the CTCF sites. Left: A 3D rendering of the chromatin domain, similar to (E), but with the loci of the CTCT sites at the ends of the domain. The scale bar is 250 nm. Right: A pairwise distance matrix of the chromatin domain with the marked boundary CTCF according to the domain.

[0284] Figure 17I shows the measured genome size distribution of chromatin domains in Chr21 in a single cell. The distribution of genome size of chromatin domains in Chr21 in a single cell, derived from simulated data that accounts for a 100 nm localization error, is shown by the black line. In this simulation, the location of the imaged gene loci is disturbed with 3D Gaussian noise with a standard deviation of 100 nm, similar to our measurement error.

[0285] Figure 17J shows the measured physical size distribution of chromatin domains in Chr21 in a single cell, defined by the radius of gyration. Similar to Figure 17I, the distribution of physical size of chromatin domains in Chr21 in a single cell, derived from simulated data considering a 100 nm localization error, is shown by the black line.

[0286] Figure 17K shows the median radius of gyration as a function of genome size for chromatin domains containing and excluding boundary loci containing the interacting CTCF / Rad21 site. Error bars indicate 95% confidence intervals induced by resampling.

[0287] Figure 17L shows the distribution of isolation scores between neighboring domains with domain boundaries present at CTCF / Rad21 junction sites and non-CTCF / Rad21 junction sites.

[0288] Figure 17M shows the median normalized end-to-end distance of domains as a function of genome size for chromatin domains containing a boundary locus with an interacting CTCF / Rad21 site, and domains without a boundary locus containing a CTCF / Rad21 site. Normalized end-to-end distance is defined as the domain end-to-end distance divided by the median distance between similar locus pairs within a single domain, separated by the same genome distance. Error bars indicate 95% confidence intervals induced by resampling.

[0289] Figure 24M shows the standard deviation matrix of interlocus spatial distances for Chr21. For each pair of regions, the standard deviation of the distance between the corresponding locus pairs in all single chromosome copies is shown.

[0290] Figure 24N shows box plots of the physical sizes of chromatin domains of different genome sizes in Chr21, as measured by the radius of gyration. For each genome size, the median (centerline), 25th–75th percentile (box), and 10th–90th percentile (whiskers) are shown. [Example 9]

[0291] Chromatin compartments within a single chromosome. Next, using a high-resolution view of the entire chromosome, we examined how chromatin loci in compartments A and B are arranged within a single cell. First, using the previously described algorithm (Figure 18A; Figure 25A), we determined the ensemble A / B compartment boundary using principal component analysis (PCA) of the Pearson-corrected matrix of proximity frequency maps for Chr21 derived from imaging data. The compartment boundary obtained from the imaging data was very similar to that determined from previously published ensemble Hi-C data (Figure 25A). Below, we used the compartment boundary obtained from the ensemble proximity frequency maps to assign A / B classifications to individual loci within individual cells.

[0292] The more than tenfold increase in resolution compared to previous studies allowed for a detailed view of the configuration of compartment A and compartment B loci within a single chromosome. A high degree of variation in the arrangement of A and B loci was observed between cells and between individual chromosome copies (Figure 18B). In some chromosomes, the A and B loci were separated within essentially non-overlapping spatial regions, while in others, substantial spatial overlap was observed between the A and B loci. Interestingly, compartment A loci within the same chromosome were sometimes separated into multiple "microcompartments" (Figure 18B).

[0293] To quantify the degree of spatial segregation of A and B loci on individual chromosomes, a local density-based method was devised, and the local densities of other A and B loci were calculated for each imaged locus (Figure 25B). As expected, compartment A loci tended to be surrounded on average by A loci, and the same was true for B loci (Figure 25C). A / B segregation scores were further defined for each individual chromosome based on the purity of the loci observed in the spatial volume containing the majority of A or B loci (Figure 18C). Complete physical segregation of A and B loci was expected to result in a segregation score of 1, and complete mixing of A and B loci was expected to result in a segregation score of 0.5 (see Example 19). For the majority of Chr21 copies in cells, the segregation score was observed to be considerably higher than the randomized control score (obtained by randomly shifting compartment boundaries along the genomic axis while keeping compartment size unchanged), centered around 0.5 (Figure 18C), indicating a tendency for A and B loci to be spatially separated within single cells. It is also noteworthy that the spatial segregation of A and B loci was often incomplete (Figure 18C). This potentially reflected incomplete spatial segregation of active and inactive chromatin, but could also be caused in part by intercellular variability of epigenetic modifications, which can make ensemble A / B compartment discrimination an incomplete surrogate for the active / inactive chromatin depiction within single cells. In particular, the degree of A / B segregation was found to be cell cycle-dependent: A / B segregation was stronger in G2 / S phase cells compared to G1 phase cells (Figure 25D), consistent with previous findings of stepwise establishment of A / B compartments during the cell cycle.

[0294] Chr21 is one of the smallest chromosomes, measuring only about 48 Mb, and was known to be divided into only a few consecutive A and B regions. To extend this finding to larger chromosomes and examine how common they are, we imaged chromosome 2 (Chr2), one of the largest chromosomes, which exhibits numerous transitions (approximately 50 transitions) between A and B compartments along its genomic sequence. Specifically, we tracked Chr2 by labeling and imaging 50 kb segments at 250 kb intervals along its genomic sequence. Using the same methodology as above, we called the A and B compartments within the p and q arms of the chromosome based on the imaging data (Figure 18D; Figure 25E) and observed quantitative agreement with the A and B compartments determined from previously published ensemble Hi-C data (Figures 25F-25G). At the single-chromosome level, we again observed a variety of spatial arrangements of the A and B loci, ranging from nearly complete spatial segregation between them to substantial spatial overlap (Figure 18E). Interestingly, some chromosomes exhibited a “sandwich” shape, with the A locus positioned between two layers of B loci, likely due to the preferential association of the B locus with the nearby nuclear lamina above and below the cell nucleus. Quantitatively, the A / B segregation score distribution across individual copies of Chr2 again showed an overall tendency for the A and B loci to segregate on individual chromosomes compared to randomized controls (Figure 18F). The degree of spatial segregation appeared to be less in Chr2 than in Chr21 (Figures 18C and 18F).

[0295] Figures 18A to 18I show the relationship between compartment structure and transcriptional activity in a single chromosome and local chromatin content. Figure 18A shows a Pearson correlation matrix for proximity frequencies normalized by genomic distance in Chr21, derived from our imaging data. Two loci are considered close if their distance is less than a 500 nm cutoff distance. The two bars below show the proximity frequency matrix (shown for compartments A and B) and A / B calling derived from G banding of each genomic locus in the chromosome.

[0296] Figure 18B shows a 3D rendering of individual copies of Chr21 in a single cell, with loci A and B shown as spheres. Flexible lines connect adjacent loci in the genomic sequence. The bars below show the A / B calling of each genomic locus in the chromosome, derived from the proximity frequency matrix. The scale bars are 1 micrometer.

[0297] Figure 18C shows the distribution of A / B segregation scores for individual copies of Chr21. To calculate the A / B segregation score, an A (or B) high-density volume is defined for each chromosome by thresholding the local A (or B) density so that 2 / 3 of the A (or B) loci are contained within the volume (note that A and B high-density volumes may overlap for chromosomes exhibiting spatial overlap between A and B loci). The purity of loci within the A (or B) high-density volume of a chromosome is defined as the proportion of all loci within the volume that are A (or B) loci, and the A / B segregation score for a chromosome copy is defined as the average purity of the A and B volumes. The histogram represents the distribution of A / B segregation scores for a randomized control, where the boundaries between consecutive A and B regions are randomly shifted along the genome sequence while keeping the number and size of A and B regions unchanged. There are approximately 7,500 chromosomes.

[0298] Figure 18D shows a Pearson correlation matrix of proximity frequencies normalized by genomic distance for the p and q arms of Chr2, derived from our imaging data, similar to Figure 18A, as well as the corresponding A / B calling and G banding.

[0299] Figure 18E shows a 3D rendering of individual copies of Chr2, similar to Figure 18B. The scale bar is 1 micrometer.

[0300] Figure 18F shows the distribution of A / B segregation scores for Chr2 in single cells, similar to Figure 18C. n = approximately 3,100 chromosomes.

[0301] Figures 25A–25G show ensemble A / B compartment analysis for Chr21 and Chr2. Figure 25A shows compartment calling based on principal component analysis for Chr21. The first principal component (PC1) is shown, calculated for the Pearson correlation matrix from proximity frequencies normalized by genomic distance derived from imaging (top) and ensemble Hi-C (bottom) experiments. PC1 values ​​greater than 0 correspond to compartment A, and PC1 values ​​less than 0 correspond to compartment B.

[0302] Figure 25B shows 3D renderings of compartment A locus (A locus) and compartment B locus (B locus), as well as the A / B density ratio in a single copy of Chr21. Left: A and B locus in a representative copy of Chr21. The A and B compartment callings from the ensemble proximity frequency map derived from the imaging are shown in the bars below. Right: The same chromosome, but each locus is color-coded according to its local A / B density ratio.

[0303] Figure 25C shows the mean A and B density scores for each imaged locus in Chr21, averaged across all imaged cells. The lower panel represents the A or B compartment calling for each locus from proximity frequency maps derived from the imaging.

[0304] Figure 25D shows a histogram of the distribution of A / B separation scores for individual copies of Chr21 in cells during the G1 and G2 / S phases of the cell cycle.

[0305] Figure 25E shows the imaging-derived proximity frequency matrix (left) and the ensemble Hi-C contact matrix (right) for Chr2. The Hi-C contact matrix is ​​binned at 50kb intervals, but only contacts from the imaged segments (selected at 250kb intervals) are shown.

[0306] Figure 25F shows the same PC analysis as Figure 25A, but for the p-arm (top) and q-arm (bottom) of Chr2.

[0307] Figure 25G shows the same mean A and B density analysis as Figure 25C, but with respect to Chr2. [Example 10]

[0308] Relationship between transcription and local A / B chromatin content. To test whether chromatin compartmentation correlates with active transcription in a single chromosome, oligonucleotide probes targeting the first introns of 86 genes present in Chr21 were designed, and sequential round hybridization was performed to image the nascent RNA transcripts of these genes, followed by chromatin tracing. Furthermore, to more accurately detect the spatial location of the genes, a 5kb genomic locus centered on the transcription start site (TSS) of each target gene was imaged. To prevent RNA probe binding to genomic DNA and vice versa, RNA probe hybridization was performed without thermal denaturation of the double-stranded genome, and the RNA molecules were digested using RNase treatment (a step also included when imaging chromatin) before chromatin tracing. Crosstalk between RNA and DNA signals was confirmed to be negligible when using this strategy (Figures 26A-26J).

[0309] Typically, a subset of the imaged genes exhibited transcriptional activity in any individual cell (Figure 18G). We examined how transcriptional activity correlated with the local chromatin environment. To characterize local A / B chromatin content, local densities near the A and B loci were calculated for each gene, and the ratio (hereinafter referred to as the A / B density ratio) was used as a metric for local enrichment of active chromatin. For approximately 80% of the genes tested, the local A / B density ratio in their TSS was found to be higher when the gene was actively transcribing than when the gene was not firing (Figure 18H). As a natural consequence, the firing rate of the genes also tended to be higher in cells where the gene's TSS had a higher local A / B density ratio (Figure 18I). These results indicated that the same gene tended to have higher transcriptional activity in cells exhibiting higher A locus enrichment and / or B locus de-enrichment in its vicinity. This increase in transcriptional activity may potentially be due to local enrichment of the transcriptional mechanism and / or de-enrichment of silencing factors. Alternatively, actively transcribing chromatin associated with the transcriptional mechanism may have a stronger tendency to interact with other active chromatin, considering that the transcriptional mechanism and cofactors can form condenses. These two possible mechanisms were found to be non-exclusive but capable of working cooperatively and reinforcing each other.

[0310] Figure 18G shows a 3D rendering of a single copy of Chr21, along with the transcription bursts of the measured genes. Spheres represent all detected nascent RNA bursts in this chromosome. The scale bar is 500 nm.

[0311] Figure 18H shows the change in A / B density ratio (measured as a logarithmic difference) at the transcription start site (TSS) of imaged genes between actively firing and non-firing states. For each gene, the median A / B density ratio is calculated at its TSS in both the chromosome in which the gene is firing and the chromosome in which it is not firing. The logarithmic differences of these values ​​for 84 genes imaged on Chr21 are ranked according to the magnitude of the change in their median A / B density ratio. 79% of the imaged genes showed an increase in A / B density ratio when they were actively firing compared to when they were non-firing.

[0312] Figure 18I shows the change in firing rate (measured as a logarithmic difference) of imaged genes when the local environment of the gene's TSS changes from a low (lower quartile) A / B density ratio to a high (upper quartile) A / B density ratio. The logarithmic differences of firing rates for 84 imaged genes on Chr21 are ranked according to the magnitude of the firing rate. 79% of the imaged genes showed higher firing rates when their TSS was in the upper quartile compared to the lower quartile of the A / B density ratio.

[0313] Figures 26A–26J show measurements of RNA and DNA FISH probe crosstalk. Figure 26A shows an exemplary cell and the fluorescence signals (bottom) of a FISH probe targeting nascent RNA in its nucleus (top) and gene (BRWD1) marked with DAPI.

[0314] Figure 26B is the same as Figure 26A, but concerns a different gene (SCAF4). The staining in Figures 26A and 26B follows the protocol described in the RNA FISH protocol in the section "Preparation of Cell Cultures and Primary / Encoding Probe Hybridization" of Example 19.

[0315] Figures 26C and 26D are identical to those in Figures 26A and 26B, respectively, except that the RNA FISH protocol was modified to include an additional RNase treatment step to remove cellular RNA before adding the FISH probe. The cells in Figures 26C and 26D were imaged under the same irradiation conditions as in Figures 26A and 26B, and their fluorescence signals are displayed with the same contrast as in Figures 26A and 26B.

[0316] Figure 26E shows the number of spots per cell with a signal-to-noise ratio greater than 3 for untreated and RNase-treated cells across the five measured genes.

[0317] Figure 26F shows an exemplary cell and its nucleus (top) marked with DAPI, and the fluorescence signal of a probe targeting the genomic locus (chr21: 15.2Mb–15.25Mb) (bottom).

[0318] Figure 26G is the same as Figure 26F, but concerns a different gene locus (chr21: 14.95Mb-15Mb). The staining in Figures 26F and 26G follows the DNA FISH protocol described in the section "Preparation of Cell Cultures and Primary / Encoding Probe Hybridization" of Example 19.

[0319] Figures 26H and 26I are identical to those in Figures 26F and 26G, respectively, except that the DNA FISH protocol was modified to omit the thermal denaturation step and thus remove accessible genomic DNA regions. Cells in Figures 26H and 26I were imaged under the same irradiation conditions as in Figures 26F and 26G, and their fluorescence signals were displayed with the same contrast as in Figures 26H and 26I.

[0320] Figure 26J shows the number of spots per cell with a signal-to-noise ratio greater than 3 for cells treated with a thermal denaturation step and cells in which this step was omitted. [Example 11]

[0321] The relationship between chromatin domains and compartments within a single chromosome. Next, we examined the interactions between single-cell chromatin domains and how these interactions correlate with compartment recognition. Due to the large size of Chr2 and the large number of compartment divisions within it, analysis on Chr2 was expected to provide more insights, and thus we focused on this chromosome.

[0322] While the majority of single-cell domains in Ch2 were “pure” domains containing all of either the A or B locus, a substantial proportion of single-cell domains crossed the ensemble A / B boundary and contained both the A and B loci (Figures 19A–19C). The presence of these “mixed” domains suggested that domain formation in single cells is not strongly related to the chromatin properties that determine compartment recognition, but that intercellular variability of epigenetic modifications may cause shifts in the active / inactive chromatin boundary in some single cells.

[0323] Next, we examined how domains interact with each other, focusing on how interdomain interactions depend on the A and B composition of the domains, as well as the genomic distance between domains. Domains were in contact with both short and long genomic segregations, and such contacts appeared as off-diagonal box features in the spatial distance map of individual chromosomes. Such contact patterns differed substantially between cells (Figure 19D). Despite this heterogeneity, domain-domain interactions were thought to be modulated by the A / B composition of their underlying chromatin (Figure 19E): the frequency of contacts between domains primarily containing B loci was, on average, higher than that between domains primarily containing A loci, and then higher than the frequency of contacts between domains biased towards chromatin with different A / B discrimination. This mean picture was consistent with a recently proposed hierarchy of A and B chromatin interaction strengths based on chromatin structure modeling of A / B compartmentalization measured by Hi-C and the overall arrangement of A and B loci measured by imaging.

[0324] Further examination of domain contact frequency as a function of genomic distance revealed a more complex picture. For simplicity, we focused on “pure” A and “pure” B domains containing single compartment-identifying loci. As expected, contact frequency decreased with genomic distance for domain pairs of all compositions (Figure 19F). However, contact frequency between B domain pairs (BB) was higher than that between A domain pairs (AA) at shorter genomic distances (up to approximately 75 Mb for Chr2), while AA domain contacts predominated over BB domain contacts at larger genomic segregations (Figure 19F). This result is consistent with the genomic distance dependence of AA-BB chromatin interactions reported in recent ensemble Hi-C studies and provides further insight into how preferential interactions between single-cell domains may give rise to these ensemble trends. In particular, at relatively large genomic distances, the probability of BB domain contact decreased to a level similar to that of A- and B-domain contact (AB), while the probability of AA domain contact remained higher than that of AB domain contact, even at large genomic segregations (Figure 19F). This resulted in a clear dominance of AA domain interactions at large genomic distances (Figure 19G). In addition, contacting domain pairs also showed varying degrees of spatial overlap; some domain pairs showed relatively superficial contact (insert in Figure 19F), while others showed strong mixing (insert in Figure 19H). Interestingly, compared to AA domain pairs, BB domains showed a substantially stronger tendency to form such mixed microspheres (Figure 19H).

[0325] Overall, these results suggested that preferential AA-BB domain interactions lead to spatial separation of chromatin compartments, and that the nature of AA-BB domain interactions differs. Differences in the nature of these interactions may be due to different molecular factors involved in AA-BB association. For example, heterochromatin factors such as HP1 are thought to be involved in BB interactions, while transcription activators or coactivators such as BRD4 and Mediator may be involved in active chromatin interactions. Further investigation is needed to determine whether these different molecular factors are responsible for the observed differences in the genomic distance dependence between AA-BB domain interactions and their tendency to mix.

[0326] Figures 19A–19H show the dependence of interdomain interactions on their A / B composition and genomic distance. Figure 19A, left, is a 3D rendering of a "mixed" chromatin domain containing both A and B loci, flanked by a "pure" domain containing only the B locus in a copy of Chr2 in a single cell. The scale bar is 500 nm. Right: The pairwise distance matrix of the same region is shown on the left. The bars and matrix below on the left show the A and B calling of the loci, and the contour lines highlight the boundaries of the chromatin domains. A / B calling is determined from the ensemble proximity frequency map of Chr2.

[0327] Figure 19B is the same as Figure 19A, but instead of a mixed domain, it concerns two pure domains, one containing the A locus entirely and the other containing the B locus entirely.

[0328] Figure 19C shows the distribution of the proportion of single-cell chromatin domains in Chr2 where the gene locus is the A locus.

[0329] Figure 19D shows the single-cell spatial distance matrix for two exemplary copies of Chr2. The first and third panels show the matrix for the two whole chromosomes, while the second and last panels show enlarged matrices for the regions highlighted in yellow in the first and third panels, respectively. The sidebar shows A / B compartment calling derived from the ensemble proximity frequency map.

[0330] Figure 19E shows the domain contact probabilities for domains with different A / B compositions in Chr2. The X and Y axes represent the proportion of loci within a domain that are A loci (0% corresponds to a pure B domain, and 100% corresponds to a pure A domain). Two domains are defined as in contact if their isolation score is less than 2. See Example 19 for the calculation of the isolation score.

[0331] Figure 19F shows the domain contact probabilities in Chr2 between two pure A domains (AA), two pure B domains (BB), and one pure A domain and one pure B domain (AB), plotted as a function of genomic distance between two interacting domains. The inset includes a 3D rendering of an exemplary domain pair exhibiting long-range interaction with an isolation score of 2. The scale bar is 500 nm.

[0332] Figure 19G is the same as Figure 19E, but concerns domain pairs with genomic distances greater than 80 Mb.

[0333] Figure 19H is the same as Figure 19F, but is limited to domain pairs exhibiting a high degree of mixing (defined by a low isolation score of less than 1). The inset includes a 3D rendering of an exemplary domain pair exhibiting long-range interaction with a high degree of mixing (isolation score = 1). The scale bar is 500 nm. [Example 12]

[0334] Chromatin imaging at the genome scale. The sequential imaging technique described above made it possible to obtain a high-resolution view of chromatin within individual chromosomes. This linear sequential imaging technique is very suitable for imaging chromatin structures that are equivalent to or smaller than diffraction-limited resolution. However, the number of imaged genomic loci increased only linearly with the number of imaging rounds in this technique. For genome-scale chromatin imaging, it was inferred that a much more efficient, non-linear scaling of the number of imaged loci would be possible with respect to the number of imaging rounds, since many genomic loci can be simultaneously resolved and localized in the nucleus.

[0335] To achieve this goal, we devised a combinatorial FISH technique, conceived from multiplexed error-robust FISH methods previously developed for transcriptome imaging, but including significant modifications specifically designed for chromatin imaging by considering both the polymeric nature of chromatin (i.e., adjacent loci in the genome sequence are spatially close) and the regional organization of chromosomes (i.e., different chromosomes tend to occupy separate spatial regions). To enable combinatorial imaging, each genomic locus was assigned a unique 100-bit binary barcode with 2 Hamming weights. That is, each barcode contained 2 "1" bits and 98 "0" bits (Figure 20A). The bit values ​​("1" or "0") in these barcodes determined the presence or absence of a signal for each locus over sequential imaging rounds. From these 100-bit Hamming weight 2 barcodes, a subset was further selected to encode target genomic loci to avoid simultaneously imaging spatially close chromatin regions with the same bit, optimizing barcode assignment so that loci with "1" bits at the same barcode location are maximally separated in genomic space (see Example 19). This strategy minimized detection errors caused by overlapping signals originating from neighboring chromatin loci. Furthermore, since the vast majority of possible 100-bit binary codes were invalid (i.e., not assigned to any target loci), this design allowed for the identification and discarding of detection errors, further improving measurement accuracy.

[0336] The barcodes were physically imprinted onto target genomic loci using a highly diverse library of encoding probes, each containing a target region for binding to one of the target loci and a readout sequence selected from 100 pre-designed readout sequences (Figure 20A). Each readout sequence corresponds to one of 100 bits, and the encoding probe set for each genomic locus (approximately 400 probes per locus) contains only two different readout sequences, corresponding to the two bits read as "1" in the barcode assigned to that locus. After encoding probe binding, the barcodes imprinted on the chromatin loci were detected by sequential hybridization of fluorescently labeled readout probes complementary to one of the 100 readout sequences (Figure 20A). In some cases, the adapter probe strategy described for high-resolution whole-chromosome tracing was also used. Two different adapter / readout probes were introduced per hybridization round and imaged in two color channels so that two bits were read in each hybridization round. This allowed us to image and identify approximately 1000 genomic loci using only 50 rounds of hybridization (Figures 20A–20C). This represented about one-tenth the number of hybridization rounds and thus one-tenth the experimental time compared to sequential imaging of the same number of loci using the same number of color channels. Since each chromosome in diploid cells has two homologs, homolog identification of the imaged loci was further assigned using a clustering algorithm that leverages the tendency of chromosomes to occupy different regions in each nucleus.

[0337] In this study, 1,041 genomic loci, each approximately 30 kb in size and uniformly encompassing 22 autosomes and X chromosomes, were selected for imaging in IMR-90 cells. Another requirement was that each chromosome contained at least 30 target loci, and therefore the number of loci to be imaged per chromosomal homolog ranged from 30 to 80, depending on the chromosome length. These 1,041 genomic loci in approximately 5,400 individual cells across five biological repeats were imaged with a detection efficiency of approximately 80% for each locus, resulting in the detection of approximately 1,700 chromatin loci in each cell (Figures 20D–20E). At the end of the combinatorial imaging process, a small subset of genomic loci were re-imaged using sequential imaging, one locus at a time. The displacement between the locus location determined by combinatorial imaging and the re-imaged location determined by sequential imaging was only about 50 nm (Figure 27A), which demonstrates both the high decoding accuracy of the combinatorial imaging method and the minimal sample degradation / deformation during the imaging process.

[0338] To obtain a population-average view of chromatin composition, the spatial distance between pairs of imaged chromatin loci was calculated in each cell, and then both the median distance and proximity frequency between loci pairs were determined across all imaged cells (Figure 20F; Figure 27B). The proximity frequencies between pairs of chromatin loci within the same chromosome, determined from the imaging data, showed a high correlation with the contact frequencies detected by Ensemble Hi-C, with a Pearson correlation coefficient of 0.89 (Figure 27C). Furthermore, the imaging results showed high reproducibility across independent biological replicates (Figure 27D).

[0339] By exploring the chromatin composition within individual cells, we found that chromosomes tend to occupy different regions within each cell (Figures 20F–20G), but also exhibit substantial overlap with one another (Figures 20G–20H). These results were consistent with and developed from observations from earlier imaging studies. Because these observations suggested a high degree of trans-chromosome interaction, further analysis focused on exploring these interactions.

[0340] Figures 20A–20H show genome-scale chromatin imaging by large-scale multiplexed combinatorial FISH. Figure 20A shows the imaging scheme. Target genomic loci were assigned error-robust barcodes, e.g., a subset of 100-bit binary barcodes with a Hamming weight of 2 (i.e., two of the 100 bits are read as "1"). The barcodes were imprinted onto the genomic loci using encoded oligonucleotide probes that recognize the locus, associate two different readout sequences with each locus, and correspond to the two bits read as "1" in the barcode assigned to the locus. Each bit is uniquely assigned to a readout sequence. Each locus is labeled with a total of 400 encoded probes, but only four are shown. By sequentially adding fluorescent readout probes complementary to the readout sequences and imaging, it is possible to determine the bits read as "1" at each locus and therefore the barcode identification of that locus.

[0341] Figure 20B shows representative images from multiple imaging rounds within the nucleus of a single cell. It shows the fluorescence signal of the chromatin locus from the readout probe and the signal of 4',6-diamidino-2-phenylindole (DAPI), which is used as a nuclear marker. The scale bar is 5 micrometers.

[0342] Figure 20C shows a magnified image (white box in B) of a small region centered on a single chromatin locus across all imaging rounds. The locus is identified based on two signaling readout probes (1 and 13). The scale bar is 300 nm.

[0343] Figure 20D is a 3D rendering of all detected chromatin loci (spheres) in a single IMR-90 cell, color-coded according to the chromosome to which they belong (chromosome-related indicators are shown below the image). Adjacent loci in the genomic sequence are connected by flexible lines. Approximately 1000 genomic loci are imaged.

[0344] Figure 20E shows the chromatin gene loci of the same cell as in Figure 20D, but highlights the two homologs of the chromosome shown.

[0345] Figure 20F shows the median distance matrix calculated from approximately 5,400 single cells. For each pair of loci, the median observed 3D spatial distance between the loci across all cells is presented.

[0346] Figure 20G shows an illustrative image illustrating the locations of multiple chromosomal regions in a single cell. Chromosomes are encoded as shown, and shaded regions represent the convex hulls surrounding all imaged loci. For clarity, only one homolog per chromosome is shown.

[0347] Figure 20H shows a spatial distance matrix for the same cell shown in Figure 20G. The spatial distances between each pair of chromatin loci are shown. The chromosome order is as shown below the matrix, and the two homologs of each chromosome are shown separately.

[0348] Figures 27A–27J show genome-scale imaging by combinatorial FISH: localization error, reproducibility, and comparison with Hi-C. Figure 27A shows the distribution of displacement between the localization of genomic loci measured during combinatorial imaging and that of the same loci individually re-imaged using sequential hybridization after combinatorial imaging was completed. Ten genomic regions in Chr6 were re-imaged across approximately 2000 cells. The median displacement was approximately 50 nm.

[0349] Figure 27B shows the proximity frequency matrix for all 1,041 genomic loci imaged by combinatorial FISH. The proximity frequency between pairs of loci was calculated by dividing the number of occurrences where the measured distance between loci was less than the 500 nm cutoff distance by the total number of measured distances between the two loci.

[0350] Figure 27C shows a correlation plot on the frequency of approach between pairs of intrachromosomal loci derived from our imaging data, binned at 500kb and centered on the target locus, and the number of contacts derived from ensemble Hi-C experiments. The Pearson correlation coefficient is 0.91. The available Hi-C data for IMR90 cells is limited regarding trans chromosome contacts, making a reliable comparison of trans chromosome interactions between our imaging data and Hi-C data impossible.

[0351] Figure 27D shows the correlation of pairwise distances between chromatin loci observed in two independent biological replicates of genome-scale imaging experiments. The Pearson correlation coefficient between replicates is 0.98. The upper right cloud represents pairwise distances across trans chromosomes, and the lower left cloud represents pairwise distances within chromosomes. [Example 13]

[0352] Enrichment of trans-chromosome A / B interactions. Next, we tested how trans-chromosome interactions depended on chromatin A / B compartment discrimination. Each of the imaged genomic loci was classified into compartment A and B based on published ensemble Hi-C data. Genomic-scale imaging data also allowed for fairly accurate A / B compartment calling, agreeing with calling based on Hi-C data by about 80% (Figure 27E), but Hi-C calling was used to classify A / B compartment discrimination of the imaged loci due to the higher genomic resolution of the ensemble Hi-C data. 38% of the imaged loci belonged to compartment A, while 62% belonged to compartment B. To examine whether the degree of trans-chromosome interactions differed for active and inactive chromatin, we rearranged genomic loci in a trans-chromosome proximity frequency matrix, placing all A loci next to each other, and then all B loci. This matrix showed that compartment A loci were, on average, more likely to interact trans-chromosomally with other compartment A loci than with individual compartment B loci (Figures 21A–21B), consistent with previous observations of trans-chromosomal interactions between active chromatins. In contrast, compartment B loci showed equal or lower trans-chromosomal affinity to each other than to compartment A loci (Figures 21A–21B). In other words, trans-chromosomal AA interactions appeared substantially more strongly than AB interactions, followed by a slightly stronger tendency than BB interactions. Similar trends were observed for various cutoff distances used to construct proximity frequency maps, although a sufficient number of cells were included in the analysis (Figures 27F–27H). Notably, this is in striking contrast to the overall hierarchy of cis interactions within the same chromosome, where BB interactions were more likely to form than AA interactions, followed by a stronger tendency than AB interactions (Figure 19E).Interestingly, however, the observed trend for trans chromosome interactions (AA>AB≈BB) was similar to that observed for cis chromosome interactions at large genomic distances in high-resolution Chr2 data (Figures 19F–19G). This trend was also observed for long-range cis chromosome interactions in genome-scale data, summed across all imaged chromosomes (Figure 21C).

[0353] Next, we examined the relationship between trans chromosome interactions and chromatin compartment discrimination at the single-cell level. Within individual cells, loci A and B adopted different spatial distributions, with loci A tending to localize more centrally in the nucleus than loci B, as expected (Figure 21D; Figure 28). To further characterize chromatin interactions in trans, we employed density-based methods similar to those previously presented, except that we considered only trans chromosome interactions here, for high-resolution whole-chromosome tracing. Briefly, for each imaged locus within each chromosome, we calculated the local densities of loci A and B derived from all other chromosomes in the same cell, and determined the ratio of these two densities (hereinafter referred to as the trans A / B density ratio) (Figures 21D-21E). This quantity provided a measure of local enrichment of trans chromosome-active chromatin near the locus. The majority (62%) of the imaged loci belonged to compartment B, and the overall bias regarding the A / B ratio was found to be less than 1. To control for this bias, the distribution of trans A / B density ratios observed for A and B loci was compared to the distribution obtained in randomized controls, where the A and B classification of the imaged loci was randomly shuffled among the imaged loci, while keeping the number of A and B loci unchanged. In particular, the trans A / B density ratio observed for A loci was substantially higher than the value observed for B loci, and subsequently higher than the value derived from the randomized controls (Figure 21E). These single-cell analyses further supported the view that trans chromosome interactions preferentially enhance interactions between active chromatin.

[0354] We further investigated whether the enrichment of trans-chromosome AA interactions required transcription. To address this, we inhibited transcription with alpha-amanitin treatment and found that this treatment did not result in a substantial reduction of enrichment for trans-chromosome AA interactions (Figures 29A-29C). This observation was consistent with and developed from previous studies showing that long-range and trans-chromosome interactions between activated betaglobin loci and other active chromatin loci are not inhibited by transcriptional inhibition.

[0355] Figures 21A–21E illustrate the enrichment of active-active chromatin interactions in trans chromosome interactions. Figure 21A shows the normalized trans chromosome proximity frequency matrix. It shows the proximity frequency between each trans chromosome locus pair (a pair of loci on different chromosomes), and locus pairs are considered close if their distance is less than the 500 nm cutoff distance. The loci are rearranged so that loci in compartment A appear first, followed by loci in compartment B, and so the upper left block represents the interactions between pairs of loci in compartment A, and the lower right represents the interactions between pairs of loci in compartment B. Each entry in the matrix is ​​normalized by the median proximity frequency of all locus pairs originating from the same pair of chromosomes, which constitute the basic level of changing interactions between pairs of chromosomes.

[0356] Figure 21B shows the distribution of trans chromosome proximity frequencies for pairs of A loci (AA; n=72,771 locus pairs), pairs of B loci (BB; n=193,753 locus pairs), and pairs containing one A and one B locus (AB; n=237,986 locus pairs), derived from the matrix shown in Figure 21A. The distribution is shown as a histogram in the upper panel and as a box plot in the lower panel, showing the median (centerline), 25th–75th percentile (box), and 5th–95th percentile (whiskers).

[0357] Figure 21C shows the median pro...

Claims

1. A method for visualizing the chromatin structure within an individual cell by imaging at least 100 distinct genomic loci within a single cell, a) Exposure of a sample containing chromatin with a population of multiple primary nucleic acid probes, wherein each population of nucleic acid probes hybridizes to a spatially separated genomic locus within the chromatin, and each genomic locus is assigned a valid codeword that forms an error-checking and / or error-correcting code space; b) Contacting each bound primary nucleic acid probe population with at least two separate fluorescent readout probes that hybridize to the read sequences of the primary nucleic acid probes, wherein each read sequence corresponds to one position of the codeword; c) Imaging the readout probe bound to the primary nucleic acid probe; and, d) Repeat steps b) and c) in one or more sequential hybridization and imaging rounds. Methods that include...

2. The method according to claim 1, wherein contact and imaging generate codewords corresponding to the binding of readout probes to read sequences of multiple nucleic acid probes.

3. The method according to claim 2, wherein genomic loci are identified based on a generated codeword and a corresponding valid codeword.

4. The method according to claim 1, wherein the three-dimensional structure of chromatin is visualized based on the spatial location of identified genomic loci.

5. The method according to claim 1, further comprising classifying the imaged genomic loci as active chromatin or inactive chromatin.

6. The method according to claim 1, further comprising imaging nascent RNA transcribed from a gene located at a genomic locus.

7. The method according to claim 6, wherein imaging of nascent RNA is performed using a pool of multiple primary nucleic acid probes, where each pool of nucleic acid probes hybridizes to a separate nascent RNA transcript, and each transcript is assigned a valid codeword having a Hamming distance of at least 2.

8. The method according to claim 6, further comprising imaging the nuclear structure.

9. The method according to claim 8, wherein the nuclear structure includes nuclear speckles and nucleoli.

10. The method according to claim 1, wherein at least 30 genomic loci are selected on each chromosome for imaging, and at least 100 distinct genomic loci are imaged in a single cell.

11. The method according to claim 1, wherein more than 1,000 genomic loci are imaged in individual cells.

12. The method according to claim 11, further comprising imaging more than 1,000 nascent RNAs transcribed from genes located at genomic loci.

13. The method according to claim 8, further comprising imaging a nuclear structure that places its chromatin composition in its natural structural and functional context.

14. The method according to claim 1, wherein the primary nucleic acid probe comprises a target sequence that hybridizes to a genomic locus and a readout sequence complementary to the readout probe, wherein each pool of nucleic acid probes comprises the same readout sequence.

15. The method according to claim 14, further comprising the use of an adapter oligonucleotide probe.

16. The method according to claim 1, wherein hybridization and imaging rounds are performed 50 or more times.

17. The method according to claim 1, wherein the size of the genomic locus is approximately 30 kb.

18. The method according to claim 1, wherein a 100-bit binary code having two or more Hamming weights is assigned to the chromatin and each genomic locus.

19. The method according to claim 18, wherein one, two, or three genomic loci are targeted for each hybridization and imaging round.

20. The method according to claim 18, wherein approximately equal numbers of genomic loci are imaged in each imaging round.

21. The method according to claim 20, wherein each genomic locus is imaged using one or two fluorescent labels.

22. The method according to claim 1, wherein the assigned codewords have an error-checking and / or error-correction code space having a Hamming distance for maximizing the distance between genomic loci on the same chromosome.

23. The method according to claim 22, wherein valid codewords in the code space are separated by a Hamming distance of at least 3, at least 4, at least 5, or at least 6.

24. The method according to claim 22, wherein a valid codeword is binary, and all valid codewords have the same number of "1" bits.

25. The method according to claim 1, wherein multicolor imaging is used in each round to enable simultaneous imaging and simultaneous determination of multiple readout probes associated with different signal-generating entities.

26. Following step a), the sample is exposed to one or more adapters in a round, where each adapter comprises a first portion substantially complementary to one of the readout sequences and a second portion containing one identification sequence; The method involves exposing a sample to one or more readout probes in one round to determine one or more identification sequences, wherein each readout probe comprises a first portion containing a sequence substantially complementary to one of the identification sequences and a second portion containing a signaling entity. The method according to claim 1.

Citation Information

Patent Citations

  • Systems and methods for determining nucleic acids

    US20170220733A1